<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Classi cation of Localization Utterances using a Spatial Ontology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohammad Fazleh Elahi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hui Shi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John A. Bateman</string-name>
          <email>batemang@informatik.uni-bremen.de</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kathleen M. Eberhard</string-name>
          <email>eberhard.1@nd.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Scheutz</string-name>
          <email>mscheutz@cs.tufts.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Psychology, University of Notre Dame</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Human-Robot Interaction Lab, Tufts University</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>SFB/TR8 Spatial Cognition, Universitat Bremen</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>13</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Dialogue systems for spatially situated tasks need to provide referential descriptions of spatially located objects and understand such descriptions from users. To construct such dialogue systems, it is useful to investigate how humans describe object locations in their immediate environment and how they ask about object locations in remote environments. In this paper, we address the semantic classi cation of the localization utterances found in the CReST corpus, which is a dialogue corpus of humans performing a cooperative, remote, search task. The aim is to explore the relation between speci c semantic con gurations and the dialogically and situationally embedded linguistic forms employed. Speci cally, we rst extracted di erent types of localization utterances from the corpus and then paired these with semantic categories provided by the linguistically motivated spatial ontology GUM. The paper concludes with a discussion of the characteristics of di erent types of localization expressions on the basis of spatial concepts and descriptions employed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The localization of objects in an indoor environment poses a considerable
challenge for dialogue-based intelligent systems, which are employed in a variety of
tasks, such as mapping and localization Kruij et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and indoor way nding
Cuayahuitl et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, the support for object localization in intelligent
systems is currently rather limited. For e ective interaction, such dialogue
systems should be able to deal with place-related knowledge acquired from verbal
expressions, such as \now you're in the long hallway", in which the place
\hallway" is described as \long". This requires them to support place-related
information extraction, representation and reasoning. Our aim here is to explore the
semantic classi cations of localization utterances, a fundamental e ort toward
the construction of sophisticated place-aware spoken dialogue systems.
      </p>
      <p>
        Spatial localization has two main aspects: location description and location
query. The recently collected human-human dialogue corpus CReST
(cooperative, remote, search task) of Eberhard et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] covers both communicative
aspects and, therefore, provides useful data. In addition, research investigating
the relation between linguistic expressions involving space and semantic
representations is of relevance (e.g., broad-coverage characterizations of semantics for
naturally occuring texts and dialogues [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). Hence, we will focus on the
relationship of natural linguistic expressions with a broader range of more formal
spatial characterizations capable of covering more of the semantics of naturally
occuring, task-based expressions, combining a data-driven and a semantic
classi cation approach.
      </p>
      <p>
        The Generalized Upper Model (GUM) is a general task and domain
independent \linguistically motivated ontology" that provides linguistic semantics
for spatial expressions. It serves as an intermediate \interface ontology"
mediating between linguistic forms and contextualized interpretations. In previous
work, Bateman et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] employed GUM to classify spatial relations in three
spatial language corpora: the Trains 93 Dialogues (Heeman &amp; Allen [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), the
HCRC Map Task (Anderson et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]), and the IBL Corpus (Lauria et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]),
and showed that the GUM spatial ontology provided a characterization of the
semantics that was of immediate use for several natural language processing
tasks. Here we follow this line of research and present the semantics of a set
of localization expressions from the CReST corpus using GUM, with focus on
static spatial con gurations and their elements.
2
      </p>
      <p>
        Corpus Analysis
The CReST corpus [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] of natural language dialogues was obtained from an
experiment involving humans performing a cooperative, remote, search task. The
experiment required individuals in a dyad to coordinate their actions via remote
audio communication in order to accomplish several tasks with target objects
(\colored boxes") that were scattered throughout an indoor search environment
(see Figure 1). Neither individual was familiar with the environment before the
experiment. One individual was designated as the director (D), and the other as
the searcher (S).
      </p>
      <p>The CReST corpus is particularly relevant for the localization problem
because the experiment directly involves several distinct localization subtasks. In
one subtask, the searcher was to report the locations of eight green boxes in the
environment and the director was to mark their locations on the map. Hence,
the director needed to learn about the boxes' locations through dialogue. In a
second subtask, the director was to direct the searcher to the cardboard box at
the furthest point in the search environment, while the searcher was to collect
blocks from blue boxes and put them into the cardboard box. Here, the director
needed to track the searcher in the environment (from dialogue only, with no
visual feedback) and then give further instructions to direct the searcher to the
blue boxes.</p>
      <p>In order to extract probable relatum and locatum from the dialogues, three
wordlists were created. The rst wordlist (BP) contains the words related to
the places in the environment (room, cubicle, o ce, hallway, doorway, etc), the
second (FO) words related to the xed objects (door, wall, steps/stairs, stage,
booth, etc), and the third (MO) words related to the movable objects
(computer, chair, table, shelf, cabinet, etc). The localization-related utterances of
the directors were rst extracted on a word-by-word basis through comparing
each word with those in the wordlists. These utterances were then processed
using a set of tools (e.g., tokenizer, part-of-speech tagger, and regular expression
based NP chunker) developed in NLTK toolkit1, in order to nd the meaningful
noun-phrase(s) (NP) of individual utterances. For instance, NP [the/DT little/JJ
tiny/JJ room/NN] is the meaningful phrase of the utterance \is it in the little
tiny room?".</p>
      <p>Analyzing the number of utterances of each group, we nd that rooms of the
wordlist BP and doors of the wordlist FO are most frequently used. Moreover,
we found that \small", \big",\tiny", " rst", etc. (see Table 1) are often used to
distinguish rooms.
1 http://nltk.org/
Group Examples
BP
FO
MO</p>
      <p>Descriptions from
the corpus
are you still in the smaller room? little, small, big, long, large,
okay now you're in the long hallway? tiny, initial, rst, third, etc
there's a open door in front of you right? open, close, single, double,
is there a door with a single door? small, two, etc
are there ling cabinets you see? ling cabinets,
are there two chairs at the desk? two chairs, etc
3</p>
      <p>
        Semantic Classi cation of Localization Expressions
The GUM ontology [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is a linguistically motivated ontology based on
grammatical evidence from a broad range of linguistic spatial expressions which provides
a uni ed account of spatial concepts and their relations. We use the GUM
concept SpatialLocating (SL) to de ne and classify the semantics of the localization
utterances in CReST for two reasons. First, GUM's SpatialModality (SM) covers
distance, direction and relationship between the locatum (L) and relatum (R)
which are crucial for localization expression. Second, GUM provides spatial
semantics to localization expressions regardless of contextual interpretations. As
can be seen from table 2, in GUM, \you were in the little cubicle, right?" and
\are there like a desk in front of the computer chair?" receive spatial
semantics SpatialLocating, where SpatialLocating is the concept that speci es the place
(as a placement relation) where an entity (as a locatum) is being positioned.
A placement relation can be a GeneralizedLocation (e.g., \on the table") or a
GeneralizedRoute (e.g., \through the hallway"). The GeneralizedLocation binds
together a relatum and a spatial relationship hasSpatialModality within a single
structured entity that may stand in a placement relation within a spatial con
guration. The utterance \There is an open door in front of you", for example, is
bound with the following GUM semantics:
Con guration:
Locatum:
Placement:
      </p>
      <p>SpatialLocating
an open door
GeneralizedLocation
hasSpatialModality: FrontProjectionExternal
relatum: you(person)</p>
      <p>Exploring the di erent forms of localization expressions, we found several
utterances (e.g, 19.5% utterances of SpatialLocating), in which the locatum (L) is
unknown and the relatum (R) is either an object or a person (e.g, the searcher).
As can be seen from table 2, the locatum of the sentence \what do you see
immediately to your left?" (Con guration: SpatialLocating, Locatum: unknown,
Relatum: person, SpatialModality: LeftProjectionExternal) is unknown. From the
analysis of spatial relations between a locatum and a relatum, we have found that
there are many utterances (e.g, 33.8% utterances of SpatialLocating) related to
the GUM concept Containment of a person (e.g., \Are you in a hallway or small
room?") or an object (e.g., \Is there like a room that has a bunch of desks with
tables and chairs?") in a spatial place (i.e., BP). This is because locating the
person rst and then the entities that can be seen from that place is essential for
solving localization problems in a search task. The GUM concepts
ProjectionRelation (e.g., \on the left"), MultipleDirectional (e.g., \it's in the bottom right or
bottom left?"), and GeneralDirectional (e.g., \on the opposite side of the ling
cabinets") are frequently used for directional relations in the corpus, together
with clauses containing words such as \face" or \look" for orientation; for
example: \Okay three steps and now when you look to your right there's another
open door." and \if you look outside the cubicle there should be a door to your
right is that correct?".</p>
      <p>
        Many utterances (e.g, 35.1% utterances of SpatialLocating) in CReST contain
what in GUM terms are called ComplexCon gurations, in which the director
describes two or more spatial locations in a single utterance. The speci cations
of such utterances usually combine several SpatialLocating con gurations with
conjunction (\SLCjSL") or disjunction (\SLDjSL") as shown in table 2. For
example, by saying \there's a cubicle on your right hand side and then straight
in front of you there's also a door?", the director required the searcher to con rm
a spatial setting which only matches with the room in the upper left of the setup
shown in Figure 1. Table 2 shows that the complex sentence can be divided into
two sentences: \there's a cubicle on your right hand side" and \there is also a
door straight in front of you" respectively.
In this paper, we analyzed location descriptions and location queries in the
CreST corpus which naturally includes di erent forms of localization utterances
in a dialogue context. First, to analyze the probable locatum and relatum, we
extracted localization utterances which are more frequent in the corpus by
classifying the environment into places (BP), xtures (FO) and objects (MO). Second,
we extracted descriptions of each group in order to explore the levels of
descriptions attached with locatum and relatum. Finally, we paired these utterances
with semantic categories provided by the GUM ontology to explore the speci c
semantic con gurations and linguistic forms employed in localization query. The
ndings include (1) the adjectives and clauses are used to describe the
locatum and relatum, (2) the entities of the locatum and relatum are often found
unknown or variable, (3) the spatial relations such as Containment and
Projectional are found frequently used, and (4) the query which describes two or more
spatial locations are used to uniquely identify a spatial setting in the
environment. Since linguistic semantics of localization expressions formalized in GUM
is domain-independent, the semantic classi cations and ndings of this research
o er a general conceptualization for relating place-sensitive natural expressions
to their spatial semantic interpretation and therefore, provide a sophisticated
foundation for the contextualization and generation of place-aware natural
expressions in situated dialogue systems, like DAISIE (see Ross &amp; Bateman [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>A. H. Anderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bader</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Bard</surname>
            , E. Boyle, G. Doherty,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Garrod</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Isard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kowtko</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>McAllister</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Sotillo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Thompson</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Weinert</surname>
          </string-name>
          .
          <source>The HCRC Map task Corpus. Language and Speech</source>
          ,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <volume>351</volume>
          {
          <fpage>366</fpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Bateman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Ross</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Tenbrink</surname>
          </string-name>
          .
          <article-title>A linguistic ontology of space for natural language processing</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>174</volume>
          (
          <issue>14</issue>
          ):
          <volume>1027</volume>
          {
          <fpage>1071</fpage>
          ,
          <year>September 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>H.</given-names>
            <surname>Cuayahuitl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Dethlefs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.-F.</given-names>
            <surname>Richter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tenbrink</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Bateman</surname>
          </string-name>
          .
          <article-title>A dialogue system for indoor way nding using text-based natural language</article-title>
          .
          <source>In Proceedings of the 11th International Conference on Intelligent Text Processing and Computational Linguistics (CICLing</source>
          <year>2010</year>
          )
          <article-title>: posters and short presentations</article-title>
          , Iasi, Romania,
          <year>2010</year>
          . March 21-27.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>K.</given-names>
            <surname>Eberhard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nicholson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kbler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gunderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Scheutz</surname>
          </string-name>
          .
          <article-title>The indiana "cooperative remote search task"(crest) corpus</article-title>
          .
          <source>In Proceedings of the Seventh Language Resources and Evaluation Conference (LREC 10)</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Heeman</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          .
          <source>The Trains 93 Dialogues. Trains Technical Note 94-2</source>
          , Computer Science Dept., University of Rochester, mar
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. G.
          <string-name>
            <surname>-J. M. Kruij</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zender</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Jensfelt</surname>
            , and
            <given-names>H. I. Christensen.</given-names>
          </string-name>
          <article-title>Clari cation dialogues in human-augmented mapping</article-title>
          .
          <source>In Proceedings of the 1st Annual Conference on Human-Robot Interaction (HRI'06)</source>
          , Salt Lake City,
          <string-name>
            <surname>UT</surname>
          </string-name>
          ,
          <year>March 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Lauria</surname>
          </string-name>
          , G. Bugmann,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kyriacou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          .
          <article-title>Training Personal Robots via Natural-Language Instructions</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          , pages
          <fpage>38</fpage>
          {
          <fpage>45</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>I.</given-names>
            <surname>Mani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Doran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Harris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hitzeman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Quimby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Richer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wellner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mardis</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. Clancy.</surname>
          </string-name>
          <article-title>SpatialML: annotation scheme, resources, and evaluation</article-title>
          .
          <source>Language Resources and Evaluation</source>
          ,
          <volume>44</volume>
          (
          <issue>3</issue>
          ):
          <volume>263</volume>
          {
          <fpage>280</fpage>
          ,
          <year>2010</year>
          . Special Issue LREC 2008:
          <article-title>Selected papers</article-title>
          , Edited by:
          <string-name>
            <given-names>N.</given-names>
            <surname>Ide</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Calzolari</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Ross</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Bateman</surname>
          </string-name>
          . Daisie:
          <article-title>Information state dialogues for situated systems</article-title>
          . In V. Matousek and P. Mautner, editors,
          <source>Text, Speech and Dialogue</source>
          , volume
          <volume>5729</volume>
          of Lecture Notes in Computer Science, pages
          <volume>379</volume>
          {
          <fpage>386</fpage>
          . Springer, Berlin / Heidelberg,
          <year>2009</year>
          . 12th International Conference, TSD 2009, Pilsen, Czech Republic,
          <source>September 13-17</source>
          ,
          <year>2009</year>
          . Proceedings.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>