<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Spatial Event Language across Geographic Domains</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Klippel</string-name>
          <email>klippel@psu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sen Xu</string-name>
          <email>sen.xu@psu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jinlong Yang</string-name>
          <email>jinlong@psu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rui Li</string-name>
          <email>rui.li@psu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Geography, GeoVISTA Center, The Pennsylvania State University</institution>
          ,
          <addr-line>PA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present first results of an analysis of a corpus of linguistic descriptions that were collected in controlled experiments. This corpus and its analysis add to the body of knowledge on formal models for spatial language, language interpretation and generation. The experiments are grounded in qualitative formalisms (RCC and Intersection Models, IM) that have a long standing tradition as means to bridge formal and linguistic descriptions of space and spatial relations. Our experiments address dynamically changing spatial relations (movement patterns/geographic events). By keeping the formal spatial characterizations identical across experiments but changing the semantics (that is, we used movement patterns across seven different geographic domains such as a hurricane in relation to a peninsula, plus two geometric figure domains) we contribute to disentangling spatial and domain specific aspects of spatial (event) language. We briefly discuss here two aspects: First, we hand examine the corpus by selecting participants that show the same conceptual behavior as identified through RCC/IM; second, we analyze the domain specific sub-corpora to address similarities and dissimilarities between individual domains.</p>
      </abstract>
      <kwd-group>
        <kwd>Event language</kwd>
        <kwd>topology</kwd>
        <kwd>corpus analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Formal models of spatial language play an import role in several disciplines addressing
questions of (natural) language processing, natural language generation, the automatic description
of spatial scenes, or the design of unifying frameworks for multimodal information systems
and processing [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1–4</xref>
        ]. While we are in the age of spatio-temporal representation and reasoning,
the four-dimensional treatment of spatial language (and information in general) is still a hotly
debated topic. With respect to language, research shows that naming of events is more
challenging than naming of object [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and it is therefore not surprising that the insights gained
from describing static spatial relations linguistically need to be carefully evaluated and
extended to the dynamic domain. This contribution is addressing this issue by combining approaches
to model events employing qualitative spatial formalisms with linguistic analysis.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>
        We have developed an experimental paradigm that allows us to evaluate the influences of
domain semantics on the conceptualization of movement patterns as well as how movement
patterns are linguistically described. Here we focus on the linguistic descriptions. Our framework
is based on a topologically defined conceptual neighborhood graph [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6–8</xref>
        ]. Figure 1 provides an
overview of the different semantic domains that we have subjected to behavioral validation. In
a nutshell: We distinguish movement patterns on the basis of formal path characteristics as
identified by the conceptual neighborhood graph. The shortest path (in each scenario) is a
single topological relation, DC (disconnected), the longest path (in each scenario) is defined as
follows: DC-EC-PO-TPP-NTPP-TPP-PO-EC-DC. To give an example, a boat that never
touches or crosses an area of shallow water will always be disconnected (DC) from it. In
contrast, a boat that makes it completely across an area of shallow water will exhibit the long path
characteristics with the start and end relation being identical (DC). Our participants have to
perform a grouping task as a way to elicit conceptual knowledge. After performing this task,
participants are presented with the groups that they created again and are asked to provide
linguistic descriptions: a short label and a longer description detailing the grouping rational.
As rich as our data set is, the flexibility of natural language has made it a challenging task to
analyze it. We are presenting two approaches. First we had a look into linguistic descriptions
for specific paths. Here we show results from four domains, two from our translation
movement patterns (geometry and hurricane) and two from scaling movement patterns (geometry
and lake). This path (DC) could be described as a hurricane not making landfall or a lake not
flooding the house. Our goal was to analyze the variety of linguistic descriptions that
participants use to this relatively simple scenario. Table 1 provides some representative examples.
The important distinction that we made for both scenarios is whether the spatial information
(about the movement patterns) in these two scenarios is linguistically encoded using spatial
language, or, whether this information is encoded using domain specific language. The two
corresponding geometry scenarios serve as a reference as they obviously do not easily allow
for using domain semantics.
      </p>
      <p>With respect to the spatial language we find very diverse ways of conveying spatial
information. We do believe that this diversity is fostered by the fact that our research is addressing
geographic events / spatio-temporal information (rather than static spatial relations). Especially
in the hurricane example we find the following strategies: relative reference frames focusing
on the end relations of the geometric characteristics of figure and ground; qualitative
distancebased descriptions; negation of what the path does not do; absolute reference; (experiment)
context specific descriptions; explicit topological descriptions; intrinsic reference induced by
the movement. Interestingly, the explicit spatial descriptions in the lake scenario seem to be
less varied, indicating a potential difference between scaling and translation movement
patterns that are indistinguishable from a topological perspective.</p>
      <p>In both scenarios we also find descriptions that are encoding the spatial event in terms of
domain specific language (to different degrees). While, for example, a statement such as “no
hit”, “weak hit”, “no landfall” are still rather explicit, a statement such as “weak hurricane”
relies heavily on background knowledge of a scenario and is open to interpretation. In case of
the lake, the descriptions are much less varied, again, and in most cases refer to a flood not
happening.</p>
      <p>
        These findings led us to explore differences between the sub-corpora (the nine different
scenarios). First, some domain-corpus properties can be extracted using AntCont [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The
token occurrences are visualized using Wordle (http://www.wordle.net/), seeFigure 2.
      </p>
      <p>From the tag clouds in Figure 2, we can see that top frequent words are mostly related to
domain specific semantics. For example, city, desert, and island are referring to objects
illustrated in each scenario. It is not surprising that participants make use of the domain semantics
for reference to objects in the scenario, as it is a direct and succinct way to describe an object
and distinguish it from surroundings. However, for the topological change depicted in different
icons, participants would have to use more complicated descriptions such as verb phrases and
prepositional phrases. This is the reason that spatial language terms such as on, middle, at,
outside, left, right, through, and ended also appear prominent in the tag cloud. Our analytical
question is: given scenarios where only domain semantic is different, how different will the
descriptions be?</p>
      <p>
        In the next step, we used the Stanford POS tagger [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We investigate the most frequently
appearing nouns, verbs, adjectives, and prepositions:
 The most frequently appearing nouns are domain specific ones (see Figure 1). Domain specific nouns
with top frequency in one corpus are often never found in other corpora, such as tornado, oil, and
desert. Nouns that can be found across domains are common referral terms, such as side, icons,
middle, and bottom.
 Frequently appearing verbs seem to be not as domain specific as nouns. Common verbs are various
forms of be, end, touch, have, and go. However, there are a few verbs that appear frequently in some
corpora but not in others. Hit and miss frequently appear in the Cannon, Hurricane, and Tornado
corpus. Cover, expand and grow frequently appear in the Desert, Oil, and Lake corpus. It is not
surprising because hit and miss can be naturally used for describing “translation” while cover, expand
and grow naturally relate to “scaling”, which is the major difference in the above two corpora sets.
There are also cases where verbs are specific to a domain. Landed used as a verb frequently appears
in Cannon and Hurricane. Flooded used as a verb appears exclusively in Lake. Sailed exclusively
appears in Ship. This shows that domain semantic also influences verb usage, but not as explicit as
nouns. More examples are recede, retreat, leave, surrounded, shrink, and disappear.
 Adjectives seem to even less domain specific. Common adjectives across all corpora are middle, same,
right, and lower. The few cases where adjective are domain specific are the use of colors. Blue, grey,
and red appears as to provide additional referral information respectively in Ship, Geometry, and
Desert corpus. Exclusively in the Ship corpus, shallow, light and dark are frequently used to refer to
the boundaries or the center of the water body. Adjectives about size were also used. Large appear
more often in Oil.
 Prepositions are the least domain specific lexical category. Few prepositions are domain specific.
      </p>
      <p>Across frequently appears in translation scenarios but not in scaling ones.</p>
      <p>In sum, POS-tagging offers possibilities to examine linguistic usages by lexical categories.
Examining the nine corpora, frequently appearing nouns are highly domain specific; a few
verbs and adjectives are domain specific and a general difference in translation vs. scaling can
be found; prepositions are least domain specific, only the word “across” is found to be
differentiable between translation scenarios vs. scaling scenarios.</p>
      <p>
        The last analysis step here involves topic modeling [
        <xref ref-type="bibr" rid="ref11 ref12">11,12</xref>
        ]. It is a method for discovering
“topics” shared among documents within a corpus. It can be viewed as cluster analysis for
documents. Applying topic modeling to all documents (one for each participant, 20 documents
per scenario) in the nine corpora (180 documents in total), we can evaluate whether documents
might be clustered based on their domain. Mallet (Machine Learning for LanguagE Toolkit)
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] is used to realize topic modeling. Setting the “number of topics” to be nine, we can see if
the nine topic models correlate with the nine domains (scenarios). Because topic models are
data-driven and don’t imply any predefined knowledge, we want to compare the topic
modeling result with domain semantics and see if they are comparable to each other. Each topic is
defined by the keywords appearing most frequently and most distinctively.
      </p>
      <p>Table 1 shows the keywords that identify each topic model. Unsurprisingly, domain specific
nouns are distributed across topic models. These topic models can be used to evaluate the
probability of one document (descriptions created by one participant) being associated with a
specific topic model (ideally catching the domain). Assigning the most probable topic ID to a
document allows for using topic models for document classification. To evaluate the
correlation between topics and domain semantics further, we use the already built topic models to
classify each document. The results and evaluations are shown in Table 2.</p>
      <p>Topic ID 0 1 2 3 4 5 6 7 8 Out of 20 documents from each
Domain domain, we evaluated the proportion
DGeesoemrtetry_translation 116 104 00 10 11 00 21 01 20 of documents being classified into
Lake 0 0 19 0 0 0 1 0 0 the same topic model, which ideally
CHGaeunorrnmioceantnrye_scaling 030 004 000 131 3016 1014 011 0110 100 (sthhoisulcdorcreosrpreosnpdoenndcetowotrhkeeddeoxmcaeipnt
Oil 5 0 1 3 0 0 11 0 0 for tornado and cannon, where most
SThoirpnado 00 00 00 00 03 00 12 105 019 of both are assigned to Topic ID 7).
The bold numbers in Table 2 shows
the topic model (see also Table 1) that most documents from a domain are assigned to. As
shown in Table 2, it is reasonable to match each topic ID to one domain semantic and the
matching proportion (sum of diagonal cells divided by total) is 70.56%. Cross-examining the
domain semantic with keywords from corresponding topic models (see Table 1) we find that a
large proportion of documents are classified correctly.</p>
      <p>However, the above matching of topic models and domain semantics may be skewed by the
high volume of domain specific nouns. Hence, as a comparison, we removed all the domain
specific nouns from all corora and rebuilt the topic models.</p>
      <p>Two observations are important: In the first part of this paper we showed an analysis by hand
that allows for relating a qualitative formal description of a movement pattern to a linguistic
description. The linguistic descriptions are varied and participants used manifold strategies to
characterize formally identical movement patterns. However, we seem to be able to clearly
reveal domain specific differences, especially if we look into whether or not domain semantics
is present. In the second part of this paper we tried to use this insight and compared the
documents from each domain (one document with all linguistic descriptions per participant, 20
documents in each domain). We found that figure and ground (moving entity and reference
entity) are the dominating linguistic features used and that these nouns allow for classifying
documents largely correctly. However, once we remove these obvious, domain specific
features, classification and identification of documents becomes very inaccurate despite the
differences we found in the first part.</p>
      <p>There could be a number of reasons for this. Instead of comparing all documents of a
particular domain, which contains linguistic descriptions of several, topologically distinguishable
paths, we may need a finer granularity for the analysis. For example, we could extract all DC
descriptions from all domains and focus only on these. Likewise, we could extract all
descriptions for movement patterns that could be labeled across in the translation scenarios and
expand-and-retreat in the scaling scenarios. We could perform this analysis for all topologically
equivalent movement patterns that we used to design our experiments.</p>
      <p>It also could be that the topic modeling approach we used needs refinement. Topic models
make use of terms and co-occurrences with documents to discover topics. It is an effective
method for knowledge discovery from large corpora without predefined knowledge. However,
we are specifically looking for spatial language usage in this study. In order to reduce the
influence of domain specific nouns, we use a crude method which is removing the domain
specific nouns. Integrating predefined knowledge (in our case, specific target language and
contexts) into topic models would allow an analysis to focus on certain term usages, which would
enhance the capability of topic modeling.</p>
      <p>To sum up, we presented a first exploratory analysis of a corpus that is the result of the
conceptualization of movement patterns in different semantic domains. The unique aspect of our
experiments is that grounding the design in qualitative spatial representation and reasoning
frameworks allows for keeping the spatial information identical across domains only changing
the semantic (domain specific) context. We are hopeful that this corpus can contribute to a
better understanding of the relation between formal/computation models and spatial language
across different domains.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bateman</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Language and space. a two-level semantic approach based on principles of ontological engineering</article-title>
          .
          <source>International Journal of Speech Technology</source>
          <volume>13</volume>
          ,
          <fpage>29</fpage>
          -
          <lpage>48</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Galton</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Spatial and temporal knowledge representation</article-title>
          .
          <source>Earth Science Informatics</source>
          <volume>2</volume>
          ,
          <fpage>169</fpage>
          -
          <lpage>187</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kordjamshidi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Otterlo</surname>
          </string-name>
          , M. von, Moens, M.-F.:
          <article-title>From language towards formal spatial calculi</article-title>
          . In: Ross,
          <string-name>
            <given-names>R.J.</given-names>
            ,
            <surname>Hois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kelleher</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . (eds.)
          <article-title>Computational Models of Spatial Language Interpretation (CoSLI</article-title>
          ) Workshop at Spatial Cognition 2010,
          <article-title>Mt</article-title>
          . Hood, Oregon, pp.
          <fpage>17</fpage>
          -
          <lpage>24</lpage>
          . CEUR Workshop Proceedings (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ross</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hois</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelleher</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . (eds.):
          <article-title>Computational Models of Spatial Language Interpretation (CoSLI</article-title>
          ) Workshop at Spatial Cognition 2010,
          <article-title>Mt</article-title>
          . Hood, Oregon.
          <source>CEUR Workshop Proceedings</source>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gentner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boroditsky</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Individuation, relativity, and early word learning</article-title>
          . In: Bowerman,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Levinson</surname>
          </string-name>
          , S.C. (eds.)
          <article-title>Language acquisition and conceptual development</article-title>
          , pp.
          <fpage>215</fpage>
          -
          <lpage>256</lpage>
          . Cambridge Univ. Press, Cambridge (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Egenhofer</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Taha</surname>
            ,
            <given-names>K.K.</given-names>
          </string-name>
          :
          <article-title>Reasoning about gradual changes of topological relationships</article-title>
          . In: Frank,
          <string-name>
            <given-names>A.U.</given-names>
            ,
            <surname>Campari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Formentini</surname>
          </string-name>
          ,
          <string-name>
            <surname>U</surname>
          </string-name>
          . (eds.)
          <article-title>Theories and methods of spatio-temporal reasoning in geographic space</article-title>
          , pp.
          <fpage>196</fpage>
          -
          <lpage>219</lpage>
          . Springer, Berlin (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Freksa</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Temporal reasoning based on semi-intervals</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>54</volume>
          ,
          <fpage>199</fpage>
          -
          <lpage>227</lpage>
          (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Randell</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cui</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohn</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>A spatial logic based on regions and connections</article-title>
          .
          <source>In: Proceedings 3rd International Conference on Knowledge Representation and Reasoning</source>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>176</lpage>
          . Morgan Kaufmann, San Francisco (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Anthony</surname>
          </string-name>
          , L.:
          <source>AntConc (version 3.2.2)</source>
          . Waseda University, Tokyo, Japan (
          <year>2011</year>
          ). available from http://www.antlab.sci.waseda.ac.jp/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Feature-rich part-of-speech tagging with a cyclic dependency network</article-title>
          .
          <source>In: Proceedings of HLT-NAACL</source>
          <year>2003</year>
          , pp.
          <fpage>252</fpage>
          -
          <lpage>259</lpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lafferty</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          :
          <article-title>Topic models</article-title>
          . In: Srivastava,
          <string-name>
            <given-names>A.N.</given-names>
            ,
            <surname>Sahami</surname>
          </string-name>
          , M. (eds.)
          <article-title>Text mining</article-title>
          . Classification, clustering, and applications, pp.
          <fpage>71</fpage>
          -
          <lpage>93</lpage>
          . CRC Press/Taylor &amp; Francis, Boca Raton,
          <string-name>
            <surname>Fla</surname>
          </string-name>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Steyvers</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffiths</surname>
            ,
            <given-names>T.L.</given-names>
          </string-name>
          :
          <article-title>Probabilistic topic models</article-title>
          . In: Landauer,
          <string-name>
            <given-names>T.K.</given-names>
            ,
            <surname>McNamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.S.</given-names>
            ,
            <surname>Dennis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kintsch</surname>
          </string-name>
          , W. (eds.)
          <article-title>Handbook of Latent Semantic Analysis</article-title>
          . Lawrence Erlbaum, Mahwah, NJ (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>McCallum</surname>
            ,
            <given-names>A.K.: MALLET.</given-names>
          </string-name>
          <article-title>A machine learning for language toolkit (</article-title>
          <year>2002</year>
          ). http://mallet.cs.umass.edu/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>