<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MIRACLE at GeoCLEF Query Parsing 2007: Extraction and Classification of Geographical Information</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Lana-Serrano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Villena-Román</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Miguel Goñi-Menoyo</string-name>
          <email>josemiguel.goni@upm.es</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Politécnica de Madrid</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Carlos III de Madrid.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DAEDALUS - Data</string-name>
          <email>jvillena@daedalus.es</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Decisions</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Language</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of MIRACLE research consortium at the Query Parsing task of GeoCLEF 2007. Our system is composed of three main modules. First, the Named Geo-entity Identifier, whose objective is to perform the geo-entity identification and tagging, i.e., to extract the “where” component of the geographical query, should there be any. This module is based on a gazetteer built up from the Geonames geographical database and carries out a sequential process in three steps that consist on geo-entity recognition, geo-entity selection and query tagging. Then, the Query Analyzer parses this tagged query to identify the “what” and “geo-relation” components by means of a rule-based grammar. Finally, a two-level multiclassifier first decides whether the query is indeed a geographical query and, should it be positive, then determines the query type according to the type of information that the user is supposed to be looking for: map, yellow page or information. According to a strict evaluation criterion where a match should have all fields correct, our system reaches a precision value of 42.8% and a recall of 56.6% and our submission is ranked 1st out of 6 participants in the task. A detailed evaluation of the confusion matrixes reveal that some extra effort must be invested in “user-oriented” disambiguation techniques to improve the first level binary classifier for detecting geographical queries, as it is a key component to eliminate many false-positives.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Linguistic Engineering, classification, geographical IR, geographic entity recognition, gazetteer, semantic
expansion, Wordnet.</p>
      <p>The goal of Geographical Information Retrieval (GIR) is to deal with those information retrieval problems that
contain some kind of spatial awareness, i.e., that include geographical references (georeferences) which are
essential for the meaning of the query, for example: “find me nice and cheap hotels near Madrid”. It is a complex
task because of its strong dependence on geographical information resources, which tend to be incomplete and
inexact. Moreover, geographical information is mainly arranged in a tree-like hierarchy, therefore queries
usually imply a multilevel search (for example: “give me documents about villages in Northern Spain”). Finally,
additional translation problems arise when dealing with multiple languages, due to the lack of specific and
specialized translation resources in a worldwide domain.</p>
      <p>GeoCLEF is the cross-language geographic retrieval track that runs as part of the Cross Language Evaluation
Forum (CLEF) campaign, whose aim is to provide with the necessary framework in which to evaluate GIR
systems for search tasks involving both spatial and multilingual aspects. This year, apart from the traditional
task, GeoCLEF 2007 offered the Query Parsing task.</p>
      <p>A geographic query is usually composed of three components, “what”, “geo-relation” and “where”. The
keywords in “what” indicate what users want to find; “where” refers to their geographic area of interest; and
“geo-relation” stands for the relationship between “what” and “where”. For instance, in the first example, “what”
would be “nice and cheap hotels”, “where” would be “Madrid”, and “geo-relation” would be “NEAR”. Note that
“Madrid” is itself ambiguous and can refer not only to the capital of Spain or the autonomous region where the
city of Madrid is located, but also other cities or administrative divisions in United States, Philippines, Mexico,
Argentina, Equatorial Guinea, Colombia, Dominican Republic, Sweden…
The key problem for GIR is to understand how to parse and extract those key components from the queries. This
is the objective of the Query Parsing task. Participants where given a set of 800,000 untagged queries and had to
detect whether each query was a geographic query or not, and, should the result be positive, had to extract the
three components: “where” (with its corresponding latitude/longitude), “geo-relation” (normalized into a
predefined relation type such as IN, NEAR, FROM, TO, NORTH_OF…) and “what” (categorized into a type of
request: “map”, “yellow page” or “information”).</p>
      <p>
        The MIRACLE team is a research consortium formed by research groups of three different universities in
Madrid (Universidad Politécnica de Madrid, Universidad Autónoma de Madrid and Universidad Carlos III de
Madrid) along with DAEDALUS, a small/medium size enterprise (SME) founded in 1998 as a spin-off of two of
these groups and a leading company in the field of linguistic technologies in Spain. MIRACLE has taken part in
CLEF since 2003 in many different tracks and tasks, including the main bilingual, monolingual and cross lingual
tasks [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as well as in ImageCLEF, Question Answering,WebCLEF and GeoCLEF [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]tracks.
This paper describes the MIRACLE participation at the Query Parsing task of GeoCLEF 2007. In the following
sections, we will first give an overview of the architecture of our system. Afterwards we will elaborate on the
different modules. Finally, the results will be presented and analyzed.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. System Description</title>
      <p>
        The objective of this module is to perform the geo-entity identification and tagging, i.e., to extract the “where”
component of the query, should there be any. It is composed of two main components: a geo-entity parser based
on a gazetteer, i.e. a database with geographical resources that constitutes the knowledge base of the system.
Our gazetteer is built up from the Geonames geographical database [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], available free of charge for download
under a creative commons attribution license. It contains over 8 million geographical names with more than 6.5
million unique features about 2.2 million populated places and 1.8 million alternate names. Those features
include a unique identifier, the resource name, alternative names (in other languages), county/region,
administrative divisions, country, continent, longitude, latitude, population, elevation and timezone. All features
are categorized into one out of 9 feature classes and further subcategorized into one out of 645 feature codes.
Geonames integrates geographical data (such as names of places in various languages, elevation or population)
from various sources, mainly the Geonet Names Server (GNS) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] gazetteer of the National Geospatial
Intelligence Agency (NGA), the Geographic Names Information System (GNIS) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] gazetteer of the U.S.
Geographic Survey, the GTOPO30 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] digital elevation model for the world developed by United States
Geological Survey (USGS) and Wikipedia, among others.
      </p>
      <p>For our purposes, all data was loaded and indexed in a MySQL database, although not all fields (such as time
zone or elevation) were to be used: the relevant fields are UFI (unique identifier), NAME_ASCII (name),
NAME_ALTERNATE (alternate names), COUNTRY, ADM1 and ADM2 (administrative region where the
entity is located), FEATURE_CLASS, FEATURE_TYPE, POPULATION, LATITUDE and LONGITUDE. To
simplify the queries, each row is complemented with the expansion of country codes (ESÆ Spain) and/or state
codes (NCÆ North Carolina) –when applicable. The final database uses 865KB.</p>
      <sec id="sec-2-1">
        <title>The geo-entity parser carries out the following tasks:</title>
        <p>
</p>
        <p>
          Geo-entity recognition: identifies named geo-entities [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] using the information stored in the gazetteer,
looking for candidate named entities matching any substring of one or more words [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] included in the
query and not included in a stopword (or stop-phrase) list [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>The stopword list is mainly automatically built by extracting those words that are both common nouns
and also georeference entities, assuming that the user is asking for the common noun sense (for
example, “Aguilera” –for Christina Aguilera– or “tanga” –thong). Specifically we have used lexicons
for English, Spanish, French, Italian, Portuguese and German, and have selected words that appear at
least with a certain frequency in the query collection. The final stopword list contains 1712 entries.
Geo-entity selection: The selected named geo-entity will be the one with the longest number of words
and, if the same, the one with higher score. The score is computed according to the type of geographic
resource (country, region, county, city…) and its population, as shown in the following table.
Feature type
Capital and other big cities
Political entities
Countries
Other cities
Other
For all cities, if country/state
name/code is also in the query</p>
        <p>Those values where arbitrarily chosen after different manual executions and subsequent analysis.
</p>
        <p>Query tagging: expands the query with information about the selected entity: name, country, longitude,
latitude, and type of geographic resource.</p>
        <p>The output of this module is the list of queries in which a possible named geo-entity has been detected, along
with its complete tagging. For example:</p>
        <p>Query| score|ufi|entity|state (code)|country (code)|latitude|longitude|feature_class|feature_type
airport {{alicante}} car rental week|2693959|2521976|Alicante||Spain (ES)|38.5|-0.5|A|ADM2
bedroom apartments for sale in {{bulgaria}}|10000000|732800||Bulgaria (BG)|43.0|25.0|A|PCLI
hotels in {{south lake tahoe}}|123925|5397664|South Lake Tahoe|California (CA)|United States
(US)|38.93|-119.98|P|PPL
helicopter flight training in southwest {{florida}}|100100000|4920378|Florida|Indiana (IN)|United
States (US)|40.16|-85.71|P|PPL

Observe that the geo-entity is specifically marked in the original query, enclosed between double curly brackets,
to help the following module to identify the rest of the components of the geographical query.
2.2.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Query Analyzer</title>
      <sec id="sec-3-1">
        <title>It consists of two subsystems: This module parses each previously tagged query to identify the “what” and “geo-relation” components of a geographical query, sorting out the named geo-entity detected by the previous module, enclosed between curly brackets.</title>
        <p>Geo-relation identifier: identifies and qualifies spatial relationships supported by a regular expression
rules based. Its output is the input list of queries expanded with information related to the identified
“geo-relation”.</p>
        <p>For instance, continuing with the previous examples, the output would be the following:
Query|geo-relation|entity|state|country|country (code)|latitude|longitude|feature_class|feature_type
airport {{alicante}} car rental week|NONE|Alicante||Spain|ES|38.5|-0.5|A|ADM2
bedroom apartments for sale #@#IN#@# {{bulgaria}}|IN|Bulgaria||Bulgaria|BG|43.0|25.0|A|PCLI
hotels #@#IN#@# {{south lake tahoe}}|IN|South Lake Tahoe|California|United
States|US|38.93|119.98|P|PPL
helicopter flight training in #@#SOUTH_WEST_OF#@# {{florida}}|SOUTH_WEST_OF|Florida|
Indiana|United States|US|40.16|-85.71|P|PPL</p>
      </sec>
      <sec id="sec-3-2">
        <title>Observe that the geo-relation is also marked in the original query. Concept identifier: analyses the output of the previous step and extracts the “what” component of a geographical query applying manually defined grammar rules based on the identified “where” and “georelation” components.</title>
        <p>2.3.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Query Type Classifier</title>
      <p>Finally, the last step is to decide whether the query is indeed a geographical query and, should it be positive, to
determine the type of query, according to the type of information that the user is supposed to be looking for:
Map type: users are looking for natural points of interest, like rivers, beaches, mountains, monuments…
Yellow page type: businesses or organizations, like hotels, restaurants, hospitals, etc.</p>
      <p>Information type: users are looking for text information, like news, articles, blogs, and so on.</p>
      <sec id="sec-4-1">
        <title>The process is carried out by a two level classifier:</title>
        <p>1. First level: a binary classifier to determine whether a query is a geographical or a non-geographical
query. This simple classifier is based on the assumption that a query is geographical if the “where”
component is not empty.
2. Second level: a multi-classification rule-based classifier to determine the type of geographical query.</p>
        <p>The multi-classifier treats the tagged queries as a lexicon of semantically related terms (words,
multiwords and query parts).</p>
        <p>The classification algorithm applies a knowledge base that consists on a set of manually defined
grammar rules, including nouns and grammatically related part-of-speech categories as well as the type
of geographical resource. The different valid lemmas are unified using Wordnet synsets.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3. Results</title>
      <p>For the evaluation, multiple human editors labeled 500 queries that were chosen to represent the whole query set.
Then all the submitted results were manually compared to those queries following a strict criterion where a
match should have all fields correct.</p>
    </sec>
    <sec id="sec-6">
      <title>Precision(1)</title>
      <p>0.428</p>
    </sec>
    <sec id="sec-7">
      <title>Recall(2)</title>
      <p>According to the task organizers, our submission achieved the best performance out of the 8 submissions of this
year, which was a good reward for our hard work.</p>
      <p>As participants in the task were provided with the evaluation data set, we have further evaluated our submission
to separately study the results for each component of the geographical queries and also the performance
level-bylevel of the final classifier.</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions and Future Work</title>
      <p>We have however some disagreements with the evaluation data provided by the organizers. Although some of
them may be actual errors, most are due to the complexity and ambiguity of the queries. Table 7 shows some
examples of queries that have been classified as geographical by our system but have been evaluated as
falsepositives. In fact, we think that it would be almost impossible to reach a complete agreement in the parsing or
classification for every case among different human editors. The conclusion to be drawn from this is that the
task to analyze and classify queries is very hard without a previous contact and without the possibility of
interaction and feedback with the user.
The analysis of the confusion matrixes for the multiclassifier that are calculated over the topics correctly
classified by the first level classifier shows that the probability that a geographical query is classified as “Yellow
Page” is very high. This could be related to the uneven distribution of topics (almost 50% of the geographical
queries belong to this class). In addition, “Information” type queries have a very low recall. These combined
facts point out that the classification rules have not been able to establish a difference between both classes. We
will focus on this issue in future participations. Moreover, we will try to incorporate some “user-oriented”
disambiguation techniques to improve the first level binary classifier, as it is a key component to eliminate many
false-positives.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>This work has been partially supported by the Spanish R&amp;D National Plan, by means of the project RIMMEL
(Multilingual and Multimedia Information Retrieval, and its Evaluation), TIN2004-07588-C03-01; and by the
Madrid’s R&amp;D Regional Plan, by means of the project MAVIR (Enhancing the Access and the Visibility of
Networked Multilingual Information for the Community of Madrid), S-0505/TIC/000267.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Charniak</surname>
            ,
            <given-names>Eugene. A</given-names>
          </string-name>
          <string-name>
            <surname>Maximum-Entropy-Inspired Parser</surname>
          </string-name>
          .
          <source>In Proceedings of NAACL-2000</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>[2] Geonames geographical database</article-title>
          . On line http:// www.geonames.
          <source>org [Visited</source>
          <volume>14</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Global</given-names>
            <surname>30 Arc-Second Elevation Data Set</surname>
          </string-name>
          . On line http://eros.usgs.gov/products/elevation/gtopo30.
          <source>html [Visited</source>
          <volume>14</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Goñi-Menoyo</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>González-Cristóbal</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Villena-Román</surname>
            ,
            <given-names>J. MIRACLE</given-names>
          </string-name>
          at
          <string-name>
            <surname>Ad-Hoc</surname>
            <given-names>CLEF</given-names>
          </string-name>
          2005:
          <article-title>Merging and Combining without Using a Single Approach</article-title>
          .
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the Cross Language Evaluation Forum</source>
          <year>2005</year>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2005</year>
          , Vienna, Austria, Revised Selected Papers (Peters,
          <string-name>
            <surname>C.</surname>
          </string-name>
          et al.,
          <source>Eds.). Lecture Notes in Computer Science</source>
          , vol.
          <volume>4022</volume>
          , Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Lana-Serrano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Goñi-Menoyo</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>González-Cristóbal</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          <article-title>MIRACLE at GeoCLEF 2005: First Experiments in Geographical IR</article-title>
          .
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the Cross Language Evaluation Forum</source>
          <year>2005</year>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2005</year>
          , Vienna, Austria, Revised Selected Papers (Peters,
          <string-name>
            <surname>C.</surname>
          </string-name>
          et al.,
          <source>Eds.). Lecture Notes in Computer Science</source>
          , vol.
          <volume>4022</volume>
          , pp.
          <fpage>920</fpage>
          -
          <lpage>923</lpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Goñi-Menoyo</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>González-Cristóbal</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lana-Serrano</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Martínez-González</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>MIRACLE's AdHoc and Geographic IR approaches for</article-title>
          <source>CLEF 2006: 7th Workshop of the Cross Language Evaluation Forum</source>
          <year>2006</year>
          ,
          <article-title>CLEF 2006</article-title>
          , Alicante, Spain, Revised Selected Papers (Peters,
          <string-name>
            <surname>C.</surname>
          </string-name>
          et al.,
          <source>Eds.). Lecture Notes in Computer Science.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] University of Neuchatel.
          <article-title>Page of resources for CLEF (Stopwords, transliteration</article-title>
          , stemmers …). On line http://www.unine.ch/info/clef
          <source>[Visited</source>
          <volume>18</volume>
          /07/2006].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>U.S. Geological</given-names>
            <surname>Survey</surname>
          </string-name>
          . On line http://www.usgs.
          <source>gov [Visited</source>
          <volume>14</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>U.S.</given-names>
            <surname>National Geospatial Intelligence Agency</surname>
          </string-name>
          . On line http://www.nga.
          <source>mil [Visited</source>
          <volume>14</volume>
          /08/2007].
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>