<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TALP at GeoQuery 2007: Linguistic and Geographical Analysis for Query Parsing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Ferers´</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Horacio Rodgır´ uez</string-name>
          <email>horacio@lsi.upc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>TALP Research Center</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Design</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Software Department</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universitat Polietc`nica de Catalunya</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our experiments on the Geographical Query Parsing pilot-task for English at GeoCLEF 2007. Our system uses some modules of a Geographical Information Retrieval system presented at GeoCLEF 2006 [3] and modiefid for GeoCLEF 2007. The system uses deep linguistic analysis and Geographical Knowledge to perform the task.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Detect whether the query is geographic or no.
• Extract the WHERE component of the query.
• Extract the GEO-RELATION (from a set of predenfied types) if present.
• extract the coordinates (LAT-LONG) of the WHERE component. This process involves
sometimes a disambiguation task.</p>
      <p>As an example, see in Table 1 the information that has to be extracted from the query ”Discount
Airline Tickets to Brazil”.</p>
      <sec id="sec-1-1">
        <title>Field</title>
        <p>LOCAL
WHAT
WHAT-TYPE
WHERE
GEO-RELATION
LAT-LONG</p>
      </sec>
      <sec id="sec-1-2">
        <title>Content</title>
        <p>YES
Discount Airline Tickets</p>
        <p>INFORMATION</p>
        <p>Brazil</p>
        <p>TO
-10.0 -55.0</p>
        <p>In this paper we present the overall architecture of our Geographical Query Parsing system and
we describe brieyfl its main components. We also present the experiments, results and conclusions
in the context of the GeoCLEF’s 2007 GeoQuery pilot task.
2
2.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>System</title>
    </sec>
    <sec id="sec-3">
      <title>Description</title>
      <sec id="sec-3-1">
        <title>Overview</title>
        <p>The system architecture has two main phases that are performed sequentially: Topic Analysis and
Question Classification.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Topic Analysis</title>
        <p>The Topic Analysis phase has two main components: a Linguistic Analysis and a Geographical
Analysis.
2.2.1</p>
        <p>Linguistic Analysis
This process extracts lexico-semantic and syntactic information using the following set of Natural
Language Processing tools: i) TnT an statistical POS tagger [1], ii) WordNet lemmatizer
(version 2.0), iii) A Maximum Entropy based NERC trained with the CONLL-2003 shared task
English data set, iv) Spear1, a modified version of the Collins parser, which performs full
parsing and robust detection of verbal predicate arguments [2] (limited to three predicate arguments:
agent, direct object (or theme), and indirect object (benefactive or instrument).</p>
        <p>We pre-processed the data-set of 800.000 queries in English from a web search-engine with
linguistic tools to obtain the following data structures:
• Sent, which provides lexical information for each word: form, lemma, POS tag
(PennTree-Bank (PTB) tag-set for English), semantic class of NE, list of EWN synsets and,
finally, whenever possible the verbs associated with the actor and the relations between
some locations (specially countries) and their gentiles (e.g. nationality).
• Sint, composed of two lists, one recording the syntactic constituent structure of the question
(basically nominal, prepositional and verbal phrases) and the other collecting the information
of dependencies and other relations between these components.
•</p>
        <p>Environment. The environment represents the semantic relations that hold between the
different components identified in the question text. These relations are organized into an
ontology of about 100 semantic classes and 25 relations (mostly binary) between them. Both
classes and relations are related by taxonomic links. The ontology tries to reflect what is
needed for an appropriate representation of the semantic environment of the question (and
the expected answer). The environment of the question is obtained from Sint and Sent. A
set of about 150 rules was built to perform this task. Refer to [4] for details.</p>
        <p>1http://www.lsi.upc.edu/~surdeanu/spear.html
2.2.2</p>
        <p>Geographical Analysis
The Geographical Analysis is applied to the Named Entities from the queries that have been
classiefid as LOCATION or ORGANIZATION by the NERC module. A Geographical Thesaurus
is used to extract geographical information about these Name Entities. This component has
been built joining four gazetteers that contain entries with places and their geographical class,
coordinates, and other information:
1. GEOnet Names Server (GNS)2: a gazetteer covering worldwide excluding the United States
and Antarctica, with 5.3 million entries.
2. Geographic Names Information System (GNIS)3, contains 2.0 million entries about
geographic features of the United States and its territories. We used a subset of 39,906 entries
of the most important geographical names.
3. GeoWorldMap4 World Gazetteer: a gazetteer with approximately 40,594 entries of the most
important countries, regions, and cities of the world.
4. World Gazetteer5: a gazetteer with approximately 171,021 entries of towns, administrative
divisions and agglomerations with their features and current population. From this gazetteer
we added only the 29,924 cities with more than 5,000 unhabitants.</p>
        <p>A subset of the most important features from this thesaurus has been manually set using 46.132
places (including all kind of geographical features: countries, cities, rivers, states,. . . ). This subset
of important features has been used to decide if the query is geographical or not geographical.
2.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Question Classification</title>
        <p>The query classicfiation task is performed through the following steps:
•
•</p>
        <p>The query is linguistically preprocessed (as described in the previous subsection) for
getting its lexical, syntactic and semantic content. See in Figure 1 the results of the process
for the former example. What is relevant in the example is the fine grained classicfiation of
’Brazil’ as country, the existence of taxonomic information, both of location type
(administrative areas@@political areas@@countries) and location content (America@@South America@@
Brazil), and coordinates (-10.0 -55.0, useful for disambiguating the location and for
restricting the search area) and the existence of a shallow syntactic tree consisting on simple tokens
and chunks, in this case built by the composition of two chunks, a nominal chunk (’Discount
Airline Tickets’) and a prepositional one (’to Brazil’).</p>
        <p>Over the sint structure, a DCG like grammar consisting of about 30 rules developed manually
from the sample of GeoQuery and the set of queries of GeoCLEF 2006, is applied for obtaining
the list of topics (each topic represented by its initial and final positions) represented by a
triple &lt;geo-relation, initial position, nfial position &gt;). A set of features (consultive operations
over chunks or tokens and predicates on the corresponding sent structures) is used by the
grammar. The following features were available:
2GNS. http://gnswww.nima.mil/geonames/GNS/index.jsp
3GNIS. http://geonames.usgs.gov/geonames/stategaz
4Geobytes Inc.: Geoworldmap database containing cities, regions and countries of the world with geographical
coordinates. http://www.geobytes.com/.</p>
        <p>5World Gazetteer: http://www.world-gazetteer.com
– chunk features: category, inferior, superior, descendents.
– token features: num, POS, word form, lemma, NE 1 (general), NE 2 specicfi.
– token semantics: synsets, concrete and generic Named Entity type predicates (Named
Entity types include: location, person, organization, date, entity, property, magnitude,
unit, cardinal point, and geographical relation.
– head of the chunk features: num, POS, word, lemma, first NE, second NE.
– head of the chunk semantic features.
– left corner of the chunk: num, POS, word form, lemma, NE 1 (general), NE 2
(specific)
– left corner of the chunk semantics: WordNet synsets.</p>
        <p>
          See Figure 2 for a sample rule. The rule can be paraphrased as follows: a sentence is
composed by two chunks followed by a gap. The rfist chunk is of type ’npb’ or ’np’, i.e. it
is a nominal phrase, basic or complex, its head cannot be a Named Entity and the limits of
the chunk provide the limits of the topic. The second chunk is a ’pp’ and it provides the list
of locations.
parse_sentence(1, DS,CT,CNES) --&gt;
cc(DS,[(cc,[npb,np]),(hne1,[nil]),(ci,[LI]),(cs,[LS])],[],(
          <xref ref-type="bibr" rid="ref1 ref1">1,1</xref>
          )),
{CT=[(LI,LS)]},
cc(DS,[(cc,[pp]),(cd,[CD])],[],(
          <xref ref-type="bibr" rid="ref1 ref1">1,1</xref>
          )),
{parse_pp(_,DS,CNES,CD,[])},
parse_gap(_,DS).
• Finally from the result of step 2 several rule-sets are in charge of extracting: i) LOCAL, ii)
WHAT and WHAT-TYPE, iii) WHERE and GEO-RELATION, and iv) LAT-LONG data.
So, there are four rule sets with a total of 25 rules. Figure 3 presents an example of a WHAT
rule. The rule selects from the list of topics one containing a generic location (e.g. the noun
’city’). In this case the selected topic is assigned to WHAT and the WHAT TYPE set to
’Map’.
        </p>
        <p>classify_question_topic(X,WHAT,’Map’):sentence_2(X,(_,CT,_,_)),
CT\==[],
sentence_1(X,S),
member((LC1,LC2),CT),
range(LC1,LC2,R),
member(LC,R),
nth(LC,S,Tk1),
is_ generic_location(Tk1,_),
concatenate_words_pos(X,R,WHAT),!.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>We performed only one experiment for the GeoQuery2007 data set. The experiment consisted in
to extracting the requested data for the GeoQuery from a set of 800.000 queries.</p>
      <p>The results of the TALP system presented at the GeoCLEF’s 2007 GeoQuery Geographical
parsing task for English are summarized in Table 1. This table has the following IR measures for
each run: Precision, Recall, and F1.</p>
      <p>In the evaluation data set, a set of 500 queries had been labeled which are chosen to represent
the whole query set (800.000). The submitted results have been manually evaluated using a strict
criterion where a correct results should have all &lt;local&gt;, &lt;what&gt;, &lt;what-type&gt; and &lt;where&gt;
fields correct (the &lt;lat-long&gt; field was ignored in the evaluation).</p>
      <p>Our run achieved the following results: 0.2222 of Precision, 0.249 of Recall, and 0.235 of F1.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This is our first approach to deal with a geographical query parsing task. Our system for the
GeoCLEF’s 2007 GeoQuery pilot task is based on a deep linguistic and geographical knowledge
analysis. Although we need to do further evaluations to compare the system with other ones it
seems that our approach could be feasible for the task.</p>
      <p>As a future work we propose the following improvements to the system: i) further evaluations
of each problem subtask, ii) apply more sophisticated geographical desambiguation algorithms.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been supported by the Spanish Research Dept. (TEXT-MESS,
TIN2006-15265C06-05). Daniel Ferers´ is supported by a UPC-Recerca grant from Universitat Polietc`nica de
Catalunya (UPC). TALP Research Center is recognized as a Quality Research Group (2001 SGR
00254) by DURSI, the Research Department of the Catalan Government.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brants. TnT -</surname>
          </string-name>
          <article-title>a statistical part-of-speech tagger</article-title>
          .
          <source>In Proceedings of the 6th Applied NLP Conference (ANLP-2000)</source>
          , Seattle, WA,
          <string-name>
            <surname>United</surname>
            <given-names>States</given-names>
          </string-name>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Collins</surname>
          </string-name>
          .
          <article-title>Head-Driven Statistical Models for Natural Language Parsing</article-title>
          .
          <source>PhD thesis</source>
          , University of Pennsylvania,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferers´</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ageno</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Rodıgr</surname>
          </string-name>
          <article-title>´ uez. The GeoTALP-IR System at GeoCLEF-2005: Experiments Using a QA-based IR System, Linguistic Analysis, and a Geographical Thesaurus</article-title>
          . In C. Peters,
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mller</surname>
          </string-name>
          , and M. de Rijke., editors,
          <source>CLEF</source>
          , volume
          <volume>4022</volume>
          of Lecture Notes in Computer Science. Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ferers</surname>
          </string-name>
          ´,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kanaan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ageno</surname>
          </string-name>
          , E. Gonaz´lez, H. Rodgır´ uez, M. Surdeanu, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Turmo</surname>
          </string-name>
          .
          <article-title>The TALP-QA System for Spanish at CLEF 2004: Structural and Hierarchical Relaxing of Semantic Constraints</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          , and B. Magnini, editors,
          <source>CLEF</source>
          , volume
          <volume>3491</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>557</fpage>
          -
          <lpage>568</lpage>
          . Springer,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>