<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SINAI at GeoCLEF 2006: Expanding the topics with geographical information and thesaurus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manuel Garc a-Vega</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel A. Garc a-Cumbreras</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Algorithms</institution>
          ,
          <addr-line>Languages, Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Information Retrieval, Geographic Information Retrieval</institution>
          ,
          <addr-line>Named Entity Recognition, GeoCLEF</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>L. Alfonso Ureaea-L pez</institution>
          ,
          <addr-line>JosØ M. Perea-Ortega</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of JaØn</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2027</year>
      </pub-date>
      <abstract>
        <p>This paper describes the rst participation of the SINAI (Intelligent Systems of Access Information) group of the University of JaØn in GeoCLEF 2006. We have developed a system made up of three main modules. The rst one is the translation subsystem, that works with queries into Spanish, Portuguese and Deutsche. The second one is the query expansion subsystem, that integrates a Named Entity Recognizer, a Gazetteer, a Thesaurus expansion module and a Geographical information module. The last subsystem is the Information Retrieval module, that works with collections and queries into English, and returns the result le. We have made several runs, that combines these modules to resolve the monolingual and the bilingual tasks. The results obtained shown that the use of geographical and thesaurus information for query expansion does not improve the retrieval, but this is the rst step to try to improve the system in the future.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The main objective of our rst participation in GeoCLEF have been the study of the problem
of this task, and the development of a system that returns relevant documents.</p>
      <p>The Cross-Language Geographic Information Retrieval (GIR) system developed at the
University of JaØn has been designed to retrieve relevant documents that contain geographic tags.</p>
      <p>For this reason, our system consist of several modules: Translation, Named Entity
RecognitionGazetteer, Geographical Information Subsystem and Thesaurus Expansion Subsystem.</p>
      <p>This paper is organized as follows: section 2 describes the whole system and each module of
the system in detail. Then, in the section 3 experiments and results are described.</p>
      <p>Finally, the conclusions about our participation in GeoClef 2006 are expounded, and some
future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System Description</title>
      <p>We propose a Geographical Information Retrieval System that is made up of ve subsystems:
² Translation Subsystem: is the query translation module. This subsystem translates the
queries into the other languages.
² Named Entity Recognition-Gazetteer Subsystem: is the query geo-expansion module.</p>
      <p>This subsystem uses the geographical information from Geographical Information module.
² Geographical Information Subsystem: is the module that stores the geographical data.</p>
      <p>This information has been obtained from Geonames1 gazetteer.
² Thesaurus Expansion Subsystem: is the query expansion module using an own
Thesaurus.
² IR Subsystem: is the Information Retrieval module. We have used the LEMUR IR
system2.</p>
      <p>The combination of these subsystems gives the results of the di erent runs, and will be shown
in section 3.
2.1</p>
      <sec id="sec-2-1">
        <title>Translation Subsystem</title>
        <p>The purpose of the Translation Subsystem is to translate the queries or topics that are not in
English.</p>
        <p>This module is used for the following bilingual tasks: Spanish-English, Portuguese-English and
German-English.</p>
        <p>For the translation we have used an own module, called SINTRAM (SINai TRAnslation
Module), that works with several online Machine Translators, and implements several heuristics. In
this case we have used an heuristic that joins the translation of a default translator (the one that
we indicate, depending of the pair of languages) with these words that have another translation
(using another translator).</p>
        <p>SINTRAM works with the following online translators:
² Epals: available at http://www.epals.com
² Prompt: available at http://translation2.paralink.com
² Reverso: available at http://www.reverso.net
² Systran: available at http://www.systransoft.com
1http://www.geonames.org
2http://www.lemurproject.org
The main goal of NER-Gazetteer Subsystem is to detect and recognize the entities in the queries, in
order to expand the topics with geographical information. We are only interested in geographical
information, so we have used only the locations detected by the NER module. The location term
includes everything that is town, city, capital, country and even continent.</p>
        <p>As we will see in the next section, the information about locations is loaded previously in the
Geographical Information Subsystem, that it is related directly to the NER-Gazetteer Subsystem.</p>
        <p>Figure 1 describes the system architecture with this relation. The NER-Gazetteer Subsystem
recognizes the entities and provides this information to the Geographical Information module.</p>
        <p>We have used the NER system that GATE3 provides.</p>
        <p>The basic operation of NER-Gazetteer module is the following:
² The rst step is the preprocessing phase. Each query is preprocessed using a tokenizer, a
sentence splitter and a POS tagger. The NER system we have used needs this information
in order to improves the named entity detection and recognition.
² The second step is the extraction of the title of each topic and the submission to the NER.</p>
        <p>The result is saved to another labelled topic le with the location entities labelled.
² The last step is the detection of the geographical places, that the NER module have not
detected. For this proposal we have use a Gazetteer, included also in GATE. We also</p>
        <p>include this information in an expanded topic, using an XML label.</p>
        <p>The NER-Gazetteer Subsystem generates some labelled topics, base on the original one, adding
the locations.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Geographical Information Subsystem</title>
        <p>
          The objective of this module is to expand the locations of the topics, using geographical
information. We have used automatic query expansion [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], a simple expansion that consists in adding
terms to the query. We have obtained the geographical information from Geonames gazetteer, a
free resource that provides geo-data such as geographical names and postal codes. Its database
contains over six million entries for geographical names whereof 2.2 million cities and villages.
        </p>
        <p>Geonames integrates geographical data such as names, altitude, population and others from
various sources.</p>
        <p>Some examples of queries that receives this module are:
² Find the capital of a country whose population is greater than X.
² Find ve cities from a country whose population is greater than X.
² Find the country name of a city.
² Find the latitude and longitude from a location.</p>
        <p>When a location is recognized by the NER subsystem we look for in the Geographical
Information Subsystem.</p>
        <p>In addition, it is necessary to consider the spatial relations found in the title ( near to( , within
X miles of( , north of( , south of( , etc.). Depending on the spatial relations, the search in the
Geographical Information subsystem is more or less restrictive.</p>
        <p>We also have to verify if the location is a city, a country or a continent. Depending on the
location type, the expansion will become of the following way:
² If the location is a continent, we expand with the capitals of countries that belong to that
continent, and with capitals if the population is greater than a number of habitants (one of
ours parameters). The expansion is not very large in order to avoid noise in the recovery
process.
² If the location is a country, we expand with the ve most important cities (with greater
population) of that country.
² If the location is a city or capital, rst we veri ed if there is some spatial relation in the
topic. If exists we use the latitude and longitude information to nd other relevant locations
and we expand the topic with them. If there is not some spatial relation, we expand the
topic only with the name from the country to which city belongs.</p>
        <p>Finally we add the locations given back by the Geographical Information Subsystem to the
topic title.</p>
        <p>To adjust the parameters of this module, for instance the number of habitants to consider a
relevant capital or the number of cities to expand, we have made experiments with the GeoCLEF
2005 framework.
2.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Thesaurus Expansion Subsystem</title>
        <p>A collection of thesauri was generated from the GeoCLEF training corpus. We were looking for
words with a very high rate of document co-location. These words will be treated like synonyms
and added to the topics.</p>
        <p>
          For that, we generated an inverse le with the entire corpus. The le has a row for each
di erent word of the corpus. Following each word appear all the current word frequencies for each
corpus le. These rows can be treated with the standard TF.IDF [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] for words comparing test.
        </p>
        <p>We probe this method with the GeoCLEF 2005 corpus and we found that a cosine similarity
great than 0.9 between words was the rate that obtain best precision/recall results.</p>
        <p>The same procedure was applied to the 2006 corpus. The thesauri collection was generated for
all the names of the topics and all the thesauri words were added to its topic. The Figure 2 shows
the calculated thesauri for the topics GC033 and GC034. We can see the topic code and the pairs
word-similarity.
The English collection dataset has been indexed using LEMUR IR system. It is a toolkit4 that
supports indexing of large-scale text databases, the construction of simple language models for
documents, queries, or subcollections, and the implementation of retrieval systems based on language
models as well as a variety of other retrieval models.</p>
        <p>The English collection include a variety of topics and geographical regions form news stories
between 1994 and 1995.</p>
        <p>
          Previously the English collection provided for GeoClef 2006 have been preprocessed, using the
English stopwords list and the Porter stemmer [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>Each topics set, the monolingual expanded and the bilingual translated and also expanded, is
run to LEMUR.</p>
        <p>
          One parameter for each experiment is the weighting function, such as Okapi [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] or TF.IDF.
Another is the use or not of Pseudo-Relevant Feedback (PRF) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <sec id="sec-3-1">
        <title>Our baseline experiment is the following:</title>
        <p>² We use the original English topics set
² We preprocess each topic (stopper and stemmer)
² Topics without expansion (geographical or thesaurus)
² Information Retrieval without PRF
4The toolkit is being developed as part of the Lemur Project, a collaboration between the Computer Science
Department at the University of Massachusetts and the School of Computer Science at Carnegie Mellon University.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experiment</title>
        <p>sinaiEnEnExp1 (best result)
sinaiEnEnExp2
sinaiEnEnExp3
sinaiEnEnExp4
sinaiEnEnExp5
SINAI has participated in monolingual task with ve experiments and the o cial results are shown
in Table 1.</p>
        <p>Best results were obtained when using our baseline experiment, without adding no type of
query expansion (experiment sinaiEnEnExp1 ), only we preprocess each topic with stopper and
stemmer and information retrieval process uses Okapi with feedback like weighting function. We
considered all tags (title, description and narrative) in information retrieval process. The
experiment sinaiEnEnExp2 is the same that the previous one but considering only title and description
tags from topics for information retrieval. We use Okapi with feedback like weighting function in
all experiments.</p>
        <p>In the experiment sinaiEnEnExp3 we expand only the title of topic with geographical
information and considered only title and description tags from topics for information retrieval with
LEMUR system.</p>
        <p>In the experiment sinaiEnEnExp4 we expand the title and the description of topics with
thesaurus information, considering only title and description tags in information retrieval process.</p>
        <p>In the experiment sinaiEnEnExp5 we expand the title and the description from topics with
geographical and thesaurus information. We considered only title and description tags for
information retrieval.
3.2</p>
        <sec id="sec-3-2-1">
          <title>Bilingual tasks</title>
          <p>In bilingual task we have participated with a total of ve experiments: two experiments for
German-English task and three experiments for Spanish-English task. The o cial results are
shown in Table 2.</p>
          <p>For German-English task we submit the experiment sinaiDeEnExp1 in which we have not
expanded the title or the description of topics, only we have translated the topic and we preprocess
it with stopper and stemmer. The retrieval process uses Okapi with feedback like weighting
function too. We considered all tags (title, description and narrative) for information retrieval
process in this experiment. We submit the experiment sinaiDeEnExp2 for German-English task
too. It is the same that previous one but only considering the title and description tags for
information retrieval.</p>
          <p>For Spanish-English, in the experiment sinaiEsEnExp1, we have not expanded the topics, only
we have translated it and we preprocess it with stopper and stemmer. We considered all tags
(title, description and narrative) for information retrieval process in this experiment. We submit
the experiment sinaiEsEnExp2 for Spanish-English task too. It is the same that previous one but
only considering the title and description tags for information retrieval.</p>
          <p>Finally, in the experiment sinaiEsEnExp3, we expand the title of topics with geographical
information, considering only the title and description tags for information retrieval.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future work</title>
      <p>In this paper, we have presented the experiment carried out in our rst participation in the
GeoCLEF campaign. We have only tried to verify if the topic expansion with geographical and</p>
      <sec id="sec-4-1">
        <title>Experiment</title>
        <p>sinaiDeEnExp1
sinaiDeEnExp2
sinaiEsEnExp1
sinaiEsEnExp2
sinaiEsEnExp3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Query Language German German Spanish</title>
        <p>Spanish
Spanish
thesaurus information increases the e ectiveness of the information retrieval process. Evaluation
results show that the use of geographical and thesaurus information does not improve the retrieval.
But this is the rst step for improving the system in the future.</p>
        <p>Several reasons exist to explain the worse results obtained with the expansion of topics:
² The NER used sometimes did not work well, because in various topics some entities are
recognized and other no. For the future we will try with another NERs.
² In topics, sometimes, appear compound locations like New England, Middle East, Eastern
Bloc, etc. that are not in Geographical Information Subsystem. Would be interesting to
create rules to control this.
² Depending on spatial relation in topic, we could improve the expansion, testing so that cases
work better to add more locations or less.</p>
        <p>Therefore, we will try to improve the NER-Gazetteer Subsystem and the Thesaurus Expansion
Subsystem to obtain one better query expansion.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments References</title>
      <p>This work has been supported by Spanish Government with grant TIC2003-07158-C04-04.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Sanchis-Arnal</surname>
          </string-name>
          .
          <article-title>A wordnet-based query expansion method for geographical information retrieval</article-title>
          .
          <source>Working Notes for the CLEF 2005 Workshop</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mart</surname>
          </string-name>
          nez-Santiago,
          <string-name>
            <surname>M.C.</surname>
          </string-name>
          <article-title>D az-</article-title>
          <string-name>
            <surname>Galiano</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Garc</surname>
            a-Vega,
            <given-names>M.T.</given-names>
          </string-name>
          <string-name>
            <surname>Mart</surname>
            n-Valdivia, and
            <given-names>L.A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureaea-L pez</surname>
          </string-name>
          . Sinai on clef 2001:
          <article-title>Calculating translation probabilities with semcor</article-title>
          .
          <source>Advances in Cross-Language Information Retrieval. Editor Carol Peters. Lecture Notes in Computer Science</source>
          . Springer-Verlag, pages
          <fpage>185</fpage>
          <lpage>192</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mart</surname>
          </string-name>
          nez-Santiago,
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Garc</surname>
          </string-name>
          a-Cumbreras,
          <string-name>
            <surname>M.C.</surname>
          </string-name>
          <article-title>D az-</article-title>
          <string-name>
            <surname>Galiano</surname>
            , and
            <given-names>L.A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureaea-L pez</surname>
          </string-name>
          . Sinai at clef 2004:
          <article-title>Using machine translation resources with mixed 2-step rsv merging algorithm</article-title>
          .
          <source>CLEF 2004 (Cross Language Evaluation Forum)</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mart</surname>
          </string-name>
          nez-Santiago,
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Garc</surname>
          </string-name>
          <article-title>a-Cumbreras, Arturo Montejo-RÆez, and</article-title>
          <string-name>
            <given-names>L.A.</given-names>
            <surname>Ureaea-L pez</surname>
          </string-name>
          .
          <article-title>Sinai at clef 2005: Multi-8 two-years-on and multi-8 merging-only tasks</article-title>
          .
          <source>Lecture Notes in Computer Science</source>
          . Springer-Verlag,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mart</surname>
          </string-name>
          nez-Santiago,
          <string-name>
            <given-names>M.T.</given-names>
            <surname>Mart</surname>
          </string-name>
          n-Valdivia, and
          <string-name>
            <given-names>L.A.</given-names>
            <surname>Ureaea-L pez</surname>
          </string-name>
          . Sinai on clef 2002:
          <article-title>Experiments with merging strategies</article-title>
          .
          <source>Advances in Cross-Language Information Retrieval. Editor Carol Peters. Lecture Notes in Computer Science</source>
          . Springer-Verlag, pages
          <fpage>187</fpage>
          <lpage>197</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Mart</surname>
          </string-name>
          nez-Santiago, Arturo Montejo-RÆez, M.C.
          <article-title>D az-</article-title>
          <string-name>
            <surname>Galiano</surname>
            , and
            <given-names>L.A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureaea-L pez</surname>
          </string-name>
          . Sinai at clef 2003:
          <article-title>Decompounding and merging</article-title>
          .
          <source>In CLEF 2003 - Workshop of the Cross-language Evaluation Forum</source>
          , Trondheim, Norway,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>An algorithm for su x stripping</article-title>
          .
          <source>In Program 14</source>
          , pages
          <fpage>130</fpage>
          <lpage>137</lpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.E.</given-names>
            <surname>Robertson</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Walker</surname>
          </string-name>
          .
          <article-title>Okapi-Keenbow at TREC-8</article-title>
          .
          <source>In Proceedings of the 8th Text Retrieval Conference TREC-8, NIST Special Publication 500-246</source>
          , pages
          <fpage>151</fpage>
          <lpage>162</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Buckley</surname>
          </string-name>
          .
          <article-title>Improving retrieval performance by relevance feedback</article-title>
          .
          <source>Journal of American Society for Information Sciences</source>
          ,
          <volume>21</volume>
          :
          <fpage>288</fpage>
          297,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton and M. J. McGill</surname>
          </string-name>
          .
          <article-title>Introduction to Modern Information Retrieval</article-title>
          .
          <string-name>
            <surname>McGraw-Hill</surname>
          </string-name>
          , London, U.K.,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>