<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UAIC: Participation in LAGI Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adrian Iftene</string-name>
          <email>adiftene@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>UAIC: Faculty of Computer Science, “Alexandru Ioan Cuza” University</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The LogCLEF track was launched in 2009 with aim to analyse and classify the user's queries and had two tasks LAGI (Log Analysis and Geographic Query Identification) and LADS (Log Analysis for Digital Societies). In this edition from 2009, we built a system in order to participate in the LAGI task. The system uses GATE or Wikipedia like external resources in order to identify geographical entities in user's queries. Because, the results obtained using these resource are comparable, the main advantage of using GATE resources comes from the short duration of execution in comparison with using of Wikipedia resources. A brief description of our system is given in this paper.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1 Introduction
1.
2.</p>
      <p>
        Tumba! - a Portuguese web search engine
The European Library (TEL) - on line search for materials in various
libraries in Europe.
Our system searches only in the subset in English logs (which are the majority of the
logs). Organizers force participating groups to respect some rules. The rules respected
by our system were:
1. A query is a geographical query if and only if it is bounded geographically.
2. A place term can be any country, a city or town, mountain, province or region
from GATE
        <xref ref-type="bibr" rid="ref1">(Cunningham et al., 2001)</xref>
        or if it is described as a place in
Wikipedia (Portuguese Wikipedia for Tumba! and English Wikipedia for our
English subset of TEL).
3. A candidate place term can map to more than one possible meaning in
      </p>
      <p>Wikipedia.
4. A place term can occur in a title (of a book, movie, team, etc.), but the title
itself (if a different text span from the place) is not to be tagged.
5. Capitalization (upper and lower case) in the query is ignored, as it is used
inconsistently in the queries.
6. Wildcards ('*') are ignored.
7. If some words of a query can be interpreted as forming a phrase, this will be
preferred over interpreting those words as isolated words put in the same
query.</p>
      <p>In order to participate in LAGI, additional to using Wikipedia resources offered by
organizers, we built another resource with geographical name entities starting from
GATE resources. This resource was loaded by our program in cache and it is used
after that in identification of geographical resources. The Figure 1 presents the system
architecture.</p>
      <p>GATE
resources</p>
      <p>Portuguese
Wikipedia</p>
      <p>English
Wikipedia</p>
    </sec>
    <sec id="sec-2">
      <title>Test Data</title>
      <sec id="sec-2-1">
        <title>Main Module</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
    </sec>
    <sec id="sec-4">
      <title>2.1 Test Data</title>
      <p>The test data consists from two files, one from Tumba! with 152 entries and one from
TEL with 108 entries. In comparison the training files had 7 entries for Tumba! file
and 9 entries for TEL file. The format for input data and for requested output data are
presented below in Table 1 and in Table 2 and are taken from TEL training data.
875336 &amp; 5431 &amp; ("central europe")
828587 &amp; 12840 &amp; ("sicilia")
902980 &amp; 482 &amp; (creator all "casanova")
196270 &amp; 5365 &amp; ("casanova")
528968 &amp; 190 &amp; ("iceland*")
470448 &amp; 8435 &amp; ("iceland")
712725 &amp; 5409 &amp; ("cavan county ireland 1870")
875336 &amp; 5431 &amp; ("&lt;place&gt;central europe&lt;/place&gt;")
828587 &amp; 12840 &amp; ("&lt;place&gt;sicilia&lt;/place&gt;")
902980 &amp; 482 &amp; (creator all "casanova")
196270 &amp; 5365 &amp; ("casanova")
528968 &amp; 190 &amp; ("&lt;place&gt;iceland&lt;/place&gt;*")
470448 &amp; 8435 &amp; ("&lt;place&gt;iceland&lt;/place&gt;")
712725 &amp; 5409 &amp; ("&lt;place&gt;cavan county ireland&lt;/place&gt; 1870")
How we can see in above tables the aim is to add &lt;place&gt; &lt;/place&gt; tags to user
queries for geographical elements.</p>
    </sec>
    <sec id="sec-5">
      <title>2.2 Resources</title>
      <p>In order to build our resource with geographical entities we start from GATE
resources and additional, we search on the web in order to add new similar entities. In
separated runs, we load our resources or resources provide by organizers: page titles
from Portuguese and English Wikipedia.</p>
    </sec>
    <sec id="sec-6">
      <title>GATE</title>
      <p>
        From GATE
        <xref ref-type="bibr" rid="ref1">(Cunningham et al., 2001)</xref>
        we use the following sets of named
entities: cities, countries, small regions, regions, mountains and provinces. In the end
the total number of entities used from GATE resources was 146.581 and the size on
disk was around 4.47 Mb. In Table 3 are examples of GATE resources.
      </p>
      <sec id="sec-6-1">
        <title>Aceh, Acores, Acquaviva, … Africa, Algarve, Antarctica, Ashmore and Cartier Islands, Asia, … 213</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Wikipedia</title>
      <p>Organizers offered to the participants’ two files with titles from Portuguese and
English Wikipedia. In comparison with GATE, the number of entries from file with
Portuguese titles from Wikipedia was 934.395 and the size on disk was 35.3 Mb, and
the number of entries from file with English titles from Wikipedia was 6.996.744 and
the size on disk was 282 Mb. In Table 4 are few lines from these files.</p>
      <sec id="sec-7-1">
        <title>Portuguese</title>
        <p>Examples
&lt;title&gt;AmericanSamoa&lt;/title&gt;
&lt;title&gt;AppliedEthics&lt;/title&gt;
&lt;title&gt;AccessibleComputing&lt;/title&gt;
&lt;title&gt;Anarchism&lt;/title&gt;…
&lt;title&gt;Astronomia&lt;/title&gt;
&lt;title&gt;Astronomia e astrofísica&lt;/title&gt;
&lt;title&gt;América Latina&lt;/title&gt;
&lt;title&gt;Albino Forjaz de Sampaio&lt;/title&gt;…
2.3</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Main Module</title>
      <p>The main module load successively the GATE or Wikipedia resources in cache using
a hash map in which the key is named entity itself, and the value is the number of
words from initial named entity. After that using this cache we will try to identify in
TEL and Tumba test data the geographical entities.</p>
    </sec>
    <sec id="sec-9">
      <title>Resources Loading</title>
      <p>From beginning when we load our geographical resources in cache, we transform
all characters from these entities in lower case. Additional, we split every name entity
in components words and load also these separated words in our cache, but we specify
the number of words from initial entity (in this way we will know if this key from our
hash map come from a simple name entity or from a composed name entity). We will
see how we will use this value when we try to identify geographical entities in user
query. Even if we lose much time at resource loading, after that the operations of
identification are very fast regardless of the number of user’s queries that we want to
process.</p>
    </sec>
    <sec id="sec-10">
      <title>Test Data Pre-Processing</title>
      <p>In order to identify in test data geographical entities from our cache, we perform
pre-processing of test data. The most important steps are:
1) In first step we parse the current line from test data and extract only the
relevant text (actually this is the initial user query).
1.</p>
      <p>2)
3)</p>
      <p>Second step has the aim to ignore special characters like +, (, ), *, “, ” or
white spaces (tab, space, return) in initial user query.</p>
      <p>Third step transform all characters extracted in previous step in lower case
characters. The result is a new form of user query, called from now new
query.</p>
    </sec>
    <sec id="sec-11">
      <title>Geographical Entities Identification</title>
      <p>The most important operation of main module is identification of the geographical
entities in this new form of user query. From now the main question is: How we
identify the geographical entities in this new query?</p>
      <p>Initial we try to see if we have in our cache the new query itself. If YES, then we
finished the process of identification of geographical entities and we skip to the next
line in test data file. This is the case of the following line from Tel test data:
4752 &amp; 11759 &amp; ("portugal)"
for which we have in our cache the all user query which is “Portugal” from GATE
countries file.</p>
      <p>If NO, then we try to split new user query in separated words if this is possible.
When we have only one word in the new query we automatically skip to the next line
in test data file. When we have more than one word, we apply the following steps:
At step 1 every individual word is searched in hash map. If current word comes
from a simple named entity we simply add “place” tags to it. This is the case of
below line:
4892 &amp; 5670 &amp;
mountain ranges)"
("climbing
on
the</p>
      <p>Himalaya
and
other
for which we have in our cache the separated word “Himalaya” from GATE
mountains file.</p>
      <p>At step 2 for every word searched in hash that comes from a composed named
entity we look successively in it left and it right in order to combine more words
with the same value in hash map.
2.1. If we have like neighbors these types of hash keys then we create a common
tag. This is the case of the following line:
13128 &amp; 11516 &amp; ("peter woods)"
for which we have in our cache separated keys “Peter” and “Woods” from
GATE cities file (first one from “Peter Tavy” and second one from “Harper
Woods”). Because both have the same value in hash (2 which represents the
number of words from initial entity) we create a common tag for both words.
2.2. If we haven’t like neighbors these types of hash keys, then we eliminate the
all these tags.</p>
      <p>During all previous steps the stop words are ignored.
We submitted two pairs of runs: one in which the main module loaded GATE like
external resource, and one in which Portuguese or English Wikipedia are loaded like
external resources.
This paper presents the UAIC system which took part in the LogCLEF 2009
competition in LAGI task. The system uses like external resources GATE files and
two files offered by organizers with titles from Portuguese and English Wikipedia.</p>
      <p>Initial we load external resources in cache using a hash map. After that, at
preprocessing part we obtain a new query from the initial user query. This new query is
used by main module in order to identify geographical entities. In this process we
searches in cache the initial query, and after that words components, with aim to have
in the end the most comprehensive succession of geographical entities.</p>
      <p>The results show how the results obtained with GATE or with Wikipedia like
external resources are comparable like quality.</p>
      <p>The author would like to thank to the students Victor Chircu and Alexandru Cristea
and their colleagues from group 4A, second year, for their help and support at
different stages of system development.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tablan</surname>
          </string-name>
          , V.:
          <article-title>GATE: an architecture for development of robust HLT applications</article-title>
          .
          <source>In ACL '02: Proceedings of the 40th Annual Meeting on Association for Computational Linguistics</source>
          ,
          <fpage>168</fpage>
          --
          <lpage>175</lpage>
          , Association for Computational Linguistics, Morristown, NJ, USA (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>