<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TextMESS at GeoCLEF 2008: Result Merging with Fuzzy Borda Ranking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jose M. Perea Ortega</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>L. Alfonso Uren~a Lopez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natural Language Engineering Lab</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dpto. de Informatica</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dpto. de Sistemas Informaticos y Computacion</institution>
          ,
          <addr-line>DSIC</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidad Politecnica de Valencia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universidad de Jaen</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describe the joint participation by the Universidad Politecnica de Valencia and the Universidad of Jaen to the GeoCLEF task. This activity has been carried out within the framework of the Spanish TextMESS project (Intelligent, Interactive and Multilingual Text Mining based on Human Language Technologies). The method employed for the participation is a result merging algorithm based on the fuzzy Borda voting scheme. This method takes as input the two document lists returned by the two systems developed by the participating groups and creates a document list where the documents are ranked according to the fuzzy Borda voting scheme. The results obtained are better than the individual systems, and also ones of the best ones of the task (second as group). However, the best result was obtained with a run which combined the baseline systems. The analysis of the results showed that the best runs were those in which only title and description were used, and unfortunately we chose to submit only a run of this type, with the base systems. The results con rm the e ectiveness of the fuzzy Borda scheme for the combination of di erent systems.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>I</kwd>
        <kwd>2 [Arti cial Intelligence]</kwd>
        <kwd>I</kwd>
        <kwd>2</kwd>
        <kwd>3 Uncertainty</kwd>
        <kwd>\fuzzy</kwd>
        <kwd>" and probabilistic reasoning</kwd>
        <kwd>I</kwd>
        <kwd>2</kwd>
        <kwd>7 Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In this paper we describe the joint participation of the groups of the Universidad Politecnica
de Valencia and Universidad de Jaen to GeoCLEF 2008. This participation has been carried
out within the framework of the Spanish project TextMESS (Spanish acronym for Intelligent,
Interactive and Multilingual Text Mining based on Human Language Technologies).</p>
      <p>We previously investigated various possibilities for the integration of our systems, focusing on
two possible choices:</p>
      <p>Identify some features in the topics that allow to determine which system is going to obtain
the best result over a determined topic (system selection);</p>
      <p>
        Combine the output of the di erent systems in a unique output (output merging ).
We carried out some preliminary experiments with the GeoCLEF topics from 2005 to 2007 and the
systems presented by the two groups in GeoCLEF 2007, in order to check whether the rst option
was feasible or not. These results proved that it was possible to use bag-of-words features to select
the best system for a given topic. However, the experiments carried out with the new systems
(those developed for GeoCLEF 2008) did not provide us with the same conclusion. Therefore, we
chose to participate with an output merging algorithm based on the fuzzy Borda voting scheme
[
        <xref ref-type="bibr" rid="ref6 ref9">9, 6</xref>
        ]. This method was previously used in the Word Sense Disambiguation task at Semeval1 with
good results [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Preliminary experiments with the data from 2005 to 2007 showed that it was
possible to achieve an improvement of 2% in Mean Average Precision (MAP) over the best
system.
      </p>
      <p>In Sections 2 and 3 we describe brie y the systems of each group (a more complete description
can be found in the corresponding report of the CLEF Working Notes). In Section 4 we describe
the fuzzy Borda ranking method, and nally we present the results and a brief discussion.
2</p>
    </sec>
    <sec id="sec-2">
      <title>SINAI-GIR System Description</title>
      <p>
        The SINAI-GIR system is made up of ve main subsystems: Translator, Collection Preprocessing
subsystem, Query Analyzer, Information Retrieval subsystem and Validator. Each translated query
is preprocessed and analyzed by the Query Analyzer, identifying their geo-entities and spatial
relations and making use of Geonames gazetteer2. This module also applies query reformulation based
on the query parsing subtask [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], generating several independent queries which will be indexed
and searched by means of the IR subsystem. On the other hand, the collection is preprocessed
by the Collection Preprocessing module and nally the documents recovered by the IR subsystem
are ltered and re-ranked by means of the Validator subsystem. Figure 1 shows the SINAI-GIR
system architecture.
      </p>
      <p>The main features of each subsystem are:</p>
      <p>
        Translator. We have used SINTRAM (SINai TRAnslation Module), our Machine
Translation system which works with di erent online machine translators and implements several
heuristics to combine di erent translations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Collection Preprocessing Subsystem. During the collection preprocessing, two indexes
are generated (locations and keywords indexes). We apply the Porter stemmer [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the
Brill POS tagger [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and a speci c Named Entity Recognizer (NER) as LingPipe3. We also
discard the English stop-words.
      </p>
      <p>Query Analyzer. It is responsible for preprocessing of English queries as well as the
generation of di erent query reformulations.</p>
      <sec id="sec-2-1">
        <title>1http://nlp.cs.swarthmore.edu/semeval 2http://www.geonames.org 3http://alias-i.com/lingpipe</title>
        <p>Validator . The aim of this subsystem is to lter the lists of documents recovered by the
IR subsystem, establishing what of them are valid, depending on the locations and the
georelations detected in the query. Another important function is to establish the nal ranking
of documents, based on manual rules and prede ned weights.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The UPV GeoWorSE System</title>
      <p>
        The system is built around the Lucene5 open source search engine, version 2:1. The Stanford
NER system based on Conditional Random Fields [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used for Named Entity Recognition and
classi cation. The access to WordNet is provided by the MIT Java WordNet Interface 6. The
toponym disambiguator is based on the method presented in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Indexing</title>
        <p>During the indexing phase, the documents are examined in order to nd location names (toponym)
by means of the Stanford NER system. When a toponym is found, the disambiguator determines
the correct reference for the toponym. Then, a modi ed lucene indexer adds to the geo index the
toponym coordinates (retrieved from GeoWordNet); nally, it stores in the wn index the toponym
together with its holonyms and synonyms. All document terms are stored in the text index. The
indices are then used in the search phase, although the geo index is not used for search: it is used
only to retrieve the coordinates of the toponyms in the document.</p>
        <sec id="sec-3-1-1">
          <title>4http://www.lemurproject.org 5http://lucene.apache.org/ 6http://www.mit.edu/ markaf/projects/wordnet/</title>
          <p>3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Searching</title>
        <p>The architecture of the search module is shown in Figure 2.</p>
        <p>
          The topic text is searched by Lucene in the text index. The toponyms extracted by the Stanford
NER are searched for in the wn index with a weight 0:25 with respect to the content terms. The
result of the search is a list of documents ranked using the Lucene's weighting scheme. At the
same time, the toponyms are passed to a module named GeoAnalyzer that creates a geographical
constraint that is used to re-rank the document list. The GeoAnalyzer may return two types of
geographical constraints:
a distance constraint, corresponding to a point in the map: the documents that contain
locations closer to this point will be ranked higher;
an area constraint, correspoinding to a polygon in the map: the documents that contain
locations included in the polygon will be ranked higher. The polygon is obtained by
calculating the convex hull of the points associated to the toponyms using the Graham algorithm
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>WordNet is used by the GeoAnalyzer module in order to extract the meronyms of the toponyms
in the topic. These meronyms allow to improve the precision of the area constraint.</p>
        <p>The objective of the GeoFilter module is to re-rank the documents retrieved by Lucene,
according to geographical information. If the constraint extracted from the topic is a distance constraint,
the weights of the documents are modi ed according to the following formula:
w(doc) = wLucene(doc) (1 + exp(
min d(q; p)))
p2P
(1)</p>
        <p>Where wLucene is the weight returned by Lucene for the document doc, P is the set of points
in the document, and q is the point extracted from the topic.</p>
        <p>If the constraint extracted from the topic is an area constraint, the weights of the documents
are modi ed according to formula 2:
w(doc) = wLucene(doc) 1 + jPqj (2)
jP j
where Pq is the set of points in the document that are contained in the area extracted from
the topic.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Fuzzy Borda Merging</title>
      <sec id="sec-4-1">
        <title>Fuzzy Borda count</title>
        <p>
          In the classical (discrete) Borda count each expert gives a mark to each alternative, according
to the number of alternatives worse than it. The fuzzy variant [
          <xref ref-type="bibr" rid="ref6 ref9">9, 6</xref>
          ] allows the experts to show
numerically how much some alternatives are preferred to the others, evaluating their preference
intensities from 0 to 1.
        </p>
        <p>Let R1; R2; : : : ; Rm be the fuzzy preference relations of m experts over n alternatives x1; x2; : : : ; xn.
Each expert k expresses its preferences by means of a matrix of preference intensities:
where each rikj = Rk (xi; xj ), with Rk : X X ! [0; 1] is the membership function of Rk. The
k
number rij 2 [0; 1] is considered as the degree of con dence with which the expert k prefers xi to
xj . The nal value assigned by the expert k to each alternative xi is the sum by row of the entries
greater than 0:5 in the preference matrix, or, formally:
0 r11
k
k
B r21
B@ : : :
k
rn1
k
r12
k
r22
: : :
k
rn2
: : : r1kn 1
: : : r2kn C
: : : : : : CA</p>
        <p>k
: : : rnn
rk(xi) =
n</p>
        <p>
          X
j=1;rikj&gt;0:5
k
rij
r(xi) =
m
X rk(xi)
k=1
(3)
(4)
(5)
The threshold 0:5 ensure the relation Rkto be an ordinary preference relation [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>The fuzzy Borda count for an alternative xi is obtained as the sum of the values assigned by
each expert:
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Application of Fuzzy Borda count to Result Merging</title>
        <p>In our approach each system is an expert: therefore, there are two preference matrices. The size
of these matrices is variable: the reason is that the document list is not the same for the two
systems. Therefore, the size of a preference matrix is Nt Nt, where Nt is the number of unique
documents retrieved by the two systems (i.e. the number of documents that appear at least in
one of the lists returned by the systems) for topic t.</p>
        <p>The systems ranks the document with weights that are not in the same range. Therefore,
the output weights w1; w2; : : : ; wn of each expert k are transformed to fuzzy con dence values by
means of the following transformation:
rikj =</p>
        <p>wi
wi + wj</p>
        <p>This transformation ensure that the preference values are in the range [0; 1]. In order to adapt
the fuzzy Borda count to the merging of the results of IR systems, we had to determine the
preference values in all the cases where one of the systems does not retrieve a document that has
been retrieved by the other one. We decided to set the preference values of these documents to
0:5. This corresponds to the idea that the expert is presented an option on which it cannot express
a preference.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>We submitted a total of 9 runs. In Table 1 we show the detail of each run in terms of the two
systems combined and the topic elds used.</p>
      <p>In Table 3 we show the Mean Average Precision (MAP) obtained for each run, together with
the MAP obtained by its composing runs.</p>
      <p>The obtained results show that the use of the fuzzy Borda merging method always allows
to improve the results of the best system. The improvement is greater if the two systems have
a similar performance (see TMESS06) and the UPV system does not use map ltering. This
behaviour is not observed when the UPV system uses map ltering. We suppose that a key
feature for obtaining greater improvements by means of fuzzy Borda is that the systems share as
few as characteristics as possible.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Further Work</title>
      <p>We combined two di erent systems by means of the fuzzy Borda voting scheme. The implemented
method allowed to improve the results of the combined systems, although the improvement was
limited. We suppose that the best results with the fuzzy Borda merging can be obtained if the
two systems share the same level of accuracy. Further work will be aimed to verify this hypothesis
and to the integration of more than two systems.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We would like to thank the TIMOM (TIN2006-15265-C06-03) and TIN2006-15265-C06-04 research
projects for partially supporting this work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Eric</given-names>
            <surname>Brill</surname>
          </string-name>
          .
          <article-title>A simple rule-based part-of-speech tagger</article-title>
          .
          <source>In Proceedings of the third Conference on Applied Natural Language Processing (ANLP'92)</source>
          , pages
          <fpage>152</fpage>
          {
          <fpage>155</fpage>
          ,
          <string-name>
            <surname>Trento</surname>
          </string-name>
          , Italy,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Davide</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>Upv-wsd : Combining di erent wsd methods by means of fuzzy borda voting</article-title>
          .
          <source>In Fourth International Workshop on Semantic Evaluations (SemEval2007)</source>
          , pages
          <fpage>434</fpage>
          {
          <fpage>437</fpage>
          .
          <string-name>
            <surname>ACL</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Davide</given-names>
            <surname>Buscaldi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <article-title>A conceptual density-based approach for the disambiguation of toponyms</article-title>
          .
          <source>International Journal of Geographical Information Systems</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <volume>301</volume>
          {
          <fpage>313</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jenny</given-names>
            <surname>Rose</surname>
          </string-name>
          <string-name>
            <surname>Finkel</surname>
          </string-name>
          , Trond Grenager, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Incorporating non-local information into information extraction systems by gibbs sampling</article-title>
          .
          <source>In Proceedings of the 43nd Annual Meeting of the Association for Computational Linguistics (ACL</source>
          <year>2005</year>
          ), pages
          <fpage>363</fpage>
          {370, U. of Michigan - Ann Arbor,
          <year>2005</year>
          . ACL.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Miguel</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Garc</surname>
            a-Cumbreras,
            <given-names>L. Alfonso</given-names>
          </string-name>
          <article-title>Uren~a Lopez, Fernando Mart nez Santiago</article-title>
          , and Jose M.
          <article-title>Perea Ortega. BRUJA System. The University of Jaen at the Spanish task of QA@CLEF 2006</article-title>
          .
          <source>In Lecture Notes in Computer Science</source>
          , volume
          <volume>4730</volume>
          <source>of LNCS Series</source>
          , pages
          <volume>328</volume>
          {
          <fpage>338</fpage>
          . Springer-Verlag,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jose</given-names>
            <surname>Luis Garc</surname>
          </string-name>
          <article-title>a Lapresta and Miguel Mart nez Panero. Borda Count Versus Approval Voting: A Fuzzy Approach</article-title>
          . Public Choice,
          <volume>112</volume>
          (
          <issue>1-2</issue>
          ):
          <volume>167</volume>
          {
          <fpage>184</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ronald</surname>
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Graham</surname>
          </string-name>
          .
          <article-title>An e cient algorith for determining the convex hull of a nite planar set</article-title>
          .
          <source>Information Processing Letters</source>
          ,
          <volume>1</volume>
          (
          <issue>4</issue>
          ):
          <volume>132</volume>
          {
          <fpage>133</fpage>
          ,
          <year>1972</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Zhisheng</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Chong</given-names>
            <surname>Wanga</surname>
          </string-name>
          , Xing Xie, and
          <string-name>
            <surname>Wei-Ying Ma</surname>
          </string-name>
          .
          <source>Query Parsing Task for GeoCLEF 2007 Report. In Proceedings of the Cross Language Evaluation Forum (CLEF</source>
          <year>2007</year>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Hannu</given-names>
            <surname>Nurmi</surname>
          </string-name>
          .
          <article-title>Resolving Group Choice Paradoxes Using Probabilistic and Fuzzy Concepts</article-title>
          .
          <source>Group Decision and Negotiation</source>
          ,
          <volume>10</volume>
          (
          <issue>2</issue>
          ):
          <volume>177</volume>
          {
          <fpage>199</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.F.</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <article-title>An algorithm for su x stripping</article-title>
          .
          <source>In Program 14</source>
          , pages
          <fpage>130</fpage>
          {
          <fpage>137</fpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>