<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UPV/BUAP Participation in WebCLEF 2006∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Measurement, Performance, Experimentation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Systems and Computation, Polytechnic University of Valencia (UPV)</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Computer Science, B. Autonomous University of Puebla (BUAP)</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Applied Computer Science</institution>
          ,
          <addr-line>UPV</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>After our first participation in the Bilingual task of WebCLEF 2005, we have emigrated to a more challenging task. In this report we are presenting the results obtained after evaluating a set of topics in the Mixed-Monolingual task of WebCLEF 2006. Our efforts were focused on the preprocessing of the EuroGOV corpus which is itself a very challenging task, due to the high variety of errors that must be treated in order to correctly interpret the content of each document to index. Moreover, we have tested a new formula for the ranking of the documents retrieved, which is based on the Jaccard formula but includes a penalization factor. Results are low but encourage to investigate whether they are the result of a bad preprocessing process and/or the malfunction of the search engine components.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Learning to deal with the high volume of data of Internet becomes more a need than a curiosity.
Currently, we are witnesses of a big explosion of information available that must be adequately
cataloged. Moreover, this information comes from all parts of the world, from very different
cultures and different languages which makes this task even more difficult. At the moment, it
would seem that only search engines such as Google and Yahoo could have enough resources for this
challenge, but new proposals would provide advances from the scientific instead of the comercial
viewpoint. Certainly, forums dedicated to the analysis of information search and retrieval, more
particularly in a cross-language environment, are needed.</p>
      <p>
        The WebCLEF concern is about the evaluation of information retrieval systems using
crosslingual web pages. The justification of the WebCLEF track is based on the fact that many issues
for which people turn to the web are essensially multilingual. In 2005 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the first edition of this
competition was done in the framework of the Cross Language Evaluation Forum (CLEF) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. At
that time, three tasks were proposed: mixed monolingual, multilingual, and bilingual
EnglishSpanish. Currently, only one of those tasks was suggested: mixed monolingual; the reason of this
action is derived from the very bad results obtained in the first edition in the multilingual task
compared with those of the mixed-monolingual task. Due to this reason, the WebCLEF forum
decided to focus this year on making robust the results in the mixed monolingual task, and later
to make efforts for improving the translation resources that needed to be applied in the other task.
      </p>
      <p>
        In 2005, we participated in the bilingual English to Spanish task [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In 2006 Mixed-Monolingual
task, we have experimented using the EuroGOV corpus, which was compiled in 2005 before the
WebCLEF campaign. This corpus consists in a crawl of governmental sites in Europe from
approximately 27 differents Internet domains. A better description of this corpus can be found in
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Therefore, we will not describe the corpus, but the way we processed it in order to obtain the
terms to index. The next section explains all the preprocessing steps we carried out. Section 3
explains the model implemented. In Section 4 we present our results and finally a discussion is
given.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Preparing the data</title>
      <p>The preprocessing phase of the EuroGOV corpus presents a big challenge, due to the written
variants a government web page could have. We have found that a big amount of documents do
not present a strict html syntax. We have written two scripts for obtaining the terms to be indexed
from each document. The first script uses regular expressions for excluding all the information
which is enclosed by the characters &lt; and &gt;. Although this script obtains very good results, it is
very slow and therefore we decided to used it only with three domains of the EuroGOV collection,
namely Spanish (ES), French (FR), and German (DE).</p>
      <p>On the other hand, we wrote a script based in the html syntax for obtaining all the terms
considered interesting for indexing, i.e., those different than script codes (javascript, vbscript,
style cascade sheet, etc), html codes, etc. This script speeded up our indexing process but it
did not took into account that some web pages are incorrectly written and, therefore, we missed
important information from those documents.</p>
      <p>Another preprocessing problem consists in the charset codification, which leads to a even more
difficult analysis. Although the EuroGOV corpus is given in UTF-8, the documents that made
up this corpus does not neccesarily keep this charset. We have seen that for some domains, the
charset codification is given in the html metadata tag, but also we found that this codification
could be wrong, perhaps because it was filled without the supervision of the creator of that page,
who may be does not know anything, and evenmore does not matter about charsets codifications.
We consider it as the most difficult problem in the preprocessing process.</p>
      <p>
        As usual in the information retrieval systems, we eliminated stopwords for each language
(except Greek). A good repository of resources for this step is suministered by Jacques Savoy from
the Institut interfacultaire d’informatique (see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). A variation on the elimination of diacritics
was done; we discuss in detail this approach in Section 4. The same process was applied to the
queries. The next section explains the model used in our runs.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Description of our model</title>
      <p>
        Nowadays, different information retrieval models are reported in literature [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Perhaps the most
popular model is the vector space model which uses the well-known tf − idf formula, however, in
practice this model is not viable. We have used a variation of the boolean model with ranking based
in the Jaccard similarity formula. We named this variation Jaccard with penalization, because it
takes into account the number of terms that a query Qi really matches when it is compared with
a document Dj of the collection. The formula used is presented as follows:
      </p>
      <p>Score(Qi, Dj) = |Dj| ∩ |Qi| − 1 − |Dj| ∩ |Qi|</p>
      <p>|Dj| ∪ |Qi| |Qi|</p>
      <p>As can be seen, the first component of this formula is the typical Jaccard approximation. The
evaluation of this formula is quite fast, and allows its implementation in real situations. The
results obtained by using this approach are presented in the next section.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>At the moment of writing this paper, individual results are only known by each team in the
competition, and therefore a comparative table with the other teams results is not presented. In
Table 1 we show the results obtained with our approximation. We evaluated different runs, varying
the use of diacritics and the preprocessing of the corpus.</p>
      <p>The WithoutDiac run eliminates all diacritics in both, the corpus and the topics, whereas the
WithDiac run only supresses the diacritics in the corpus. We can observe a expected reduction of
Mean Reciprocal Rank (MRR), but it is not significatively high with respect to the first run. This
is clearly derived from the amount of diacritics introduced in the topics of evaluation, which is not
very high. An analysis of the queries in real situations may be interesting in order to determine
whether the topics set is realistic. The last run (CDWithoutDiac) eliminates diacritization in
both, the topics and corpus, but also tries a charset detection for each document to be indexed.
Unfortunately, from the table we can observe that we did not success in our attempt.</p>
      <p>Run
WithoutDiac
WithDiac
CDWithoutDiac
We have proposed a new approach for the ranking formula in a information retrieval system based
on the Jaccard formula, but with a penalization factor. After evaluating this approach in the
approximately 75% of queries from the WebCLEF competition, we obtained low results. We
observed that our major problem was related to the preprocessing phase. The charset decodification
must be improved in order to correctly interpret the data.</p>
      <p>An evaluation of the use of diacritization in the task has shown that results are not
significatively different, which may be suggesting that the set of queries provided for the evaluation does
not have a high number of diacritics. More investigation would determine whether this behaviour
is realistic or must be tuned in further evaluations.</p>
      <p>Even if a comparison with other results of the competition is still pending in order to determine
how low are the results we obtained, we assume that our results are not very good from observing
those from the last competition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jacques</given-names>
            <surname>Savoy</surname>
          </string-name>
          :
          <article-title>Information about multilingual retrieval</article-title>
          , http://www.unine.ch/info/clef/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Jim´enez-</article-title>
          <string-name>
            <surname>Salazar</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso: BUAP-UPV</surname>
            <given-names>TPIRS</given-names>
          </string-name>
          :
          <article-title>A System for Document Indexing Reduction on WebCLEF</article-title>
          ,
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the Cross-Language Evaluation Forum, CLEF 2005, Lecture Notes in Computer Science</source>
          , Vol.
          <volume>4022</volume>
          , Springer-Verlang,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton: Automatic Text</surname>
          </string-name>
          <string-name>
            <surname>Processing</surname>
          </string-name>
          , Addison-Wesley,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. Y.</given-names>
            <surname>Berger</surname>
          </string-name>
          <article-title>: Report on CLEF-2005 evaluation campaign: Monolingual, bilingual, and GIRT information retrieval</article-title>
          . In C. Peters, Clough,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Magnini</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          (Ed.),
          <source>Working notes of CLEF</source>
          <year>2005</year>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sigurbj</surname>
          </string-name>
          ¨ornsson, J. Kamps, and M. de Rijke:
          <article-title>EuroGOV: Engineering a Multilingual Web Corpus</article-title>
          ,
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the CrossLanguage Evaluation Forum, CLEF 2005, Lecture Notes in Computer Science</source>
          , Vol.
          <volume>4022</volume>
          , Springer-Verlang,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sigurbj</surname>
          </string-name>
          ¨ornsson, J. Kamps, and M. de Rijke:
          <article-title>WebCLEF 2005: Cross-Lingual Web Retrieval</article-title>
          ,
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the CrossLanguage Evaluation Forum, CLEF 2005, Lecture Notes in Computer Science</source>
          , Vol.
          <volume>4022</volume>
          , Springer-Verlang,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>