<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MIRACLE's Approach to Multilingual Web Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ángel Martínez-González</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Luis Martínez-Fernández</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>César de Pablo-Sánchez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Villena-Román</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis Jiménez-Cuadrado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paloma Martínez</string-name>
          <email>paloma.martinez@uc3m.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Carlos González-Cristóbal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Politécnica de Madrid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Carlos III de Madrid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DAEDALUS - Data</string-name>
          <email>amartinez@daedalus.es</email>
          <email>jmartinez@daedalus.es</email>
          <email>jvillena@daedalus.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Decisions</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Language</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Martínez</institution>
          ,
          <addr-line>José L.; Villena, Julio; Fombella, Jorge; G. Serrano, Ana; Martínez, Paloma; Goñi, José M.; and González</addr-line>
          ,
          <institution>José C. MIRACLE Approaches to Multilingual Information Retrieval: A Baseline for Future Research. Comparative Evaluation of Multilingual Information Access Systems (Peters</institution>
          ,
          <addr-line>C; Gonzalo, J.; Brascher, M.; and Kluck, M., Eds.). Lecture Notes in Computer Science, vol. 3237, pp. 210-219. Springer, 2004</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2003</year>
      </pub-date>
      <fpage>21</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>For MIRACLE participation on WebClef 2005, a set of independent indexes was constructed for each top level domain of the EuroGOV collection. Each of these indexes contains information extracted from the document, like URL, title, keywords, detected named entities or HTML headers. These indexes are queried to obtain partial document rankings, which are combined with various relative weights to test the value of each index. The trie based indexing and retrieval engine developed by the MIRACLE team is now fully functional and has been adapted to the WebClef environment and employed in this campaign. Other tools, such as the Named Entities Recognizer based on a finite automaton, have also been developed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>or HTML headers. The MIRACLE group has taken part in the two main tasks (Mixed Monolingual and</p>
      <sec id="sec-1-1">
        <title>Multilingual).</title>
        <p>2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Experiment design</title>
      <p>For MIRACLE participation on WebClef 2005 we decided not to follow a full text approach. Instead, a set of
independent indexes was constructed for each top level domain of the collection. Partial results are obtained by
applying the probabilistic ranking formula BM25 [12] to these indexes. Finally, these partial results are
combined to get the final result. In different experiments, different weights are given to each set of partial results,
so as to evaluate the relative importance of the different information sources that have been indexed.</p>
      <sec id="sec-2-1">
        <title>The generated indexes were the following:</title>
      </sec>
      <sec id="sec-2-2">
        <title>H1 index, containing document titles and H1 HTML headers (if you are not familiar with the HTML standard, see [11]).</title>
      </sec>
      <sec id="sec-2-3">
        <title>H2 index, containing headers H2 to H6.</title>
      </sec>
      <sec id="sec-2-4">
        <title>PN index, containing named entities (proper nouns) found by a detection module.</title>
      </sec>
      <sec id="sec-2-5">
        <title>Ky index, document keywords given in a META HTML element.</title>
      </sec>
      <sec id="sec-2-6">
        <title>Url, containing parsed parts of the document url, removing the querystring and taking characters such as</title>
        <p>‘.’ ,‘/’ or ’–‘ as delimiters.</p>
        <p>Consequently, the total number of indexes was 85 (5 indexes/domain * 17 domains). Another index called
LINKS was initially planned but finally not included due to lack of time. This index contained the words in the
anchor (&lt;A&gt;) elements of documents that point to the indexed document. Note that, unlike the other indexes,
LINKS needs two passes over the collection, so that links pointing to a document not in the collection are
discarded. Although this index was not included, the tool for building it is available for future participations.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Developed tools</title>
      <p>
        MIRACLE toolbox [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ] consists of a set of independent modules that perform extraction (XML parsing),
preprocessing (word segmentation, filtering, stop word removing, stemming), indexing and retrieval of
documents. In last year participation, a trie based indexing and retrieval engine was under development, but not
yet finished, so a Xapian based front-end was used. This trie base engine is now fully functional and has been
used for all WebCLEF experiments (an also in MIRACLE participation in other tracks). Some of the functional
characteristics of this engine are:
•
•
•
•
•
•
•
•
•
      </p>
      <sec id="sec-3-1">
        <title>Several variants of probabilistic and vectorial ranking formulas can be selected, as an option of the retrieval program, with no need of reindexing. There is only one index format, which contains all the necessary information for each ranking algorithm.</title>
      </sec>
      <sec id="sec-3-2">
        <title>Indexing time, which is the most critical factor limiting the number of different experiments that can be performed in the available time, is optimized rather than retrieval times or index size. An important improvement has been achieved in this aspect over the previously used Xapian based tool.</title>
      </sec>
      <sec id="sec-3-3">
        <title>The index can be incrementally built. Deletion of terms or documents is in principle also supported, but inefficient.</title>
      </sec>
      <sec id="sec-3-4">
        <title>Relevance feedback is also supported, although not used in these experiments.</title>
      </sec>
      <sec id="sec-3-5">
        <title>The other tools created by MIRACLE for WebCLEF are explained bellow, in the order they are applied to the collection:</title>
        <p>•
•</p>
      </sec>
      <sec id="sec-3-6">
        <title>Document extraction: the collection is given in a few huge files with a format close to XML. The</title>
        <p>documents must be extracted.</p>
        <p>HTML parser: based on the El-Kabong HTML processing library [2]. The content or attributes of
special tags such as headers, anchors or META tags are extracted. The body is then extracted to
plain text format.
•
•
•
•</p>
      </sec>
      <sec id="sec-3-7">
        <title>Named Entities Recognition: named entities are filtered from plain text using a multilingual</title>
        <p>recognizer in current development. Recognition is based on the evaluation of predicates in a Finite</p>
      </sec>
      <sec id="sec-3-8">
        <title>State Automaton. We have explicitly considered Spanish, Portuguese, Italian, French, English,</title>
      </sec>
      <sec id="sec-3-9">
        <title>Swedish and Dutch. Simple and multiword proper nouns are detected by means of cues such as capitalization, words that introduce named entities (“Sr.”, “president”, “river”), connectors (“van”, “de”) and punctuation signs (“Paris-Dakar”, “Madigan’s”, etc.).</title>
      </sec>
      <sec id="sec-3-10">
        <title>After WebCLEF we have evaluated the tool with data from the CONLL 2002 shared task and achieved the following results, only for recognition:</title>
      </sec>
      <sec id="sec-3-11">
        <title>Combination of partial results: relevance rankings from different indexes are mixed by means of an ad-hoc script that calculates the average relevance allowing to easily assign different weights to different indexes.</title>
      </sec>
      <sec id="sec-3-12">
        <title>Query language detection: in the case of the baseline mixed monolingual run, no metadata such as</title>
        <p>the target language of the query was allowed to be used, so this module tries to guess the target
language from the words of the query title. For the multilingual task we have used a list of stop
words, while for the mixed monolingual task a list of locations and names of the inhabitants of a
country or region. A simple vote algorithm has been used.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Description of the submitted runs</title>
      <p>The MIRACLE team has taken part in the two main tasks (Mixed Monolingual and Multilingual), submitting
five runs for each one of them. A baseline run, using no metadata is mandatory. The other four runs (which will
be referred as extended in this paper) use supplied metadata (the target domain). In the baseline runs, the
language identification tool is employed to guess the target language from the words in the query title. For each
query, only the indexes of the top level domain corresponding to the target language and the international INT
domain are queried. For example, if the target language of a query is known of has been guessed to be Spanish,
only ES and INT domains are considered, even though there are documents in Spanish in the other domains.</p>
      <sec id="sec-4-1">
        <title>In the five Monolingual runs submitted, partial results were combined in the following ways:</title>
        <p>•
•
•
•
•</p>
      </sec>
      <sec id="sec-4-2">
        <title>Monobase: this is the baseline run. Relevance of documents is averaged over the five partial results, giving al of them the same weight.</title>
      </sec>
      <sec id="sec-4-3">
        <title>MonoExt: extended run, combining the results in the same way as in MonoBase.</title>
      </sec>
      <sec id="sec-4-4">
        <title>MonoExtH1PN: extended run; only H1 and PN indexes are considered, giving both of them the same weight.</title>
      </sec>
      <sec id="sec-4-5">
        <title>MonoExtUrlKy: extended run; only Url and Ky indexes are considered, giving both of them the same weight.</title>
      </sec>
      <sec id="sec-4-6">
        <title>MonoExtAH1PN: extended run. All indexes are considered for retrieval, but the H1, PN and Ky</title>
        <p>indexes are considered more relevant than the rest, so a weight factor with value 2 is applied for these
partial lists.</p>
        <p>In the Multilingual runs Multibase, MultiExt, MultiExtH1PN, MultieExtUrlKy and MultiExtAH1PN, partial
results are mixed in the same way as in the corresponding monolingual runs. All this information is summarized
in the table below.
The following graphics are based on the evaluation results provided by WebCLEF organizers. The different
parameters included in these results will be explained bellow as they appear in a figure.</p>
        <p>The average success at n is defined as the portion of topics where the known-entity was found at a rank less than
or equal to n. The following figures show this average success rate as a function of n for the monolingual and
multilingual tasks.</p>
        <p>Average success in monolingual runs
0,45</p>
        <p>0,4
0,35</p>
        <p>0,3
0,25</p>
        <p>0,2
0,15</p>
        <p>0,1
0,05
0
5
0,4
0,3
0,2
0,1
Average success in multilingual runs</p>
      </sec>
      <sec id="sec-4-7">
        <title>MultiBase</title>
      </sec>
      <sec id="sec-4-8">
        <title>MultiExt</title>
      </sec>
      <sec id="sec-4-9">
        <title>MultiExtAH1PN</title>
      </sec>
      <sec id="sec-4-10">
        <title>MultiExtH1PN</title>
      </sec>
      <sec id="sec-4-11">
        <title>MultiExtUrlKy</title>
        <p>1
5
10
20
50</p>
      </sec>
      <sec id="sec-4-12">
        <title>The expected conclusion is confirmed: titles and named entities are the most valuable sources of information to</title>
        <p>find known-items. In Multilingual runs results are worse and the effect of different combinations of results are
not so significant.</p>
        <p>The Mean Reciprocal Rank (MRR) is 1 divided by the rank given to the known-entity or 0 if the relevant
document has not been retrieved. The parameter called DFA is defined as the difference between the MRR score
for a given topic and the average MRR score over the submitted runs of all participants. In the figures bellow,
DFA values as a function of the topic are given for the best extended monolingual and multilingual runs:</p>
      </sec>
      <sec id="sec-4-13">
        <title>MonoExtH1PN and MultiExtAH1PN.</title>
        <p>The DFA value averaged over all topics is -0,04 for MonoExtH1PN and +0,02 for MultiExtAH1PN: both of
them are very close to 0, the mean value over runs submitted by all participants. Our results in the multilingual
task are, although worse in absolute terms than the monolingual results, better if considered relative to the other
participants, even though our approach was quite simple, with no query or document translation. Elements
without translation such as named entities are less noisy and especially valuable for known-item search.
Although not shown in the figures above, our results were rather variable with the target language of the topic.
The results in languages such as Greek or Russian were much poorer than other languages, even though the
techniques used are language independent (with the partial exception of named entities recognition). This
suggests we have had some sort of problem with character sets and encodings, which should be corrected for
future participations.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and future work</title>
      <p>Obviously, in this first year of WebCLEF track, there were no previous results available and the selection of
experiments was somehow blind. Nevertheless the foundations for future campaigns have been settled and
several valuable conclusions have been drawn. We have at our disposal a set of software tools that we plan to use
and further improve in future campaign in order to pursue more ambitious aims.</p>
      <p>We believe that a full text index, combined appropriately with the more specific indexes would probably
improve the results. In the next campaign, we are also planning to introduce some sort of query translation
mechanism. Another improvement would be to consider the hyperlink structure of the collection; a voting
algorithm could be used to estimate the relative importance of web pages and this way detect home pages.</p>
      <sec id="sec-5-1">
        <title>Finally, we are considering experimenting with automatic web classification using neural networks.</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This work has been partially supported by the Spanish R+D National Plan, by means of the project RIMMEL
(Multilingual and Multimedia Information Retrieval, and its Evaluation), TIN2004-07588-C03-01.
Special mention to our colleagues of the MIRACLE team should be done (in alphabetical order): Ana María
García-Serrano, José Carlos González-Cristóbal, Ana González Ledesma, José Miguel Goñi-Menoyo, José Mª</p>
      <sec id="sec-6-1">
        <title>Guirao, Sara Lana-Serrano and Antonio Moreno Sandoval.</title>
        <p>[2] El-Kabong HTML. A speedy HTML processing library. On line http://www.ekhtml.sourceforge.net/
[visited 28/07/2005].</p>
      </sec>
      <sec id="sec-6-2">
        <title>Martínez, J.L.; Villena-Román, J.; Fombella, J.; García-Serrano, A.; Ruiz, A.; Martínez, P.; Goñi, J.M.;</title>
        <p>and González, J.C. (Carol Peters, Ed.): Evaluation of MIRACLE approach results for CLEF 2003.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Working Notes for the CLEF 2003 Workshop, 21-22 August, Trondheim, Norway.</title>
        <p>[9] de Pablo, C.; Martínez-Fernández, J. L.; Martínez, P.; Villena, J.; García-Serrano, A. M.; Goñi, J. M.; and</p>
      </sec>
      <sec id="sec-6-4">
        <title>González, J. C. miraQA: Initial experiments in Question Answering. CLEF 2004 proceedings (Peters, C. et al., Eds.). Lecture Notes in Computer Science, vol. 3491. Springer, 2005 (to appear). [10] Porter, Martin. Snowball stemmers and resources page. On line http://www.snowball.tartarus.org. [Visited 13/07/2005]</title>
        <p>[11] Ragget, D.; Le Hors A. and Jacobs I. (Ed.). HTML 4.01 Specifications. W3C Recommendation 24</p>
      </sec>
      <sec id="sec-6-5">
        <title>December 1999. On line http://www.w3.org/TR/html4/ [visited 15/07/2005].</title>
        <p>[12] Robertson, S.; Walker, S.; Hancock-Beaulieu, M.M. and Gatford, M. Okapi at trec 3. Text Retrieval</p>
      </sec>
      <sec id="sec-6-6">
        <title>Conference, 2003.</title>
        <p>[13] SYSTRAN 5.0 translation resources. On line http://www.systransoft.com. [Visited 13/07/2005]
[14] University of Neuchatel. page of resources for CLEF (Stopwords, transliteration, stemmers, …). On line
http://www.unine.ch/info/clef/. [Visited 13/07/2005]
[15] Villena, Julio; Martínez, José L.; Fombella, Jorge; G. Serrano, Ana; Ruiz, Alberto; Martínez, Paloma;</p>
      </sec>
      <sec id="sec-6-7">
        <title>Goñi, José M.; and González, José C. Image Retrieval: The MIRACLE Approach. Comparative</title>
      </sec>
      <sec id="sec-6-8">
        <title>Evaluation of Multilingual Information Access Systems (Peters, C; Gonzalo, J.; Brascher, M.; and Kluck,</title>
      </sec>
      <sec id="sec-6-9">
        <title>M., Eds.). Lecture Notes in Computer Science, vol. 3237, pp. 621-630. Springer, 2004.</title>
      </sec>
      <sec id="sec-6-10">
        <title>Xapian: an Open Source Probabilistic Information Retrieval library. On line http://www.xapian.org.</title>
        <p>[Visited 13/07/2005]</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Aoe</surname>
            , Jun-Ichi; Morimoto, Katsushi; Sato,
            <given-names>Takashi.</given-names>
          </string-name>
          <article-title>An Efficient Implementation of Trie Structures</article-title>
          .
          <source>Software Practice and Experience</source>
          <volume>22</volume>
          (
          <issue>9</issue>
          ):
          <fpage>695</fpage>
          -
          <lpage>721</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Goñi-Menoyo</surname>
          </string-name>
          , José M; González, José C.;
          <string-name>
            <surname>Martínez-Fernández</surname>
          </string-name>
          , José L.; and
          <string-name>
            <surname>Villena</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>MIRACLE's Hybrid Approach to Bilingual and Monolingual Information Retrieval</article-title>
          .
          <article-title>CLEF 2004 proceedings</article-title>
          (Peters,
          <string-name>
            <surname>C.</surname>
          </string-name>
          et al.,
          <source>Eds.). Lecture Notes in Computer Science</source>
          , vol.
          <volume>3491</volume>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>199</lpage>
          . Springer,
          <year>2005</year>
          (to appear).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Goñi-Menoyo</surname>
          </string-name>
          , José M.;
          <string-name>
            <surname>González</surname>
          </string-name>
          , José C.;
          <string-name>
            <surname>Martínez-Fernández</surname>
          </string-name>
          , José L.;
          <string-name>
            <surname>Villena-Román</surname>
          </string-name>
          , Julio; GarcíaSerrano, Ana; Martínez-Fernández, Paloma; de Pablo-Sánchez,
          <article-title>César;</article-title>
          and
          <string-name>
            <surname>Alonso-Sánchez</surname>
          </string-name>
          ,
          <article-title>Javier. MIRACLE's hybrid approach to bilingual and monolingual Information Retrieval</article-title>
          .
          <source>Working Notes for the CLEF 2004 Workshop (Carol Peters and Francesca Borri, Eds.)</source>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>150</lpage>
          . Bath, United Kingdom,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Goñi-Menoyo</surname>
          </string-name>
          , José Miguel;
          <article-title>González-Cristóbal, José Carlos</article-title>
          and
          <string-name>
            <surname>Fombella-Mourelle</surname>
            ,
            <given-names>Jorge.</given-names>
          </string-name>
          <article-title>An optimised trie index for natural language processing lexicons</article-title>
          .
          <source>MIRACLE Technical Report</source>
          . Universidad Politécnica de Madrid,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Martínez-Fernández</surname>
          </string-name>
          , José L.;
          <string-name>
            <surname>García-Serrano</surname>
            , Ana; Villena,
            <given-names>J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Méndez-Sáez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <article-title>MIRACLE approach to ImageCLEF 2004: merging textual and content-based Image Retrieval</article-title>
          .
          <article-title>CLEF 2004 proceedings</article-title>
          (Peters,
          <string-name>
            <surname>C.</surname>
          </string-name>
          et al.,
          <source>Eds.). Lecture Notes in Computer Science</source>
          , vol.
          <volume>3491</volume>
          . Springer,
          <year>2005</year>
          (to appear).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>