<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MIRACLE at ImageCLEFphoto 2007: Evaluation of Merging Strategies for Multilingual and Multimedia Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julio Villena-Román</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Lana-Serrano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Luis Martínez-Fernández</string-name>
          <email>joseluis.martinez@uc3m.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Carlos González-Cristóbal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Carlos III de Madrid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universidad Politécnica de Madrid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DAEDALUS - Data</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Decisions</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Language</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Expansion, WordNet</institution>
          ,
          <addr-line>Word Sense Disambiguation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of MIRACLE research consortium at the ImageCLEF Photographic Retrieval task of ImageCLEF 2007. For this campaign, the main purpose of our experiments was to thoroughly study different merging strategies, i.e. methods of combination of textual and visual retrieval techniques. While we have applied all the well known techniques which we had already used in previous campaigns, for both textual and visual components of the system, our research has primarily focused on the idea of performing all possible combinations of those techniques in order to evaluate which ones may offer the best results and analyze if the combined results may improve (in terms of MAP) the individual ones. The system includes three main modules. On one hand, apart from the search engine (Xapian or Lucene), the textual retrieval module includes parsers, stemming, stopword filtering, proper noun detection and semantic expansion components. On the other hand, the visual retrieval module is based on two well-known content-based engines: GIFT and FIRE. Finally, the merging module allows to use different operators (AND, OR, LEFT, RIGHT) to combine the outputs of the two previous subsystems and to calculate the result relevance based on different metrics (max, min, avg, max-min). We finally submitted 110 multilingual textual (text-based) runs, 22 visual (content-based) runs and 21 mixed runs. Results in general show a poor performance for all groups, due to the characteristics of the image collection and the difficulty of the defined topics. The most interesting conclusion is that the defined merging strategies are successful as our best mixed experiment outperforms both the textual and visual experiments in which it is based, using the LEFT operator for the combination along with the max-min metric for computing the relevance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        This paper describes our participation at the ImageCLEF Photographic Retrieval task of ImageCLEF 2007.
Briefly, the goal of this task (fully described in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]) is: given a multilingual statement describing a user specific
information need, find as many relevant images as possible from the given multilingual document collections
containing images as well as text. The reference database for this campaign is the IAPR TC-12 Benchmark [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
created under Technical Committee 12 (TC-12) of the International Association of Pattern Recognition (IAPR
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]). This collection contains 20,000 photos (mainly colour photographs) taken from locations around the world
and comprises a varying cross-section of still natural images, annotated with semi-structured captions in English
and German.
      </p>
      <p>Participants are provided with a list of 60 topics which include a short textual title representing the research goal
in 15 different languages (Danish, Dutch, English, Finnish, French, German, Italian, Japanese, Norwegian,
Polish, Portuguese, Russian, Spanish, Swedish, and Simplified and Traditional Chinese) and, in addition, three
image examples for each topic. The objective is to retrieve as many relevant images as possible from the given
visual and multilingual topics.</p>
      <p>For this campaign, the main purpose of our experiments was to thoroughly study the different merging strategies,
i.e. methods of combination of visual and textual techniques. While we have applied all the well known
techniques which we had already used in previous campaigns [11] [12], for both textual and visual components
of the system, our research has focused on the idea to perform all possible combinations of those techniques in
order to evaluate which ones may offer the best results.</p>
      <p>All experiments are fully automatic, thus avoiding any manual intervention. None of the experiments
incorporates relevance feedback, due to time constraints. We finally submitted 110 multilingual textual
(textbased) runs, 22 visual (content-based) runs and 21 mixed runs (using a combination of both), elaborated on in
the following sections.</p>
    </sec>
    <sec id="sec-2">
      <title>2. System Description</title>
      <p>Based on our experience in previous campaigns, to be able to execute a large number of runs that exhaustively
cover all the combinations of the different techniques, we developed a flexible system, composed of a set of
small components which can be easily added in different configurations and are executed sequentially to build
the final result set.
the textual (text-based) retrieval module, which indexes the IAPR TC-12 image descriptions to look
for those descriptions that are more relevant to the text of the topic;
the visual (content-based) retrieval module, that provides the IAPR TC-12 images which are more
similar to the given topic images;
and, finally, the merging module, which uses different operators to combine the outputs of the two
previous subsystems to provide the final results.









2.1.</p>
    </sec>
    <sec id="sec-3">
      <title>Textual Retrieval</title>
      <p>Since MIRACLE has taken part in ImageCLEF (or, in general, in CLEF), different linguistic and statistical
techniques have been developed to be used in the text-based part of the different tasks [11] [12]. This year, the
main goal was to make an exhaustive study of these diverse methods, combining them in all possible ways and
testing the results achieved. A flexible and configurable system has been developed for this purpose.
The list of components includes:</p>
      <p>Proper Noun Detection. A heuristic based module to detect the appearance of entities in the text, based
on a finite state automaton.</p>
      <p>
        Linguistic Analyzer. A component to obtain morphosyntactic analysis (POS tagging) and
lemmatization. The target language considered in MIRACLE experiments has been English (although
the IAPR TC-12 collection was also available in German and Spanish) so the Charniak parser [
        <xref ref-type="bibr" rid="ref1">2</xref>
        ] has
been integrated. The combination of this parser with (Euro) Wordnet [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] has been used to provide
lemmatization.
      </p>
      <p>
        Stemming. Of course, a component to obtain stems for words has been included as part of the indexing
process. The implementation considered for this component has been the Porter’s stemmer [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Stopword Detection. One of the usual steps to take into consideration in the indexing process is the
exclusion of semantic empty words, also called stopwords [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Semantic Expansion. Words appearing in the query can be expanded using Wordnet [
        <xref ref-type="bibr" rid="ref4">5</xref>
        ] and an
implementation of a Word Sense Disambiguation algorithm [13]. Two working modes have been
considered, one where each word is expanded with all the words considered as synonyms by Wordnet,
and another one where a disambiguation process is carried out to try to filter out synonyms for the
wrong senses.
      </p>
      <p>
        Search Engine. Two different search engines have been used, Lucene [1] and Xapian [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
The former methods have been combined in all possible ways and applied to the text extracted from the fields
(either individually or combined) in which image descriptions are divided. As described in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], image captions
include the following fields: Title, Description, Notes, Location, Date, Image and Thumbnail. However, the
Description field is left out in the 2007 collection.
      </p>
      <p>The final goal of this system configuration was to evaluate which combination of techniques was the best,
considering last year relevance assessments. However, there were many possible combinations (although some
of them were illogic, like applying semantic expansion to stems). Therefore only those with the best results on
2006 data where selected.
2.2.</p>
    </sec>
    <sec id="sec-4">
      <title>Visual Retrieval</title>
      <p>
        For this part of the system, we resorted to two publicly and freely available Content-Based Information Retrieval
systems: GIFT (GNU Image Finding Tool) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and FIRE (Flexible Image Retrieval Engine) [
        <xref ref-type="bibr" rid="ref3">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ]. They are
both developed under the GNU license and allow to perform query by example on images, using an image as the
starting point for the search process and relying entirely on the image contents.
      </p>
      <p>In the case of GIFT, the complete image database was indexed in a single collection, downscaling each image to
32x32 pixels. For each topic, a visual query is made up of all the images contained in the topic. Then, this visual
query was used in GIFT to obtain the list of the most relevant images (i.e., images which are more similar to
those included in the visual query), along with the corresponding relevance values.</p>
      <p>On the other hand, we used the results of the FIRE system (kindly provided by the organizers), with no further
processing.
2.3.</p>
    </sec>
    <sec id="sec-5">
      <title>Result Merging</title>
      <p>The textual and image result lists are then merged and combined by applying different techniques, which are
characterized by an operator and a metric for computing the relevance (score) of the result. Table 1 shows the
defined operators: union (OR), intersection (AND) and external join (LEFT/RIGHT JOIN). Each of these
operators selects which images are part of the final result set. Then, results are reranked by computing a new
relevance measure value based on their corresponding input results, by using different metrics shown in Table 2.
score = max(a, b)
score = min(a, b)
score = avg(a, b)
score = max(a, b) + min(a, b) *</p>
      <p>min(a, b)
max(a, b) + min(a, b)</p>
    </sec>
    <sec id="sec-6">
      <title>3. Experiments and Results</title>
      <p>Experiments are defined by the choice of different combinations of the previously described modules, operators
and score computation metrics. All experiments are fully automatic, avoiding any manual intervention. None of
the experiments incorporates relevance feedback, due to time constraints.</p>
      <p>We finally submitted a wide set of experiments: 110 multilingual textual (text-based) runs, 22 visual
(contentbased) runs and 21 mixed runs (using a combination of both). The different name schemas for the run identifiers
are described in the following tables: textual retrieval experiments (Table 3), visual retrieval experiments (Table
4) and mixed experiments (Table 5).
(2) Method: TI (use image title index), DE (use image description index), NO (use image notes index), LO (use
image location index), NP (perform proper noun detection), S (apply stemming)
(3) Combination: OR (images in any result set), AND (images in both result sets), max (use maximum value),
min (use minimum value), avg (use average value)
(4) Topic language: DE (German), ES (Spanish), FR (French), JA (Japanase), PT (Portuguese), RU (Russian),
SV (Swedish)
(2) Combination: OR (images in any result set), AND (images in both result sets), max (use maximum value),
min (use minimum value), avg (use average value)
After the evaluation by the task organizers, results for each kind of experiment are presented in the following
tables. Each table shows the run identifier, the mean average precision (MAP), the precision at 10, 20 and 30
first results and the number of relevant images which have been retrieved (out of 3,416 relevant images).
Regarding text-based experiments (Table 6), the best results are achieved with the experiment where the title
(TI), description (DE) and location (LO) fields of the image caption have been indexed using Xapian (Xa) and
applying the proper noun detection module (NP) along with the Porter stemming algorithm (S). For this
experiment a MAP of 0.1995 is obtained. Last year, our top ranked text-based experiment had a MAP of 0.2005
with a similar combination of techniques. The only difference is that, while in 2006 topic narratives where also
used to build the query, this year there was no narrative provided for the topics. Implications of this are that no
improvement has been reached.
(1) out of 3,416 relevant images
In general, our experiments show Lucene turns out to give worse results than Xapian (the first run with Lucene
search engine has a MAP of 0.1871), although this conclusion has to be further studied. Moreover, no strategy
for merging results from both search engines have led to any improvement in MAP with respect to the baseline
experiments, neither with the OR operator (which was supposed to increase the number of results at the expense
of a loss of precision) nor the AND operator (which was supposed to increase precision).
(1) Only the first two results for each language are shown
1896
1902
1616
1616
1942
1946
1672
1567
1518
1533
1771
1791
1461
1446
The best results correspond to Spanish, our mother tongue in which we have a strong expertise, with a MAP
similar to the monolingual experiments. We are negatively surprised by the low precision with French and
German languages that may be attributed to a deficient stemming module. However, this issue has to be further
researched.</p>
      <p>The 2007 collection did not include the description field and our experiments have been produced using 2006
data collection (we did not notice this fact until the submission was over). Therefore our results are not really
comparable to the results provided by the rest of participants, shown in Table 8. However, our baseline
experiment would be among the best results this year and our group would be ranked 4th out of 16. Our
conclusion is that for the IAPR TC-12, textual descriptions of images are very convenient for the retrieval
process and should not be dropped out in for the next years.
Table 9 in next page shows the results of the visual-based experiments. In this case, MAP is very low due to the
complexity of the visual retrieval task in this domain. The best experiment is the combination of results with
GIFT and FIRE, and it is better than using GIFT alone. However, we regretfully did not execute the experiment
with FIRE alone, which would have been shown whether the combination of both systems improves the final
result. Also note that, although the number of relevant images retrieved using AND is lower than when OR is
used, the precision is higher.
Results from other participants, shown in Table 10, clearly outperform our own results. Our MAP is 56% lower
than the MAP of the best experiment and our group is ranked 7th out of 13 participants.
The best experiment uses a combination of all textual and visual results obtained in the previous experiments
with the LEFT operator along with the Max-Min (mm) scoring metric. Moreover, the mm metric clearly
outperforms the others (max, min, avg) even with the same operator over the same result sets, as seen in
Mix2LEFT experiments. The most interesting conclusion is that this merging strategy is successful as the MAP
is higher than in both the textual and visual partial experiments with which the experiment is built.
Finally, although our mixed experiments can really be compared to the others groups’ for the reasons already
stated before, our MAP is considerably lower than the MAP of the best mixed experiment (Table 12), with a
decrease to 70%. However, MIRACLE ranked 5th (out of over 10 participants), which is indeed considered to be
an acceptable position.</p>
      <p>cut-EN2EN-F50</p>
      <sec id="sec-6-1">
        <title>EN-EN-AUTO-FB-TXTIMG_MPRF</title>
      </sec>
      <sec id="sec-6-2">
        <title>DE-EN-AUTO-FB-TXTIMG_MPRF cut-EN2EN-F20 EN-EN-AUTO-FB-XTIMG_QTXT_COMBPREFFKTXT 0.3020 0.4300 0.3733 0.3306</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>4. Conclusions and Future Work</title>
      <p>Results in general show a poor performance for all groups, due to the characteristics of the image collection and
the difficulty of the defined topics.</p>
      <p>For the textual retrieval, the experiment where the title, description and location fields of the image caption have
been indexed using Xapian and applying the proper noun detection module and the stemming component has
produced the best results, with a MAP of 0.1995, which shows that no improvement has been reached with
regards to last year experiments when our top ranked text-based experiment had a MAP of 0.2005 with a similar
combination of techniques. In general, we obtain worse MAP with Lucene rather than Xapian, although this
topic has to be further investigated. Regarding the visual retrieval, our results are not good compared to other
groups’, but it is not strange as our areas of expertise do not specifically include image processing and we
resorted to publicly available “black-box” engines. In addition, no specific conclusion can be obtained as we
omitted to include the result for FIRE itself, so we could not evaluate if the merging strategy is successful or not.
However, our main interest was not in experiments where only text or image content is used in the retrieval
process. Instead, our challenge was to test whether the text-based image retrieval could improve the analysis of
the content of the image, or vice versa. The most interesting conclusion is that merging strategies are successful
as our best mixed experiment outperforms both the textual and visual experiments in which it is based, using the
LEFT operator for the combination along with the Max-Min (mm) metric for computing the relevance. This
result shows that our hypothesis was right. Our combination of a “black-box” search using publicly accessible
content-based retrieval engines with a text-based search has turned out to provide better results than other
presumably “more complex” techniques. This simplicity may be a good starting point for the implementation of
a real system.</p>
      <p>We are sure that there may still be some place for improvement with a more careful study of the parameters for
the merging strategy (both the combination operator and the score computing metric). In addition, none of the
experiments incorporates relevance feedback, due to time constraints. However, we are already carrying out
some experiments that incorporate this technique which are showing promising results, overcoming our
submissions.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work has been partially supported by the Spanish R+D National Plan, by means of the project RIMMEL
(Multilingual and Multimedia Information Retrieval, and its Evaluation), TIN2004-07588-C03-01; and by the
Madrid’s R+D Regional Plan, by means of the project MAVIR (Enhancing the Access and the Visibility of
Networked Multilingual Information for the Community of Madrid), S-0505/TIC/000267.
[1] Apache Lucene project. On line http://lucene.apache.org [Visited 10/08/2007].</p>
      <p>On
line</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Charniak</surname>
            ,
            <given-names>Eugene. A</given-names>
          </string-name>
          <string-name>
            <surname>Maximum-Entropy-Inspired Parser</surname>
          </string-name>
          .
          <source>In Proceedings of NAACL-2000</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>CLEF</given-names>
            <surname>2005 Multilingual Information</surname>
          </string-name>
          <article-title>Retrieval resources page</article-title>
          . On line http://www.computing.dcu.ie/ ~gjones/CLEF2005/Multi-8/ [Visited 10/08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ; Keysers; D.; Ney,
          <string-name>
            <given-names>H.</given-names>
            <surname>FIRE - Flexible Image</surname>
          </string-name>
          Retrieval Engine:
          <article-title>ImageCLEF 2004 Evaluation</article-title>
          . In CLEF 2004, LNCS 3491,
          <string-name>
            <surname>Bath</surname>
          </string-name>
          , UK, pp
          <fpage>688</fpage>
          -
          <lpage>698</lpage>
          ,
          <year>September 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Eurowordnet</surname>
          </string-name>
          :
          <article-title>Building a Multilingual Database with Wordnets for several European Languages</article-title>
          .
          <source>March</source>
          (
          <year>1996</year>
          ). On line http://www.illc.uva.nl/EuroWordNet/ [Visited 10/08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[6] FIRE: Flexible Image Retrieval System. ~deselaers/fire</article-title>
          .
          <source>html [Visited</source>
          <volume>10</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          http://www-i6.
          <article-title>informatik.rwth-aachen</article-title>
          .de/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] GIFT: The GNU Image-Finding Tool</article-title>
          . On line http://www.gnu.org/software/gift/ [Visited 10/08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Grubinger</surname>
            , Michael; Clough, Paul; Hanbury, Allan; Müller,
            <given-names>Henning.</given-names>
          </string-name>
          <article-title>Overview of the ImageCLEF 2007 Photographic Retrieval Task</article-title>
          .
          <source>Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Grubinger</surname>
            , Michael; Clough, Paul; Müller, Henning; Deselaers,
            <given-names>Thomas.</given-names>
          </string-name>
          <article-title>The IAPR-TC12 benchmark: A new evaluation resource for visual information systems</article-title>
          . In International Workshop OntoImage'
          <year>2006</year>
          <article-title>Language Resources for Content-Based Image Retrieval, held in conjunction with LREC'06</article-title>
          , pages
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
          , Genoa, Italy, May
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10] IAPR: On line http://www.iapr.org/ [Visited 10/08/2007] [11] [12] [13]
          <string-name>
            <surname>Martínez-Fernández</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Villena-Román</surname>
          </string-name>
          , Julio; García-Serrano, Ana M.;
          <string-name>
            <surname>Martínez-Fernández</surname>
          </string-name>
          ,
          <article-title>Paloma. MIRACLE team report for ImageCLEF IR in CLEF 2006</article-title>
          .
          <source>Proceedings of the Cross Language Evaluation Forum</source>
          <year>2006</year>
          , Alicante, Spain.
          <fpage>20</fpage>
          -
          <issue>22</issue>
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Martínez-Fernández</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Villena-Román</surname>
          </string-name>
          , Julio; García-Serrano, Ana M.;
          <string-name>
            <surname>González-Cristóbal</surname>
            ,
            <given-names>José</given-names>
          </string-name>
          <string-name>
            <surname>Carlos</surname>
          </string-name>
          .
          <article-title>Combining Textual and Visual Features for Image Retrieval</article-title>
          .
          <source>Accessing Multilingual Information Repositories: 6th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2005</year>
          , Vienna, Austria,
          <source>Revised Selected Papers. Carol Peters et al (Eds.). Lecture Notes in Computer Science</source>
          , Vol.
          <volume>4022</volume>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          ISSN:
          <fpage>0302</fpage>
          -
          <lpage>9743</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Montoyo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Método basado en marcas de especificidad para WSD</article-title>
          .
          <source>In Proceedings of SEPLN (Sociedad Española para el Procesamiento del Lenguaje Natural)</source>
          ,
          <source>nº 24</source>
          ,
          <year>September 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>Martin.</given-names>
          </string-name>
          <article-title>Snowball stemmers and resources page</article-title>
          . On line http://www.snowball.tartarus.
          <source>org [Visited</source>
          <volume>10</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15] University of Neuchatel.
          <article-title>Page of resources for CLEF (Stopwords, transliteration</article-title>
          , stemmers …). On line http://www.unine.ch/info/clef
          <source>[Visited</source>
          <volume>10</volume>
          /08/2007].
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <article-title>Xapian: an Open Source Probabilistic Information Retrieval library</article-title>
          . On line http://www.xapian.
          <source>org [Visited</source>
          <volume>10</volume>
          /08/2007].
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>