<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AVEIR at ImageCLEFphoto 2008: on the fusion of runs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sabrina Tollari</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcin Detyniecki</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marin Ferecatu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Herv´e Glotin</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philippe Mulhem</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Massih-Reza Amini</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ali Fakeri-Tabrizi</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Gallinari</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hichem Sahbi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhong-Qiu Zhao</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TELECOM ParisTech, UMR CNRS 5141 LTCI</institution>
          ,
          <addr-line>Paris</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universit ́e Joseph Fourier, UMR CNRS 5217 LIG</institution>
          ,
          <addr-line>Grenoble</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universit ́e Pierre et Marie Curie-Paris6, UMR CNRS 7606-LIP6</institution>
          ,
          <addr-line>Paris</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universit ́e du Sud Toulon-Var, UMR CNRS 6168 LSIS</institution>
          ,
          <addr-line>Toulon</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this working note, we present the submission of the AVEIR consortium, composed of 4 French laboratories, to ImageCLEFphoto 2008. The submitted runs correspond to different fusion strategies applied to four individual ranks, each proposed by an AVEIR consortium partner. In particular, we study the complete, and partial, average of the ranking values, the minimum of these values, and a random based diversification. We first briefly describe the individual run of each partner, then we describe the fusion runs. The official results classed one of the runs, the MEAN fusion, as the third best in the automatic text-image run category. This run gives better results than the best partner run.</p>
      </abstract>
      <kwd-group>
        <kwd>Rank Fusion</kwd>
        <kwd>Image Retrieval</kwd>
        <kwd>Multimodal Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>AVEIR (Automatic annotation and Visual concept Extraction for Image Retrieval) is the name
of a project supported by the French National Agency of Research (ANR-06-MDCA-002). A
consortium of four French CNRS research laboratories are involved in the project:
LSIS Laboratoire des Sciences de l’Information et des Syst`emes at the Universit´e du Sud
Toulon</p>
      <p>Var (USTV),
LTCI Laboratoire Traitement et Communication de l’Information at the TELECOM ParisTech.</p>
      <p>The overall goal of the project is to enrich image retrieval systems with semantic indexation
and annotation, and with symbolic relational description, all being automatically extracted and
built from the textual and image content extracted from documents or web pages. This semantic
and symbolic information are, then, used to reduce the visual ambiguity in images and to enhance
the retrieval of images from large databases. The project develops 3 research axes. The first axis
focuses on image analysis, feature extraction and visual feature representations. The second axis
is concerned with automatic labeling of image components or objects with textual concepts. The
third axis considers image retrieval and evaluation of the proposed algorithms. For more details
please refer to http://aveir.lip6.fr.</p>
      <p>
        In order to compare the state of the art approaches, each of the partners participated
individually to ImageCLEFphoto (cf. [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
        ]).
      </p>
      <p>The particularity of the 2008 ImageCLEFphoto edition was its focus on diversity. The
evaluation was based on two measures: precision at 20 and instance recall at rank 20 (also called cluster
recall or S-recall), which calculates the percentage of different classes or clusters represented in
the top 20. The idea behind these measures was to focus on relevant but diverse - in terms of
clusters - images.</p>
      <p>In order to analyze if combining different runs improves the diversity, a submission under the
label AVEIR was proposed. In this paper we briefly discuss the former submission, in particular
the different fusion strategies, and the results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of individual runs</title>
      <p>
        Although each of the partners had its own diversification strategy, for the fusion we used the non
diversified runs. In table 1, we briefly describe each the used runs. For more details please refer
to the specific papers:
LIG histo 3 p o 1.5 4 0 NOCLUST EN-EN-AUTO-TXTIMG [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]: this run is based on
the linear combination of the scores provided by a language model using Dirichlet smoothing
on the text and by a Jeffrey-Divergence correspondence on the images.
      </p>
      <p>
        UPMC-LIP6 r3tfidf VCDTWN EN-EN-AUTO-TXTIMG [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: the text processing is based
on standard TF-IDF with cosine similarity. Forest of Fuzzy Decision Trees (FFDT) trained
on VCDT ImageCLEF task 2008 are used for a visual concept filtering of the textual results.
      </p>
      <p>
        The matching of the concepts and the topics text used WordNet
LSIS EN-EN-AUTO-TXTIMG-AUTO GLOZHA ar 12 NOCLUST [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: the visual
features are entropic features. Lots of SVMs are trained and generated with different parameters
using the sample images provided. Then the first 20 images of the LIG run are used as the
positive samples for each topic, and the others as the negative samples to construct the
validation set for selecting the best one among the generated SVMs.
      </p>
      <p>
        PTECH-EN-EN-AUTO-TXTIMG-AMKNR [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: the run uses a combination of text and
image descriptors. For a given topic, a separate query is performed for each modality (text
and image). The results are merged by a minimum rank criterion: each image keeps the
best rank.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Description of AVEIR runs</title>
      <p>The AVEIR consortium proposed, for ImageCLEFphoto2008, 4 runs each with a different fusion
strategy. Since for each partner’s run, we have at most 1000 images ranked by topic, some images
are sometimes not ranked.</p>
      <p>AVEIR LIG LIP6 LSIS PTECH EN-EN-AUTO-TXTIMG MIN: for each image, the
fusionrank corresponds to the minimum rank observed on each of the 4 partner’s runs. This
strategy corresponds to creating a rank by alternatively choosing an image from each of the
partners’ runs. The first image of the fusion rank corresponds to the first image of the first
partner; the second image corresponds to the first image of the second partner; the fifth
corresponds to the second image of the first partner, and so on.</p>
      <p>AVEIR LIG LIP6 LSIS PTECH EN-EN-AUTO-TXTIMG MEAN: for each image, the
fusion-rank corresponds to the average rank observed on each of the 4 partner’s runs. This
strategy corresponds to a compromise taking into account all the systems. Images not present
in one of the ranked lists are considered as having rank 1001.</p>
      <p>AVEIR LIG LIP6 LSIS PTECH EN-EN-AUTO-TXTIMG MEAN2on4: here only
images that were ranked by at least two partners where considered. The fusion-rank correspond
to the average of the available ranks. The idea behind this strategy is to avoid fusionning
images returned only by one partner.</p>
      <p>AVEIR LIG LIP6 LSIS PTECH EN-EN-AUTO-TXTIMG MEAN DIVALEA40: the
first 40 images of the MEAN run were randomly shuffled. The objective of this run is to
observe how randomness affects diversity and to provide a baseline for the instance recall.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results and discussion</title>
      <p>Figure 1 compares Precision and Cluster Recall when considering the first n retrieved images. The
average precision at 20 (P20) and the average cluster recall at 20 (CR20), of the best 4 runs from
each participating group (25 groups and 100 runs), was respectively P20= 0.32 and CR20= 0.35.
All the fusion strategies are above these scores. This may be explained by the fact that some of
the partners’ runs performed very well.</p>
      <p>When comparing MEAN and MEAN DIV (for n &lt; 40) on figure 1, we conclude that a random
diversification worsens the results as well for the precision as for the cluster recall. In terms of
precision, the best fusion strategy is the MEAN, the worse being the MIN. In other words, from
the precision point of view, it is more interesting to base the fusion on a compromise. In fact, the
MIN strategy considers an image as very good as long as one of the partners, independently of
the others, ran it high. The best images, when using the MEAN strategy, correspond to images
that were highly ranked by all the systems.</p>
      <p>Surprisingly, from the cluster precision perspective, in average (over all the topics), there is
not much difference between the runs, although the MIN slightly outperforms the other strategies
(in particular when considering the very first images). If we look at figure 2(a), we discover that
there are topics for which the MIN strategy is better and topics for which the MEAN is better.
Although the reasons behind this behaviour needs further research, it explains why in average
there is no difference between the two strategies.</p>
      <p>Table 2 compares the best individual run with the AVEIR fusion runs. Only the MEAN strategy
shows an improvement with respect to the best of the individual runs. The Mean Average Precision
is clearly improved. There is no improvement in the cluster recall, actually there is a slight drop.
The explanation lies in the behaviour per topic. On figure 2(b), we observe that for some topics
the best individual run outperforms any fusion, while for others the fusion improves beyond the
best individual run. The fusion is not correlated to best individual run. Furthermore, the topics
that have a high cluster recall score (i.e. with a score higher than 0.5 for as well for the MEAN
as for the Best Individual) are better served by the best individual run, while the ones with a low
score are better with the MEAN fusion. The compromise pays, in terms of diversity, when the
problem is difficult. This may be explained by our previous observation that the MEAN improves
the precision. In fact, for difficult topics, the MEAN brigs new images up, increases the precision
and the cluster recall, since a new relevant image belongs with a high probability to a new class.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this working note, we presented the submission, of the AVEIR consortium, to ImageCLEFphoto
2008. The particularity of this year edition was its focus on diversity. The evaluation was based
on the relevance, measured by the precision at 20 and by the diversity measured by the cluster
recall at rank 20. The idea behind these two measures was to focus on relevant but diverse images.</p>
      <p>The submitted runs correspond to different fusion strategies applied to four individual ranks,
each proposed by a partner. In particular we study the complete, and partial, average of the
rank values (MEAN and MEAN2on4), the minimum of these values (MIN), and a random based
diversification (DIVALEA40). The official results1 classed one of the runs, the MEAN fusion, as
the third best. Our experiments showed why this fusion particularly improves the precision and,
even more, the mean average precision. We also observed that it only slightly affects the diversity.</p>
      <p>Furthermore, the MIN fusion - which corresponds to alternating images from each individual
run - despite its weak precision at 20, improves slightly the overall diversity. The weak precision
at 20 may be partially explained by the disparity, in terms of quality, of the runs. In fact, low
scoring runs bring non-relevant images, lowering the precision, but also keeping the diversity of
the runs’ information.</p>
      <p>Finally, the experiments also pointed out that the diversity is strongly affected by the relevance,
in particularly for difficult queries. Although the experiments showed that, in terms of diversity,
the best individual run performs better than any type of fusion, we observed that for low precision
topics it is more interesting to perform a MEAN fusion (that increases the mean average precision)
and that for high precision topics it is more interesting to fusion with the MIN (as long as the
runs have similar performance). In other terms diversity comes after a good relevance.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment</title>
      <p>This work was supported by the French National Agency of Research (ANR-06-MDCA-002).
5
15
59
21491</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Marin</given-names>
            <surname>Ferecatu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hichem</given-names>
            <surname>Sahbi</surname>
          </string-name>
          . TELECOM ParisTech at ImageClefphoto 2008:
          <article-title>Bi-modal text and image retrieval with diversity enhancement</article-title>
          .
          <source>In Working Notes of ImageCLEFphoto2008</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] Herv´e Glotin and
          <string-name>
            <surname>Zhong-Qiu Zhao</surname>
          </string-name>
          .
          <article-title>Affinity propagation promoting diversity in visuo-entropic and text features for clef photo retrieval 2008 campaign</article-title>
          .
          <source>In Working Notes of ImageCLEFphoto2008</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Philippe</given-names>
            <surname>Mulhem</surname>
          </string-name>
          . LIG at ImageCLEFphoto
          <year>2008</year>
          . In Working Notes of ImageCLEFphoto2008,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Sabrina</given-names>
            <surname>Tollari</surname>
          </string-name>
          , Marcin Detyniecki, Ali Fakeri-Tabrizi,
          <article-title>Massih-Reza Amini, and Patrick Gallinari</article-title>
          . UPMC/LIP6 at ImageCLEFphoto 2008:
          <article-title>on the exploitation of visual concepts (VCDT)</article-title>
          .
          <source>In Working Notes of ImageCLEFphoto2008</source>
          ,
          <year>2008</year>
          .
          <article-title>0.2 0.4 0.6 0.8 cr20 TELECOM ParisTech (b) Best individual run vs best fusion strategy (MEAN)</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>