<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LIG at ImageCLEFphoto 2008</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Philippe Mulhem</string-name>
          <email>Philippe.Mulhem@imag.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>UJF, UMR CNRS 5217, Laboratoire d'informatique de Grenoble</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This working notes describe the runs and results obtained by the LIG at ImageCLEFphoto 2008. The submitted runs are: two runs (text only and text+image) without diversification on classes, and two runs (text only and text+image) with class diversification were submitted. The text retrieval is based on language model of Information Retrieval, and the image part is processed using RGB histograms on 9 image blocks with a similarity value based on Jeffrey divergence. Results using text+image are obtained by a linear combination of normalized results on text and image. The diversification is based on clusters, according to the cluster given in the queries. When the cluster name is not directly extracted from the images (like city or country), we apply a visual clustering. Not surprisingly, the cluster recall at 20 (i.e., cr(20)) results are higher for the runs that include diversification. On the other hand, the precision at 20 and the mean average precision results are higher without diversification on our runs, for both text only and image+text results.</p>
      </abstract>
      <kwd-group>
        <kwd>Text indexing and retrieval</kwd>
        <kwd>Image indexing and retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This paper describes the runs that where submitted at ImageCLEFphoto 2008 by the LIG
(Laboratoire d’Informatique de Grenoble). The runs submitted deal with text only and text+image
retrieval. The main idea behind our submissions is to make a first step into considering language
models for text and for image. Another aspect of this work is to check the impact of image
content on the quality of the retrieval. When considering clusters of images to ensure diversification
of the results, we applied simple solutions based on the analysis of the images description when
available, and we relied on visual clustering of the images when the clustering was not expressed
in the location field of the image description. Among the four runs submitted, our best results are
below the average ImageCLEFphoto results for the map (-5%) and the precision at 20 documents
(-3%), and the results obtained for the clustering recall at 20 documents are above the average
(+19%). This fact shows that the diversification based on simple features may impact positively
the results.</p>
      <p>The remaining of the paper is organized as follows. Section 2 describes the runs submitted
by focusing on the text processing, the image processing and the fusion and the processing of
text+image, section 3 details the results obtained, and we conclude in part 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the runs</title>
      <p>As descibed above, the runs submitted consider on ons sidetext only and text+image indexing
and retrieval, and the use of diversification or not on other side. We describe in this section these
different aspects.
2.1</p>
      <sec id="sec-2-1">
        <title>Text processing</title>
        <p>
          When considering text retrieval, it has been shown by [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] that language models (LM) of
Information Retrieval, inspired by speech recognition, give results that are close to, or outperform,
existing approaches (like Vector Space Model [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] or probabilistic models based on BM25 for
instance [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]). Different kinds of Language Models already exist, and we chose to use a language
model that expresses the probability for the query Q to be generated from a document model D
by : P(Q|D). Such probability is computed using the probability of any term ti to be generated
by the document, P(ti|D). One strong aspect of the LM is the use of smoothing leading to consider
that a document that does not contain a term does not have a zero probability of generating this
term. The smooting used in our experiments is the Dirichlet smoothing, that has been shown
in citezhailaferty04 to be a good smothing method. Such smoothing is described using the
following formula:
        </p>
        <p>P (ti|D) =
where
tf(ti,D)+μ. ttff((t∗i,,CC))</p>
        <p>tf(ti,D)+μ
• C is the corpus of documents considered,
• tf(x,D) denotes the term frequency of the term x in the document D,
• tf(x,D) denotes the term frequency of the term x in the corpus C,
• μ is a parameter defined experimentally that indicates the importance of the corpus on the
smoothing.</p>
        <p>A usual information retrieval preprocessing is performed: we remove stopwords from the image
descriptions and from the queries, and we stem the texts using a Porter stemmer. We use the
concatenation of the TITLE, DESCRIPTION and NOTES fields, in lower case, of the xml description
of the images as their text description.</p>
        <p>For the query processing, we used the query and the narrative as the text query, and the smae
preprocessing is applied on this text.</p>
        <p>As explained above, the query processing evaluates P(Q|D) as:</p>
        <p>P (Q|D) = P (tq,1, ..., tq,nq|D)
= Y P (tq,i|D)
i
(1)
The results are ranked according to the decreasing order of this probability.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Image processing</title>
        <p>The processing of the images computes histograms on image blocks. The result obtained on 9
blocks gave better results than taking the whole image in previous of our studies on personal
photographs.</p>
        <p>Each image of the corpus is split into a 3x3 regular grid. For each of the blocks bI,i, 1 ≤ i ≤ 9
of an image I, we extract an RGB histogram, HI,i, 1 ≤ i ≤ 9. The choice of the RGB color space
is due to the fact that it gives similar results than other color spaces that require more processing
time to be computed. Each histogram has 512 bins, according to a 8×8×8 regular split of the RGB
cube. Such size is a tradeof between too small histograms unable to discriminate visual elements,
storage uasge and matching duration. Then, one global histogram GHI with 512×9=4602 bins is
created for one image, by concatenating the 9 HI,i histograms. GHI is then normalized to 1.</p>
        <p>The matching function between two images I and J uses the Jensen-Shannon divergence. This
divergence is a symmetrical version of the Kullback-Leibler divergence. The formula of the
JensenShannon divergence between two histograms GHI and GHJ is:
The similarity Sim(I, J) between two images I and J is defined as 1 − J S(GHI ||GHJ ).</p>
        <p>In imageCLEFphoto, the visual part of a query is composed of 3 images, each image
representing one sample of what is expexted. We consider that the relevance of images in the corpus
depends on the best relevance value wrt. each of the query images. The similarity between a set
of images ISq = {Iq,i} and one image J is then defined as the maximum of the similarity between
J and each of the images of ISq.</p>
        <p>The results are then ranked according to the decreasing order of this similarity.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Image+Text processing</title>
        <p>Several ways may be used to represent mixed textual and visual data for retrieval. We consider
that even if the internal representation for these two media are somewhat similar (distribution of
probabilities), their inner nature refrain us to apply a kind of early fusion of them by concatenating
the two distributions. Another reason is that “adding” two distributions does not necessarily create
another one. That is why, in our run, we consider a late fusion defined as a linear combination of
the two results (text and image) obtained.</p>
        <p>
          In fact, the linear distribution is applied on normalized results, using for the text (resp. the
image) the minimum and the maximum matching values to ensure the normalized results to be in
[
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]. This linear combination is refined using the following heuristics: we assume that the best
visual results (i.e., very similar to one of the query images), are often relevant, which is not the
case for the text. So, we extend the matching function using a condition on the matching:
matching(Q, D) =
(1,
α.Simt(Qt, Dt) + (1 − α).Simv(Qv, Dv) othewise
if Simv(Qv, Dv) &gt; tv
(3)
with Qv (resp. Qt) the visual (resp. textal) part of the query, and Dv (resp. Dt) the visual (resp.
textal) part of the document.
        </p>
        <p>The results are then ranked according to the decreasing order of this similarity.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Diversification processing</title>
        <p>As one task for the imageCLEFphoto this year was dedicated to study the capacity of the systems
to provide diverse results for a query, and not only near duplicate results. To achieve this goal,
we defined clusters according to several criteria:
• country: based on the LOCATION field of the corpus images description, we generated a set
of clusters Ccountry, that groups images by country name. Ccountry contains 23 classes, with
on average 952 images per cluster.
• city: based on the LOCATION field of the corpus images description, we generated a set of
clusters Ccity, that groups images by city name and country when available. One cluster of
Ccity contains all the images with no city. Ccity contains 511 classes, with on average 38
images per cluster.
• clusters on date where also built. Because such clusters were not used during the query
processing we only mention them;
• visual: we used a KNN clustering, namely Cvisual, based on the visual description of the
images (4608 dimension histograms). The target number of cluster is 500, i.e. similar to the
number of clusters for the cities. There is 40 images per cluster on average.</p>
        <p>Depending on the query un ser consideration, a clustering is chosen for the diversification
process:
• Queries having a cluster name city are diversified using the Ccity clusters;
• Queries having a cluster name state or country are diversified using the Ccountry clusters;
• Queries having another cluster name are diversified using visual clusters Cvisual.
We describe now the way the diversification is processed, according to a clustering Cx:
1. From the top of the initial (not diversified) result list Lo, we generate a diversified list Ld
that contains only the best representative of one cluster of Cx according to the ranking of
Lo, until the diversified list contains 20 elements.
2. In a second step, we loop on Lo from its top, and we add to Ld the elements of Lo that do
not already belong to Ld.</p>
        <p>This process ensures that the 20 first results are from different clusters.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        We describe in this section the official results obtained by our approach. According to the
parameters described in the previous parts, the λ of the text language model is set to 1500, using
the Zettair system [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. For the runs text+image, we set the threshold tv to 0.99 and the α used
in the linear combination to 0.55 . We apply the diversification process explained earlier. To be
able to study the advantages or drawbacks of the diversification, we submitted diversified and not
diversified runs.
      </p>
      <p>The figure 1 presents the evolution of the precision for our results, and the figure 2 presents
the cluster recall for our four submitted runs. The runs without diversification (NOCLUST) are
presented with dotted lines in these two figures, where the plain lines correspond to diversified runs.</p>
      <p>We analyse first the results regarding the precision of figure 1. Without diversification, we see
in this figure that the text+image configuration outperforms the text only runs (+34% averaged
on the 6 values given by the official results).</p>
      <p>For the text only run, we see that for the precision after 15, 20 and 30 documents the precision
value is roughly constant (0.2171, 0.2026 and 0.2103). This means that at 15 documents the
systems gives 3.2 relevant documents, and at 30 documents the system gives 6.3 relevant documents
on average. On the other hand, the text+image run shows a monotonic decreasing of the precision.
Considering the diversified results, the text+image also outperfoms the text only based system
by 46.32% (averaged on the 6 values given by the official results). For both text+image and text
only results, we see that the precision is increased between 20 and 30 documents. We explain this
fact by our diversification process that manages in a different way the first 20 documents, and
then only adds the retrieved documents: because the first documents have more chances to be
relevant, such documents can be placed from the 20th document of the diversified list, increasing
the precision.</p>
      <p>The diversification process lowers the precision values obtained, for both the text+image and the
text only results. This can be explained by the fact that the diversification pushes up documents
that are not in the first results of the initial list, and these documents are potentially less relevant.
For the text+image results however, the difference after 30 documents is not large (0.2684 versus
0.2436), which means that on average only one additional relevant document is retrieved without
diversification compared to the results after diversification.</p>
      <p>We study then the cluster recall resuts presented in figure 2.</p>
      <p>As expected, the results obtained after diversification for text only (resp. text+image) outperform
the results obtained without diversification (+9.6% for text only retrieval, +5.7% for text+image,
averaged on the 8 values of the official results). At 20 documents, the difference is +16.7%
for the text only, and +10.8% for the text+image run. We see on the figure2 that, here also,
the text+image runs give better results than text only runs. However, after 50 documents, the
difference beween diversified and undiversified results is very small for the text+image runs: +2.0%
on average. In this case, the diversification is almost useless with respect to the cluster recall.
For the text only runs, a difference of 3.0% is achieved after 100 documents, leading to conclude
that the undiversified text results benefit more from the diversification than the texte+image
results. This is validated through the study of the average of the differences for diversified results
versus undiversified results up to the 30th document (i.e., for 5, 10, 15, 20 and 30): the text only
diversified run outperforms by +19.9% the undiversified text only, and the text+image diversified
run outperforms by +11.2% the text+image undiversified results.
The work described here lists the runs and results obtained by the LIG at ImageCLEFphoto 2008.
We experimented text and text+image retrieval. For the text, we considered a language model
using a Dirichlet smooting. For the image, we extracted RGB histograms of 9 blocks of the images,
ane the matching was performed using a Jeffrey divergence. For the retrieval and text+image, we
used a linear combination of the matching values for the text and for the images of the query. We
applied a simple diversification scheme in a way to force multiple clusters to be given for the 20
first results. We submitted results with and without diversification.</p>
      <p>When comparing our different runs, we found out that text+image runs consistently outperform
text only runs. For the precision values, undiversified results outeperfom diversified runs. For the
cluster recall, diversified runs outperfom undiversified results. So, we conclude that precision and
recall do not benefit both from the diversification process. Such antagonism between recall and
precision is new in IR, but we think that studying the effect of our diversification process, and
testing other ways to diversify (other clusters, smarter clustering, smarter reranking of the results)
may help to increase both precision and recall, at least below 30 documents.</p>
      <p>The results obtained are below the average of the runs with and without diversification for the
text only, according to the precision and the cluster recall. For the text+image runs, the results
are below the average when considering the non-diversified run, but above the average when using
diversification for the cluster recall.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgement</title>
      <p>This work was partially supported by the French National Agency of Research
(ANR-06-MDCA002).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Jay</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ponte</surname>
            and
            <given-names>W. Bruce</given-names>
          </string-name>
          <string-name>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>A language modeling approach to information retrieval</article-title>
          .
          <source>In Proceedings of the ACM SIGIR</source>
          , pages
          <fpage>275</fpage>
          -
          <lpage>281</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S E</given-names>
            <surname>Robertson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M M</given-names>
            <surname>Hancock-beaulieu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M</given-names>
            <surname>Gatford</surname>
          </string-name>
          .
          <source>Okapi at trec-3</source>
          . pages
          <fpage>109</fpage>
          -
          <lpage>126</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Gerard</given-names>
            <surname>Salton and M.J. McGill</surname>
          </string-name>
          .
          <article-title>Introduction to modern information retrieval</article-title>
          .
          <source>McGraw-Hill</source>
          ,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] The Zettair search engine</article-title>
          . http://www.seg.rmit.edu.au/zettair/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>