<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Diversity in Photo Retrieval: Overview of the ImageCLEFPhoto Task 2009</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Monica Lestari Paramita</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Sanderson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Clough</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>m.paramita</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>m.sanderson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>p.d.clough}@sheffield.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms Measurement</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Performance</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Experimentation</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Sheffield</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ImageCLEF Photo Retrieval Task 2009 focused on image retrieval and diversity. A new collection was utilised in this task consisting of approximately half a million images with English annotations. Queries were based on analysing search query logs and two different types were released: one containing information about image clusters; the other without. A total of 19 participants submitted 84 runs. Evaluation, based on Precision at rank 10 and Cluster Recall at rank 10, showed that participants were able to generate runs of high diversity and relevance. Findings show that submissions based on using mixed modalities performed best compared to those using only concept-based or content-based retrieval methods. The selection of query fields was also shown to affect retrieval performance. Submissions not using the cluster information performed worse with respect to diversity than those using this information. This paper summarises the ImageCLEFPhoto task for 2009.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Performance Evaluation</kwd>
        <kwd>Image Retrieval</kwd>
        <kwd>Diversity</kwd>
        <kwd>Clustering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Evaluation Objectives for 2009</title>
      <p>
        The Photo Retrieval task in 2009 was focused at studying diversity further. Using resources from Belga, we
provided a much larger collection, containing just under half a million images, compared to 20,000 images
provided in 2008. We also obtained statistics on popular queries submitted to the Belga website in 2008 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
which we exploited to create representative queries for this diversity task. We experimented with different ways
of specifying the need for diversity which was given to participants, and this year decided to release half of the
queries without any indication of diversity required or expected. We were interested in addressing the following
research questions:
•
•
•
      </p>
      <p>Can results be diverse without sacrificing relevance?
How much will knowing about query clusters a priori help increase diversity in image search results?
Which approaches should be used to maximize diversity and relevance for image search results?
These research questions will be discussed further in section 4.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Framework</title>
      <p>
        One of the major challenges for participants of the 2009 ImageCLEFPhoto task was a new collection which was
25 times larger than that used for 2008. Query creation was based completely on query log data, which helped to
make the retrieval scenario as realistic as possible [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We believe this new collection will provide a framework
in which to conduct a more thorough analysis of diversity in image retrieval.
2.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Document Collection</title>
      <p>The collection consists of 498,920 images with English-only annotations (i.e. captions) describing the content of
the image. However, different to the structured annotations of 2008, the annotations in this collection are
presented in an unstructured way (Table 1). This increases the challenge for participants as they must
automatically extract information about the location, date, photographic source, etc of the image as a part of the
indexing and retrieval process. The photos cover a wide-ranging time period, and there are many cases where
pictures have not been orientated correctly, thereby increasing the challenge for content-based retrieval methods.
20090126 - DENDERMONDE, BELGIUM: Lots of people
pictured during a commemoration for the victims of the
knife attack in Sint-Gilles, Dendermonde, Belgium, on
Monday 26 January 2009. Last friday 20-Year old Kim De
Gelder killed three people, one adult and two childs, in a
knife attack at the children's day care center "Fabeltjesland"
in Dendermonde. BELGA PHOTO BENOIT DOPPAGNE
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Query Topics</title>
      <p>
        Based on search query logs from Belga, 50 example topics were generated and released as two query types (as
mentioned previously). From this set, we randomly chose 25 queries to be released with information including
the title, cluster title, cluster description and image (example) as shown in Table 2. We refer to these queries as
Query Part 1. In this example, participants can notice that this result about ‘Clinton’ requires 3 different clusters,
which are ‘Hillary Clinton’, ‘Obama Clinton’ and ‘Bill Clinton’. Results covering other aspects of “Clinton”,
such as Chelsea Clinton or Clinton Cards, will not be counted towards the final diversity score. More
information about these clusters and the method used to produce them can be found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Given that one might argue that the diversity result in Query Part 1 could be relatively easy to produce as
detailed information about the different sub-topics is provided as part of the query topic and there are often in
practice instances when little or no query log information is available to indicate possible clusters, we released
25 queries containing no information about the kind of diversity expected (referred to as Query Part 2). An
example of this query type is given in Table 3. It should be noted that information about the cluster titles and
description were also based on Belga’s query logs. However, we did not release any of this information to the
participants.
&lt;top&gt;
&lt;num&gt; 26 &lt;/num&gt;
&lt;title&gt; obama &lt;/title&gt;
&lt;image&gt; belga30/06098170.jpg &lt;/image&gt;
&lt;image&gt; belga28/06019914.jpg &lt;/image&gt;
&lt;image&gt; belga30/06107499.jpg &lt;/image&gt;
&lt;/top&gt;
The list of 50 topics used in this collection is given in Table 4. Since Belga is a press agency based in Belgium,
there are a large number of queries which contain the names of Belgian politicians, Belgian football clubs and
members of the Belgian royal family. Other queries, however, are more general such as Beckham, Obama, etc.
There are some queries which are very broad and under-specified (e.g. Belgium); others are highly ambiguous
(e.g. Prince and Euro).
* = ambiguous, ** = under-specified queries, bold queries: queries with more than 677 (median) relevant
documents
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Relevance Assessments</title>
      <p>
        Relevance assessments were performed using the DIRECT (Distributed Information Retrieval Evaluation
Campaign Tool)2, a system which enables assessors to work in a collaborative environment. We hired 25
assessors to be involved in this process and assessments were divided into 2 phases: in the first phase, assessors
were asked to identify images relevant to a given query. Information about all relevant clusters to the topic was
given to assessors to ensure they were aware of the scope of relevant images for a query. The number of relevant
images for each query resulting from this stage is shown in Figure 1.
Having queries from different types shown in Table 4, we then analysed the number of relevant documents in
each type. This data, shown in Table 5, illustrates that under specified queries have the highest average number
of relevant documents.
After a set of relevant images were found, for the second stage different assessors were asked to find images
relevant to each cluster (some images could belong to multiple clusters). Since topics varied widely in content
and diversity, the number of relevant images varied from 1 to 1,266 for each cluster. Initially, there were 206
clusters created for the 50 queries, but this number dropped to 198 as there were 8 clusters with no relevant
images which had to be deleted. There are an average number of 208.49 relevant documents for each cluster,
with a standard deviation of 280.59. The distribution of clusters is shown in Figure 2.
2 http://direct.dei.unipd.it
The method for generating results from participant’s submissions was similar to that used in 2008 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
precision of each run (P@10) was evaluated using trec_eval and cluster recall (CR@10) was used to measure
diversity. Since the maximum number of clusters was set to 10 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we focussed evaluation on P@10 and
CR@10. The F1 score calculates the harmonic mean of these two measures.
3
      </p>
    </sec>
    <sec id="sec-7">
      <title>Overview of Participation and Submissions</title>
      <p>A total of 44 different institutions registered for the ImageCLEFPhoto task (the highest number of applications
ever received for this task). From this number, 19 institutions from 10 different countries finally submitted runs
to the evaluation. Due to the large number of runs received last year, we limited the number of submitted runs to
5 per participant. A total of 84 runs were submitted and evaluated (some groups submitted less than 5 runs).
3.1</p>
    </sec>
    <sec id="sec-8">
      <title>Overview of Submissions</title>
      <p>The participating groups for 2009 are listed in Table 8. From the 24 groups participating in the 2008 task, 15
groups returned and were involved this year (Returning). We also received four new participants who joined this
task for the first time (New).</p>
      <p>Participants were asked to specify the query fields used in their search and the modality of the runs. Query fields
were described as T (Title), CT (Cluster Title), CD (Cluster Description) and I (Image). The modality was
described as TXT (text-based search only), IMG (content-based image search only) or TXT-IMG (both text and
content-based image search). The range of approaches is shown in Tables 6 and 7 and summarised in Figure 3.</p>
      <p>Modality
Number of Runs
Institution
University of Alicante
Hungarian Academy of Science, Budapest
Computer Science, Trinity College, Dublin
Computational Linguistics at Concordia (CLAC)
Lab, Concordia University, Montreal
Interactive Information Access
Computer Science Faculty, Daedalus, Madrid
Multimedia IR, University of Glasgow
Lab. Informatique Grenoble
Language Tech
Institution for InfoComm Research
LEAR Team
Intelligent Systems, University of Jaen
Intelligent System Group, Daedalus, Madrid
NLP, AI.I.Cuza U. of IASI
Electronics and Computer Science, University of
Southampton
Department of Computer Science, Laboratoire
d’Informatique de Paris 6
System and Information Sciences Lab, France
Wroclaw University of Technology
XEROX Research</p>
      <p>Country
Spain
Hungary
Ireland
Canada
Netherlands
Spain
UK
France
Mexico
Singapore
France
Spain
Spain
Canada
UK
France
France
Poland
France</p>
      <p>Runs
5
5
4
4
5
5
5
4
5
5
5
4
3
5
4
5</p>
    </sec>
    <sec id="sec-9">
      <title>Results</title>
      <p>This section provides an overview of the results based on the type of queries and modalities used to generate the
runs. As mentioned in the previous section, we used P@10 to calculate the fraction of relevant documents in the
top 10 and CR@10 to evaluate diversity, which calculates the proportion of subtopics retrieved in the top 10
documents as shown below:</p>
      <p>Cluster − recall at K ≡</p>
      <p>UiK=1 subtopics(di )</p>
      <p>nA
F1 =
2 x (P10 x CR10)</p>
      <p>(P10 + CR10)
The F1 score was used to calculate the harmonic mean of P@10 and CR@10, to enable the results to be sorted by
one single measure:
4.1</p>
    </sec>
    <sec id="sec-10">
      <title>Results across all Queries</title>
      <p>The top 10 runs computed across all 50 queries (ranked in descending order of F1 score) are shown in Table 9.
Looking at the top 10 runs, we observe that highest effectiveness is reached using mixed modality (text and
image) and using information from the query title, cluster title and the image content itself. The scores for P@10,
CR@10 and F1 in this year’s task are notably higher than the evaluation last year. Moreover, the number of
relevant images in this year’s task was higher. Having two different types of queries, we analysed how
participants dealt with the different queries. Tables 10 and 11 summarise the top 10 runs in each of query types.
Different compared to results presented previously, it is interesting to see that the top run in Queries Part 1 used
only text retrieval approaches. Even though the CR@10 score was lower than most of the runs, it obtained the
highest F1 score due to a high P@10 score. The uses of tags vary within results, but the top 9 runs consistently
use both title and cluster title. We therefore conclude that the use of title and cluster title do help the participants
to achieve a good score in both precision and cluster recall.</p>
      <p>In the queries part two, participants did not have access to cluster information. We specifically intended this to
see how well the system finds diverse results without any hints. The results of the top runs in queries part 2 is
shown in Table 11.
* submitted results for 24 out of 25 queries. Score shown is the average of the submitted queries only.
It is shown in the table that the top 9 runs use information from example images, which shows that example
images and their annotations might have given useful hints to detect diversity. To analyse this further, we
divided the runs which used the Image field and those which did not, and found that the average CR@10 scores
were 0.5571 and 0.5270 respectively. We conclude that having example images helps to identify diversity and
present a more diverse set of results.</p>
      <p>Comparing the CR@10 scores in the top 10 runs of Queries Part 1 and Queries Part 2, the scores in the latter
group were lower, which implied that systems did not find as many diverse results when cluster information was
not available. The F1 scores from these top 10 were also lower, but they only differed slightly compared to the
Queries Part 1. We also calculated the magnitude of difference between results for different query types (shown
in Table 12). This indicates that on average runs do perform lower in Query Part 2, however the difference is
small and not sufficient to conclude that runs will be less diverse if cluster titles are not available (p=0.146).
It is important to understand that not all the runs in Query Part 1 use the cluster title. To analyse how useful the
“Cluster Title” (CT) information is, we divided the runs of Query Part 1 based on the use of CT field. The mean
and standard deviation of P@10, CR@10 and the F1 scores is shown in Table 13 (the highest score shown in
italics).
Number
of Runs
Mean
0.6845
0.6641
0.6315</p>
      <p>SD
0.2
0.2539
0.2185
Table 13 provides more evidence that the Cluster Title field has an important role in identifying diversity. When
Cluster Title is not being used, the F1 scores of both Query Part 1 and Query Part 2 do not differ significantly.
Figure 3 shows a scatter plot of F1 scores for each query type. Using a two-tailed paired t-test, the scores
between Queries Part 1 and Queries Part 2 were found to be significantly different (p=0.02). There is also a
significant correlation between the scores: the Pearson correlation coefficient equals 0.691.
We evaluated the same test on the runs using Cluster Title only to the runs in Query Part 2, and found that they
are also significantly different (p=0.003), the Pearson correlation coefficient equals 0.745. However, when the
same evaluation was being performed on runs not using Cluster Title, the difference in scores was not significant
(p=0.053), although obtaining a Pearson correlation coefficient of 0.963.</p>
      <p>SD
0.2088
0.2208
0.2185</p>
      <p>Mean
0.5467
0.5583
0.5415
We also analysed whether the number of clusters have any effect on the diversity score. To measure this factor,
we calculated the mean CR@10 for all of the runs. These scores are then plotted based on the number of clusters
contained in each specified query. This scatter plot, shown in Figure 5, has a Pearson correlation coefficient of
-0.600, confirming that the more clusters a query contains, the lower the CR@10 score is.
4.2</p>
    </sec>
    <sec id="sec-11">
      <title>Results by Retrieval Modality</title>
      <p>In this section, we will present an overview result of runs using different modalities.</p>
      <p>Modality
TXT-IMG
TXT
IMG</p>
      <p>Number of</p>
      <p>Runs
Mean
0.713
0.698
0.103
According to Table 15, both the precision and cluster recall scores are highest if systems use both low-level
features based on the content of an image and its associated text. The mean of the runs using image content only
(IMG) is drastically lower based on the P@10 score; however the gap decreases when considering only the
CR@10 score. Further research should be carried out to improve runs using content-based approaches only, as
the best run using this approach had the lowest F1 score (0.218) compared to TXT (0.351) and TXT-IMG
(0.297).
4.3</p>
    </sec>
    <sec id="sec-12">
      <title>Approaches Used by Participants</title>
      <p>Having known that the mixed modality performs best, we were also interested to see the best combination of
query fields to maximize the F1 score of the runs. We therefore calculated the mean of each combination and
modality and the result is shown in Table 16 with the highest score for each modality shown in italic.
It is interesting to note that the highest F1 score was different for each modality. A combination of T-CT-I had
the highest score in TXT-IMG modality. In the TXT modality, a combination of T-I scored the highest, with
TCT-I following on the second place. However, since only one run used the T-I, it was not enough to provide a
conclusion about the best run. Calculating the average F1 score regardless of diversity shows that the best runs
are achieved using a combination of Title, Cluster Title and Image. Using all tags in the queries resulted in the
worst performance.
5</p>
    </sec>
    <sec id="sec-13">
      <title>Conclusions</title>
      <p>This paper has reported the ImageCLEF Photo Retrieval Task for 2009. Still focusing on the topic of diversity,
this year’s task introduced new challenges to the participants, mainly through the use of a much larger collection
of images than used in previous years and by other tasks. Queries were released as two ‘types’: the first type of
queries included information about the kind of diversity expected in the results; the second type of queries not
providing this level of detail.</p>
      <p>The number of registering participants in this year was the highest of all the ImageCLEFPhoto tasks since 2003.
Nineteen participants submitted a total of 84 runs, which were then categorised based on the query fields used to
find information, and the modalities being used. The result showed that participants were able to present a
diverse result without sacrificing precision. In addition, results showed the following:
•
•
•</p>
      <p>Information about the cluster title is essential for providing diverse results, as this enables participants
to correctly present images based on each cluster. When the cluster information was not being used, the
cluster recall score is proven to drop, which showed that participants need better approach to predict the
diversity need in it.</p>
      <p>A combination of Title, Cluster Title and Image was proven to maximize the diversity and relevance of
the search engine.</p>
      <p>Using mixed modality (text and image) in the runs managed to achieve the highest F1 compared to
using only text or image features alone.</p>
      <p>Considering the increasing interest of participants in ImageCLEFPhoto, the creation of the new collection was
seen as a big achievement in providing a more realistic framework for the analysis of diversity and evaluation of
retrieval systems aimed at promoting diverse results. The findings from this new collection were found to be
promising and we plan to make use of other diversity algorithms in the future to enable evaluation to be done
more thoroughly.</p>
    </sec>
    <sec id="sec-14">
      <title>Acknowledgments</title>
      <p>We would like to thank Belga Press Agency for providing us the collection and query logs and Theodora
Tsikrika for the preprocessed queries which we used as the basis for this research.</p>
      <p>The work reported has been partially supported by the TrebleCLEF Coordination Action, within FP7 of the
European Commission, Theme ICT-1-4-1 Digital Libraries and Technology Enhanced Learning (Contract
215231).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <year>2009</year>
          . Queries Submitted by Belga Users in
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Paramita</surname>
            ,
            <given-names>M. L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Developing a Test Collection to Support Diversity Analysis</article-title>
          . SIGIR 2009 Workshop: Redundancy, Diversity, and Interdependent Document Relevance,
          <source>July 23rd</source>
          , Boston, Massachusetts, USA.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Arni</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Grubinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Overview of the ImageCLEFPhoto 2008 Photographic Retrieval Task. Cross Language Evaluation Forum</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>