<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The CLEF 2011 Photo Annotation and Concept-based Retrieval Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefanie Nowak</string-name>
          <email>research@stefanie-nowak.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karolin Nagel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Judith Liebetrau</string-name>
          <email>judith.liebetrau@idmt.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer Institute for Digital Media Technology (IDMT) Ehrenbergstr.</institution>
          <addr-line>31, 98693 Ilmenau</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ImageCLEF 2011 Photo Annotation and Concept-based Retrieval Tasks pose the challenge of an automated annotation of Flickr images with 99 visual concepts and the retrieval of images based on query topics. The participants were provided with a training set of 8,000 images including annotations, EXIF data, and Flickr user tags. The annotation challenge was performed on 10,000 images, while the retrieval challenge considered 200,000 images. Both tasks di erentiate among approaches that consider solely visual information, approaches that rely only on textual information in form of image metadata and user tags, and multi-modal approaches that combine both information sources. The relevance assessments were acquired with a crowdsourcing approach and the evaluation followed two evaluation paradigms: per concept and per example. In total, 18 research teams participated in the annotation challenge with 79 submissions. The concept-based retrieval task was tackled by 4 teams that submitted a total of 31 runs. Summarizing the results, the annotation task could be solved with a MiAP of 0.443 in the multimodal con guration, with a MiAP of 0.388 in the visual con guration, and with a MiAP of 0.346 in the textual con guration. The conceptbased retrieval task was solved best with a MAP of 0.164 using multimodal information and a manual intervention in the query formulation. The best completely automated approach achieved 0.085 MAP and uses solely textual information. Results indicate that while the annotation task shows promising results, the concept-based retrieval task is much harder to solve, especially for speci c information needs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>With the increasing amount of digital information on the Web and on
personal computers, the need for systems that are capable of automated indexing,
searching, and organising multimedia documents incessantly grows. Automated
systems have to retrieve information with high precision in order to be accepted
by industry and end-users. Often, multimedia retrieval systems are evaluated on
di erent test collections with di erent performance measures, which makes the
comparison of retrieval performance impossible and limits the bene t of the
approaches. Benchmarking campaigns counteract these tendencies and establish an
objective comparison among the performance of di erent approaches by posing
challenging tasks and by distributing test collections, topics, and measures.</p>
      <p>This paper presents an overview of the ImageCLEF 2011 Photo Annotation
and Concept-based Retrieval Tasks. The two tasks aim at the automated
detection of visual concepts in consumer photos and the retrieval of photos based on
a certain topic. Section 2 introduces the tasks and the evaluation methodology.
Following, Section 3 discusses the visual concepts and query topics which
simulate the user's information need in image search. Then, Section 4 describes the
test collection and the relevance assessment process. Section 5 summarizes the
approaches of the participants. Following, the results for the annotation task and
the concept-based retrieval task are presented and discussed in Section 6 and
Section 7, respectively. Finally, Section 8 summarizes and concludes this paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task Description</title>
      <p>
        The Photo Annotation and Concept-based Retrieval Tasks pose an image
analysis challenge which consists of two sub tasks. The annotation task aims at the
automated annotation of consumer photos with multiple concepts. It is similar
to the visual concept detection and annotation task (VCDT) as it was posed
in the last years [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. This year, the participants are asked to annotate a test
set of 10,000 Flickr images with 99 visual concepts. To solve this task, an
annotated training set of 8,000 images is provided. The evaluation considers a fully
assessed test collection to compare the approaches of the participants. The
second challenge poses a concept-based retrieval task. The participants are asked to
retrieve (up to) the 1,000 most relevant images in ranked order for a given topic
out of a test collection of 200,000 images. In total 40 topics, each consisting of a
logical connection of concepts from the annotation task, are provided. Concept
detectors may be trained on the training set of the annotation task (8,000
images annotated with 99 visual concepts). The assessment incorporates a pooling
strategy with crowdsourced relevance assessments. Both tasks can be solved by
following three di erent approaches:
1. Automatic annotation with visual information only (\V")
2. Automatic annotation based on Flickr user tags and image metadata (\T")
3. Multi-modal approaches that consider visual information and/or Flickr user
tags and/or EXIF information (\M")
      </p>
      <p>
        The participants can choose one task or participate in both. Both tasks make
use of a subset of the MIR Flickr 1 Million image dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The MIR Flickr
collection supplies all original tag data provided by the Flickr users (further
denoted as Flickr user tags). These Flickr user tags are made available for the
textual and multi-modal approaches of both subtasks. For most of the photos,
the EXIF data is included and may be used.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Evaluation Objectives</title>
        <p>The main evaluation objectives of the two tasks in 2011 lie in the exploitation of
di erent knowledge sources, the bene t of annotation approaches as part of the
retrieval process, and the automated prediction of subjective concepts such as
sentiments. Moreover, participants need to deal with an unbalanced amount of
data per concept, a varying number of labels per image, the diversity of image
content per concept, and the di erent qualities of image metadata.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Ontology</title>
        <p>
          The novel sentiment concepts are included in the Photo Tagging ontology [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] of
the last years. The hierarchy allows making assumptions about the assignment
of concepts to documents. Additionally, other relationships between concepts
determine possible label assignments. The ontology restricts, for instance, the
simultaneous assignment of some concepts (disjoint items) or de nes that one
concept postulates the presence of other concepts. The ontology allows the
participants to incorporate semantic knowledge in their annotation algorithms, and
to make assumptions about probable concept combinations.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Evaluation Measures</title>
        <p>
          In the annotation task, the evaluation sticks to the concept-based and
examplebased evaluation paradigm. For the concept-based evaluation, the Mean
interpolated Average Precision (MiAP) is utilized, while the example-based evaluation
applies the example-based F-Measure (F-Ex). Additionally, we introduce a novel
performance measure called Semantic R-Precision (SR-Precision) which is based
upon the example-based R-Precision, but incorporates the Flickr Tag
Similarity (FTS) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to determine the semantic relatedness of visual concepts in the
case of misclassi cations. R-Precision calculates Precision at perfect Recall in an
example-based evaluation scenario. The SR-Precision variant assigns
misclassication costs based on the semantic relatedness among misclassi ed concepts.
The semantic relatedness is derived from the FTS measure. In contrast to the
Ontology Score with Flickr Context Similarity (OS-FCS) which was used in
2010 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], the SR-Precision is able to incorporate ranked predictions instead of
forcing the systems to provide binary decisions. However, this measure requires a
normalization of classi er scores over di erent classi er outputs to deliver
meaningful results. This requirement was not explicitly posed to the participants and
therefore algorithms might not be optimally parameterized for this measure.
        </p>
        <p>The concept-based retrieval task evaluates performance on a test collection
with incomplete relevance judgments. All submissions of the participants are
pooled by using a pool depth of 100 documents per topic and run. Finally, the
runs are evaluated with the Mean uninterpolated Average Precision (MAP),
Precision@10 (P@10), Precision@20 (P@20), Precision@100 (P@100), and
conceptbased R-Precision (R-Prec) with the trec eval 8.1 program.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Incorporation of user needs in the evaluation</title>
      <p>Topics and visual concepts are strongly related to user needs and de ne the
use cases of a system. While concepts are modality-independent (i.e., the event
\birthday" might be detectable in the visual modality (birthday cake, people
celebrating) as well as in the auditory modality (people singing a birthday song)),
visual concepts are solely described by the visual content of a photo and are
therefore language independent. This section introduces the visual concepts that
are applied in the annotation task and the derivation of query topics for the
retrieval task based on these concepts, query logs, and related work. Please note
that the process of collecting images and de ning visual concepts is di erent
from related work. While usually the concept lexicon exists before images are
collected, in the case of the ImageCLEF VCDT test collection, this process is
decoupled and the images have been collected rst. This approach is much closer
to reality and poses new challenges, as objects are not necessarily centred in the
image and the distribution of images per concept varies considerably.
3.1</p>
      <sec id="sec-3-1">
        <title>De nition of visual concepts</title>
        <p>
          The test collection for the annotation task contains manual annotations for 99
visual concepts. These concepts describe the scene (indoor, outdoor, landscape...),
depicted objects (car, animal, person...), the representation of image content
(portrait, gra ti, art...), events (travel, work...), or quality issues (overexposed,
underexposed, blurry...). This year, a special focus is laid on the detection of
sentiment concepts. All in all, 49 concepts of the 53 concepts used in 2009 [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] were
utilized again. The concept Canvas as well as the concepts No Visual Season,
No Visual Place, and No Visual Time were discarded in this year's challenge.
The 41 concepts which were added in 2010 are all reused. In 2011, nine novel
sentiment concepts were added to the test collection. For the de nition of
sentiments, we follow the approach of Russell [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], who de nes an emotional space with
two dimensions (arousal and valence) on which emotional adjectives/sentiments
can be placed. Valence spans from the negative pole \misery" to the positive pole
\pleasure" on the x-dimension, while arousal spans from \passive" to \active" in
the y-dimension. In this model, adjectives are grouped into eight a ect concepts
in circular order. The model was slightly adapted and an additional concept
funny was included. The eight sentiments are structured according to their
degrees in the circle as proposed by Russel. Partly, the wording is changed as to
better t the sentiments to describe images (e.g., an image cannot be excited or
astonished, but it may look exciting to a human being). Starting with happy at
0 , the circle is further composed of funny (about 30 ), euphoric (70 ), active
(90 ), scary (150 ), unpleasant (180 ), melancholic (210 ), inactive, (270 ), and
calm/comforting (330 ).
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>De nition of topics for the concept-based retrieval task</title>
        <p>
          Based on the visual concepts, 40 topics for the concept-based retrieval task
were constructed. We conceived that each topic contains a di erent number of
relevant images, and that the topics comprise a range of di culty levels. For
the de nition of relevant topics, we followed two approaches: First, we adapted
topics from the ImageCLEF Wikipedia retrieval task [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], as these topics
were designed based on web-query logs and because they range from simple to
semantic (hence highly di cult) topics as described in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. A total of 17 topics
were directly applicable to our test data. Second, we examined interesting queries
for the test collection. Based on the output for each query, it was decided if the
chosen topic comprises an adequate occurrence in the test collection. The 40
resulting topics and their source are shown in Table 1. Sample images of the
dataset were taken for clari cation and provided as examples for the topics.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Ground Truth Acquisition</title>
      <p>The relevance assessments for the annotation task and the concept-based
retrieval task were acquired with a crowdsourcing approach using Amazon
Mechanical Turk1 (MTurk). MTurk is an online marketplace which distributes mini-jobs
to an unde ned crowd of people. At MTurk these mini-jobs are called HITs
(Human Intelligence Tasks). The workers at MTurk, called turkers, can choose the
HITs they would like to perform and submit the results to MTurk. The requester
of the work collects the results from MTurk and approves or rejects the work
of the turkers. Experiences with MTurk from ImageCLEF 2010 show the
applicability of crowdsourcing for ground truth acquisition of image labels. This
year, additional quality assurance mechanisms were incorporated to reduce the
impact of spammers on the annotations.
4.1</p>
      <sec id="sec-4-1">
        <title>Design of the annotation HIT template</title>
        <p>The assessment of the sentiment concepts was performed by asking the turkers
what sentiments an image conveys. The HIT template includes a de nition of
sentiments, synonymous sentiments, and example images (see Figure 1). The
de nitions are derived from WordNet 3.02 and the Free Dictionary3. Each survey
comprises ten images. The image is depicted on the left, while on the right the
adapted circumplex model of Russel (see Section 3.1) is visualised, as illustrated
on the example of one image in Figure 2. The option no sentiment should be
chosen if no sentiment ts to the image. After selecting this checkbox, the turkers
were asked to give a mandatory reason why no sentiment ts. We included
this question in the survey to prevent turkers from clicking at this checkbox
without thinking about the task. For all other sentiments, several choices could
be selected at the same time. Additionally, the turkers were asked which reason
let them decide for a sentiment: the motif or the overall impression of the image.
They could choose on a ve-point scale with the scales \motif" { \mostly motif"
{ \both equally" { \mostly overall impression" { \overall impression".</p>
        <p>The HIT template includes an automated veri cation procedure. For all ten
images that belong to one HIT, it is veri ed that the survey is completely lled
out before the submission of the task works. In the case of missing answers, the</p>
        <sec id="sec-4-1-1">
          <title>1 www.mturk.com 2 http://wordnetweb.princeton.edu/perl/webwn, last accessed 20.07.2011 3 http://www.thefreedictionary.com/, last accessed 20.07.2011</title>
          <p>turkers see the corresponding questions marked in red. This procedure ensures
that it is not too easy to answer randomly and submit spam, and it helps reducing
our work to lter out incomplete answers that need to be republished. While it
does not assure that all random annotators are excluded (as turkers still can
randomly answer each question), this at least assures that it also costs some
amount of work to cheat compared to the time that is needed to answer honestly.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Assessment statistics of the annotation task</title>
        <p>The ground truth was acquired in di erent annotation batches. The pretest
included 400 images of the training set arranged in surveys of ten images per HIT.
Each HIT was annotated three times by a total of 22 turkers in an average
annotation time of 3 minutes and 12 seconds and paid with 0.05$. The purpose of
the pretest was to understand if the template design and the task were
understandable and if the turkers were able to solve the task. Results show that in
about 50% of the images the turkers are agreeing on the sentiment (or choosing
neighbouring sentiments) while in the other 50%, they chose opposite sentiments
(like happy and melancholic). The rest of the training set of the Photo
Annotation Task was annotated in altogether 4,225 HITs. Each HIT contained nine
photos of the training set and one photo of the pretest as gold standard. The
gold standard was built by a majority vote of the pretest images excluding the
no sentiment concept and randomly placed into each HIT survey. Each HIT was
annotated ve times and rewarded with 0.07$. On average, they were completed
in 2 minutes and 36 seconds by a total of 258 distinct turkers. The test set was
assessed in 5,560 HITs which each included nine images and one gold standard
image of the training set. Each image of the test set was annotated ve times.
The HITs of the test set were divided into two batches (in order not to pose too
many HITs at the same time) and annotated by 156 distinct turkers. Each HIT
was rewarded with 0.07$, again. For the rst batch of 2,745 HITs, each HIT was
annotated on average in 2 minutes and 8 seconds, while the 2,815 HITs of the
second batch were annotated in average in 1 minute and 44 seconds.</p>
        <p>The veri cation of the work of the turkers is di cult, as the task of sentiment
annotation is very subjective. Therefore, we followed several strategies on how
to compare the annotations. The veri cation of the HITs of the training and
test set with the gold standard images lead to a direct acceptance of 3,204 and
4,358 HITs, respectively, allowing a deviation of at most 90 on the a ect circle.
For HITs that did not pass the gold standard test, we compared the results of
the HIT to the four answers of other turkers for the same HIT. For all images
of the HIT, the deviation to the annotations of the other HITs was computed
and the HITs were accepted when the deviation was equal or less to 90 on the
a ect circle per image. A total of seven out of the 10 images had to t. With
this procedure all remaining HITs could be accepted.</p>
        <p>The nal construction of the ground truth considers the majority vote for
each image. In the case that no clear answer was given, we decided to discard
any sentiment information for that image. In total, about 15% of the training
set images and 14% of the test set images have no sentiment information.
Interestingly, the no sentiment option was rarely chosen by the turkers. For none of
the images of the training set and only for one image of the test set a majority
of people decided for this concept.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Design of the topic HIT templates</title>
        <p>In the relevance assessment of the concept-based retrieval task, the turkers were
asked to mark all relevant images on the HIT template for a given topic. Each
HIT template includes a de nition of the topic and example images (see
Figure 3). A HIT contains 22 images plus two gold standards images, which were
used as a means of reliability control for the assessments. For each topic, we
selected one image that ts the de nition and one image that is not relevant
for the given topic. Special attention was taken in the design of the irrelevant
images per topic. Instead of using images that are clearly out of the scope of the
topic, images that match the meaning of the topic quite close, but not exactly,
were chosen; see Figure 4 for sample images of the topic sh in the water. The
gold images were placed randomly in the HIT templates.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Assessment statistics of the retrieval task</title>
        <p>The number of HITs per topic is dependent on the total number of distinct
images that were retrieved by the runs of the participants. Each HIT contained 24
images and was assessed by three turkers; so in total 7,868 HITs were processed.
Accepted HITs were paid with 0.03$. Each topic was processed by at least ve
(topic 29) and at most 41 (topic 12) distinct turkers. The average grading time
varies per topic between 31 seconds (topic 25) and 1 minute and 19 seconds
(topic 15). To increase the reliability of the relevance judgments, the results
were subjected to a post screening procedure. The assessments of a turker were
rejected if the gold standard images were not marked correctly. These HITs
were published again until for all HITs three reliable results were available. In
the next step, the assessments of the three turkers per HIT were compared with
each other. As the task of selecting relevant images for a topic is a subjective
task, its veri cation is di cult. In an additional step, we visualized the images
that were assessed as relevant and estimated the number of false assignments and
missing assignments. Depending on these results, the number of votes that were
necessary to de ne an image as relevant were chosen for each topic. For most
topics, a majority vote from at least two of the three assessors was necessary. A
minority vote was used for only ve topics, while all assessors had to agree on
relevance for three topics.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Participation</title>
      <p>Altogether, 48 groups registered for the challenge. 42 groups signed the license
agreement and were provided with the test collections. For the annotation task,
18 groups submitted results in altogether 79 runs. The number of runs was
restricted to a maximum of 5 runs per group. In total, there were 46 submissions
using only visual information, 8 submissions using only textual information, and
25 submissions utilising multi-modal approaches. For the retrieval task, 4 groups
submitted results in a total of 31 runs. The maximum number of runs per group
was set to 10. The submissions include 14 visual runs, 7 textual runs, and 10
multi-modal runs. The runs can be subdivided into 16 runs that retrieved all
images in a completely automated fashion and 15 runs that included a manual
intervention in the query generation step or relevance feedback. All participants
that submitted to the retrieval task also took the challenge in the annotation
task. The teams and their approaches are brie y introduced in the following:</p>
      <p>
        BPACAD [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: The team of the Computer and Automation Research
Institute of the Hungarian Academy of Science submitted one textual, two visual
and three multi-modal runs to the annotation task. Their approach is based
on a kernel weighting procedure using visual Fisher kernels and a Flickr-tag
based Jensen-Shannon divergence based kernel. Classi cation uses a linear SVM
trained for each concept separately.
      </p>
      <p>BUFFALO: The team of the University at Bu alo, New York, USA
submitted ve visual runs for the annotation task. They follow two approaches: the
rst considers a local linear coordinate method to learn concepts with a
regression method. The second uses a combination of GIST and colour features and
classi es the images by a neural network.</p>
      <p>
        CAEN [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]: The group of University of Caen, France participated with four
visual runs in the annotation task. The proposed approach uses visual image
features, such as SIFT, HOG, Texton, LAB, SSIM, and Canny, and aggregated
them by a Bag-of-Words (BoW) model into a global histogram. Fisher Vectors
and contextual information were used as enhancement of the BoW-models. The
classi cation considers SVM models trained for each concept separately.
      </p>
      <p>
        CEALIST [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: The team from the Laboratory of Vision and Content
Engineering, France submitted one textual, one visual, and three multi-modal runs
to the annotation task. The textual descriptor is based on semantic similarity
between tags and visual concepts. Two distances were used: one based on the
Wordnet ontology and one based on social networks. The visual component
considers various local and global features, such as Fisher vectors as well as colour
and edge features. Late fusion was used to combine visual and textual modalities.
      </p>
      <p>DBIS: The team of the Technical University of Cottbus, Germany submitted
ve runs in the visual con guration to the annotation task. They use various
features and investigate the in uence of several parameters in clustering on the
annotation performance.</p>
      <p>
        HHI [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]: The team of Fraunhofer HHI, Berlin, Germany submitted ve
visual runs to the annotation task. Their approach is based on the BoW model. A
feature fusion of the opponent SIFT descriptor and the GIST descriptor was done
in order to improve the classi cation performance of scene-based concepts. HHI
investigates a sampling of informative images in the training procedure, which
resulted in qualitative as well as runtime performance gains. A post-classi cation
processing step is incorporated, which re nes classi cation results based on rules
of inference and exclusion between concepts.
      </p>
      <p>
        IDMT [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]: The group of Fraunhofer IDMT, Ilmenau, Germany submitted
one textual and four multi-modal runs to the annotation task. Their approach
focuses on the fusion of multi-modal information and the exploitation of Flickr
user tags. As visual features, they employ RGB-SIFT features in a codebook
approach and classify the images with a one-against-all strategy using a SVM
with RBF kernel.
      </p>
      <p>
        ISIS [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]: The Intelligent Systems Lab of the University of Amsterdam,
The Netherlands participated with ve runs in the annotation task (3V, 2M)
and ten runs (10V) in the retrieval task. All runs of the annotation task use
several colour SIFT features with Harris-Laplace and dense sampling, and apply
the SVM classi er. The multi-modal runs further include binary vectors for the
most frequent Flickr user tags. In the retrieval task, three runs are computed
completely automated and seven include a manual intervention by following
two approaches. In the fully automated runs, a combination of the provided
positive example images and random irrelevant images were used to train the
concept detector. For the human topic mapping, a human reads the topic and
then selects the relevant concept(s). The probability scores of these concepts are
then combined using either summation or multiplication. In the human topic
inspection approach, relevance feedback was used to improve results.
      </p>
      <p>
        LAPI [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]: The group of Laboratorul de Analiza si Prelucrarea Imaginilor,
Universitatea Politehnica Bucuresti, Bucharest, Romania submitted two runs
using a visual-only approach. They combine colour and structural features and
adopt a Linear Discriminant Analysis for classi cation. Post-processing considers
joint probabilities of concept occurrences in the training set for label elimination.
      </p>
      <p>
        LIRIS [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]: The group of Universite de Lyon, CNRS, France participated in
the annotation task with two textual, one visual, and two multi-modal runs. They
consider two textual descriptors: one is based on a semantic distance between the
text and an emotional dictionary, the other one contains the valence and arousal
meanings by making use of the A ective Norms for English Words dataset. In
the visual approaches, di erent visual features including colour, texture, shape,
and high level aesthetic features are applied. Performance is compared using
di erent fusion strategies as well as Adaboost and SVM classi ers.
      </p>
      <p>
        MEIJI [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]: The group of Meiji University, Kanagawa, Japan submitted
ve runs (2V, 1T, 2M) to the annotation task and ten completely automated
runs (2V, 2T, 6M) to the retrieval task. Their approach is based on visual word
co-occurrence using the BoW model and global colour features as well as textual
features derived by tf-idf weigths of Flickr user tags. Classi cation is performed
by an adaptation of the so-called confabulation model.
      </p>
      <p>
        MLKD [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]: The Machine Learning and Knowledge Discovery group of
the Aristotle University of Thessaloniki, Greece participated in the annotation
task with ve runs (1V, 1T, 3M) and in the retrieval task with two automated
and eight semi-automated runs (2V, 4T, 4M). They approach the photo
annotation challenge with multi-label learning algorithms based on Random Forests
as base classi er. The visual features consider seven local descriptors with two
sampling strategies. The textual models are based on a Boolean BoW
representation including word stemming, stop words removal, and feature selection. The
multi-modal approach considers a hierarchical late-fusion of the modalities. For
the concept-based retrieval task two approaches were used: one based on the
concept relevance scores in a manual con guration and one automated approach
which is based solely on the sample images using textual information.
      </p>
      <p>
        MRIM [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]: The team of Grenoble University, France submitted four runs
(3V, 1M) to the annotation task. Classi cation considers multiple SVM classi ers
with RBF kernel. In the visual runs, several global and local colour and texture
descriptors are applied and dimension reduction techniques are investigated. The
multi-modal run additionally considers Flickr user tags as simple textual features
in a late fusion of SVM classi er scores.
      </p>
      <p>MUFIN [21]: The Faculty of Informatics, Masaryk University, Brno, Czech
Republic participated with four multi-modal runs in the annotation task. Their
approach is based on a free-text annotation system that assigns arbitrary words
to web images by visual and textual neighbour searching. For the textual search,
the EXIF data and image descriptions were used, while the visual search
considers di erent MPEG-7 descriptors. The search considers the Pro media dataset
to nd the nearest neighbours and transfers its annotations to the ImageCLEF
test collection including a removal of stopwords and names. The resulting words
were transformed into the xed set of 99 visual concepts of the annotation task
with the help of WordNet and the provided ontology.</p>
      <p>NII [22]: The team of the National Institute of Informatics, Tokyo, Japan
participated with ve visual runs in the annotation task. Their models are using
global and local features. As for global features, colour moments, colour
histogram, edge orientation histogram, and local binary patterns are applied. As
for local features, keypoint detectors such as Harris Laplace, Hessian Laplace,
Harris A ne, and Dense Sampling are used to extract SIFT-descriptors.
Classication is performed with a SVM classi er.</p>
      <p>REGIMVID [23]: The research group on Intelligent Machines,
University of Sfax, Tunisia submitted one textual run to the annotation task and one
textual, automated run to the retrieval task. Their approach focuses on the
exploration of Flickr tags to extract contextual relationships of tag relations.
Therefore, two types of contextual graphs are modeled: an inter-concepts graph
and a concept-tags graph.</p>
      <p>TUBFI [24]: The joint submission of the Machine Learning Group, Berlin
Insitute of Technology and Fraunhofer FIRST Berlin, Germany consists of four
visual and one multi-modal run to the annotation task. Classi cation considers
non-sparse multiple kernel learning and multi-task learning. Di erent extensions
of the BoW models with respect to sampling strategies and BoW mappings were
proposed. The multi-modal run further considers frequent Flickr user tags based
on a soft mapping for textual BoWs and Markov random walks over tags.</p>
      <p>UNIKLU: The team of the Institute of Information Technology,
AlpenAdria University, Klagenfurt, Austria participated with four visual runs in the
annotation task. They made use of the LIRE framework and applied several
features such as SIFT, SURF, MSER, CEDD, FCTH, and colour histograms
and classi ed the images with a linear SVM. Two of the runs incorporate an
automated post-processing of the classi cation results.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Annotation Task: Results</title>
      <p>This section illustrates the results for the annotation subtask. First, the overall
results of all teams are presented, independent of the con guration. In the
following subsections the results per con guration are highlighted. The results for
all runs can be found at the Photo Annotation Task website4.</p>
      <p>The task was solved best with a MiAP of 0.443 (TUBFI), followed by a MiAP
of 0.437 (LIRIS) as illustrated in Table 2. Both runs make use of multi-modal</p>
      <sec id="sec-6-1">
        <title>4 http://www.imageclef.org/2011/photo</title>
        <p>Team Rank SR-Prec. Con g.</p>
        <p>ISIS</p>
        <p>CAEN
BPACAD</p>
        <p>HHI</p>
        <p>LIRIS
TUBFI
MLKD
MRIM</p>
        <p>IDMT
BUFFALO</p>
        <p>DBIS
CEALIST</p>
        <p>MEIJI
UNIKLU</p>
        <p>MUIN
LAPI</p>
        <p>NII
REGIMVID
information. Table 3 depicts the overall rankings for the results of the evaluation
per example. The best results in terms of F-Ex are achieved in a multi-modal
con guration with 0.622 (ISIS), followed by 0.600 F-Ex (CAEN) which makes
use of a visual con guration. In terms of SR-Precision, the best run scores with
0.742 SR-Precision (ISIS) in a multi-modal run, followed by 0.729 SR-Precision
(BPACAD) in a visual run.</p>
        <sec id="sec-6-1-1">
          <title>Results for the visual con guration</title>
        </sec>
        <sec id="sec-6-1-2">
          <title>Results for the textual con guration</title>
          <p>The results for the textual runs are presented in Table 5. The best run scores
with 0.346 MiAP (BPACAD), followed by 0.326 MiAP (IDMT, MLKD). In the
example-based evaluation the best run scores with 0.525 F-Ex (IDMT) and 0.677
SR-Precision (IDMT) followed by 0.506 F-Ex (MLKD) and 0.676 SR-Precision
(CEALIST, LIRIS).</p>
        </sec>
        <sec id="sec-6-1-3">
          <title>Results for the multi-modal con guration</title>
        </sec>
        <sec id="sec-6-1-4">
          <title>Comparison of achievements with di erent information sources</title>
          <p>Last year only two runs considered the textual con guration. In contrast, this
year eight textual runs were submitted by seven teams. This allows for a more
reliable analysis of the performance of textual runs in image annotation. The
performance of textual runs is close to the results that can be achieved in the
visual con guration. The best visual run achieves a MiAP of 0.388 in contrast to
the best textual run, which scores with a MiAP of 0.346. The di erence of 4.2% is
rather small, especially when considering that not for all images EXIF data and
Flickr user tags exist. In the example-based evaluation, the di erence between
visual and textual runs is more signi cant. Visual runs score better by about 9%
in terms of F-Ex and about 6% in terms of SR-Precision. Results in the
multimodal con guration outperform classi cation with single modality information
in the visual con guration by 5.5% and the textual con guration by about 10%
in terms of MiAP. For the example-based measures F-Ex and SR-Precision,
di erences are very small with 1% for the visual con guration. Comparing the
multi-modal to the textual con guration, di erences are signi cant and lie by
about 10% and 6.5% for F-Ex and SR-Precision, respectively.</p>
        </sec>
        <sec id="sec-6-1-5">
          <title>Annotation performance per concept</title>
          <p>In Table 7, the results for each concept are summarized independent of the
con guration. On average, the concepts could be detected with a MiAP of 0.48
considering the best run per concept out of all con gurations. In general, 79
concepts were best detected with a multi-modal approach, 17 concepts were detected
best with a visual approach, and 3 concepts were detected best by a textual
approach. High performance is achieved for the concepts Neutral-Illumination,
No-Persons, No-Blur, and Outdoor. Following, the concepts Sky, Day, Clouds,
and Plants were annotated with high scores. The worst annotation quality was
achieved for the concept abstract followed by the concepts work, gra ti,
technical, old-person, and boring.</p>
          <p>In the evaluation in 2010, a great di erence in prediction quality among
the concepts from 2009 and the ones newly introduced in 2010 of 0.57 MiAP
and 0.37 MiAP could be seen. This di erence is still present in this evaluation
cycle. The concepts from 2009 (number 1-49) could be detected with a MiAP
of 0.57 considering the best prediction for each concept out of all runs, while
the 2010 concepts (numbers 50-90) improved minimally to a MiAP of 0.38. The
new sentiment concepts (numbers 91-99) can be detected with a MiAP of 0.39.
Although these are arguably very subjective concepts, the detection algorithms
are capable of identifying a strong trend of sentiments correctly. Especially, the
sentiments calm and inactive could be detected very well, while the sentiments
scary and euphoric were annotated worst. However, one has to note that the
2010 concepts occur on average in 7.9% of the training set images, while the 2009
and 2011 concepts are visible in 18% and 14% of the training set, respectively.
Therefore, the algorithms had more example images to learn the sentiments in
comparison to the more object-based concepts introduced in 2010.
7</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Concept-based Retrieval Task: Results</title>
      <p>In the following, the results of the concept-based retrieval task are presented
and discussed. The participation of four teams was lower with respect to the
annotation task. Despite, 31 runs have been submitted in all con gurations,
consisting of 10 multi-modal, 7 textual, and 14 visual runs. Approximately half
of the systems used a completely automated processing.</p>
      <p>Table 8 depicts all runs and indicates the con guration and the degree of
automation of each run. Results are sorted in terms of MAP. The MAP value
ranges from 0.1640 to 0.0013. Overall, the task was solved best with a MAP of
0.164 by the MLKD group, who used a multi-modal con guration with manual
query formulations. It can clearly be seen that the approaches using manual
processing achieve better results than the automated versions. The best performing
automated run achieves a MAP of 0.0849 (MLKD). The multi-modal and
textual con gurations of MLKD work signi cantly better than their visual ones.
MLKD provides also the best MAP value for a textual con guration, which is
only 0.0094 points lower than the overall best run. The best working visual
conguration, also with manual query formulation, was submitted by ISIS, achieving
a MAP of 0.0997. MEIJI provided solely automated runs, for which the
multimodel approaches outperform the textual and visual runs. REGIMVID provided
one automated, text-based con guration, which achieves a MAP value of 0.0042.</p>
      <p>Table 9 presents the best run per team in the three con gurations. If a
team submitted an automated and a manual run in the same con guration, the
results for both runs are illustrated. The direct comparison of automated and
manual runs per team shows that the manual runs work best, independent from
the con guration (textual, visual, multi-modal). This is mostly apparent in the
big di erence { nearly factor two { of the textual con guration of MLKD. The
results of the two approaches submitted by ISIS support this interpretation.
The manually processed run outperforms the automated approach with a MAP
of 0.0997 vs. 0.0430. The performance di erence between the two multi-modal
con gurations submitted by MKLD and MEIJI can partially be explained by the
di erent degree of automation. MEIJI uses a fully automated system resulting</p>
      <p>Team</p>
      <p>MAP
ISIS 0.0997</p>
      <p>ISIS 0.0430
MLKD 0.0361
MEIJI 0.0017
MLKD 0.1546
MLKD 0.0849</p>
      <p>MEIJI 0.0227
REGIMVID 0.0042
in a MAP value of 0.0444 and MLKD uses a manually query formulation which
achieves a MAP value of 0.1640 (best run overall).</p>
      <p>The boxplots of the MAP scores of all approaches per topic in Figure 5
allow for a more detailed examination. Most topics show a wide variation among
the obtained MAP values, which indicates the di erences in performance per
topic. This can be seen, e.g., for the topics 33 (cars and motion blur ) and 28
( reworks). For topic 33, the lowest MAP scores are 0. The highest value (0.5218)
for this topic is achieved by ISIS with an automated, visual con guration. MAP
scores for topic 28 are in the range of 0 to 0.423. A closer look reveals that
only for two topics (33 and 5) MAP values over 0.5 are reached. This is most
surprising in the case of topic 5 (riders on horse), due to the consistent low MAP
values of the other approaches. The six runs performing signi cantly better than
the rest are provided by MLKD and use textual and multi-modal information.
The gure also shows that the MAP scores for topic 18 (on the far right) are
homogeneously low. It can be concluded that all approaches had great di culties
identifying relevant images for female old person. The topics 13, 21 and 26 were
also hard to identify (MAP close to 0) for most of the approaches, but better
performing outliers clearly exist. For topic 13 (female person(s) doing sports )
some approaches of MLKD reach values above 0.14. The same observation, with
values close to 0.1, applies to topic 21 (scary dog(s)). The most striking e ect
can be observed for houses in mountains (26): Only two ISIS approaches were
able to achieve signi cantly higher MAP values than 0. The best sentiment topic</p>
      <p>MAPs, sorted by descending Median of single Topics</p>
      <p>No. of Topic
in terms of MAP is topic 34 (unpleasant insects). Relatively high scores above
0.4 are achieved by several approaches.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>The ImageCLEF 2011 Photo Annotation and Concept-based Retrieval Tasks
posed two image analysis challenges that could be solved with three general
con gurations: textual-based analysis, visual-based analysis, and multi-modal
analysis. The aim of the annotation task was to automatically annotate images
with 99 concepts in a multi-label scenario. The task attracted a considerable
number of international teams with a nal participation of 18 teams that
submitted a total of 79 runs. The results show that the annotation task could be
solved reasonably well, with the best multi-modal run achieving a MiAP of 0.443
in the multi-modal con guration, a MiAP of 0.388 in the visual con guration,
and a MiAP of 0.346 in the textual con guration. For the evaluation per
example, the best multi-modal run achieves 0.62 F-Ex, the best visual run scores with
0.61 F-Ex, and the best textual run with 0.53 F-Ex. All in all, the multi-modal
approaches got the best scores for 79 out of 99 concepts, followed by 17 concepts
that could be detected best with the visual approach and 3 that won with a
textual approach. In general, the multi-modal approaches outperformed visual
and textual con gurations for nearly all performance measures of the teams that
submitted results for more than one con guration.</p>
      <p>The concept-based retrieval task asked participants to retrieve the most
relevant images given certain topics. The topics were constructed based on user
needs and query logs, and consist of a Boolean connection of several visual
concepts. In total, 4 teams participated in this novel challenge and submitted 31
runs. 10 runs belong to the multi-modal con guration, 14 runs were submitted in
the visual con guration, and 7 runs are based on textual information. The best
multi-model con guration obtained a MAP value of 0.164, the textual con
guration scored best with 0.1546 MAP, and the best visual run achieves a score of
0.0997 MAP. The task was solved by 16 completely automated approaches and
14 runs which include manual intervention. It was observed that most manually
processed runs work best, independent from the con guration (textual, visual,
or multi-modal). They achieve MAP values in the range of 0.164 and 0.0295,
whereas the automated solutions range between scores of 0.0849 and 0.0013
MAP. A closer examination showed that all approaches had great di culties to
identify relevant images for the topic female old person. Also the topics 5, 13,
21, and 26 were hard to identify as well, but here some approaches were able
to reach higher MAP values. Especially, the topic riders on horse shows very
strong outliers with high MAP values above 0.5. The obtained MAP scores of the
remaining topics vary widely which points to a large variation in the di culty
level of topics. Some con gurations are able to achieve MAP values higher than
0.5 for individual topics. Considering the topic female old person, all runs show
nearly the same low performance. This seems to be an extremely critical topic
for concept-based image retrieval.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>We would like to thank the CLEF campaign for supporting the ImageCLEF
initiative. This work was partly supported by grant 01MQ07017 of the German
research program THESEUS funded by the Ministry of Economics.
21. Budikova, P., Batko, M., Zezula, P.: MUFIN at ImageCLEF 2011: Success or</p>
      <p>Failure? In: Working Notes of CLEF 2011, Amsterdam, The Netherlands. (2011)
22. Le, D.D., Satoh, S.: NII, Japan at ImageCLEF 2011 Photo Annotation Task. In:</p>
      <p>Working Notes of CLEF 2011, Amsterdam, The Netherlands. (2011)
23. Amel, K., Benammar, A., Amar, C.B.: REGIMvid at ImageCLEF2011:
Integrating Contextual Information to Enhance Photo Annotation and Concept-based
Retrieval. In: Working Notes of CLEF 2011, Amsterdam, The Netherlands. (2011)
24. Binder, A., Samek, W., Kawanabe, M.: The joint submission of the TU Berlin and
Fraunhofer FIRST (TUBFI) to the ImageCLEF2011 Photo Annotation Task. In:
Working Notes of CLEF 2011, Amsterdam, The Netherlands. (2011)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF 2009 Large-Scale Visual Concept Detection and Annotation Task</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Tsikrika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            , Muller, H., KalpathyCramer, J.,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Caputo</surname>
          </string-name>
          , B., eds.
          <source>: Multilingual Information Access Evaluation Vol. II Multimedia Experiments: Proceedings of the 10th Workshop of the Cross-Language Evaluation Forum (CLEF</source>
          <year>2009</year>
          ),
          <source>Revised Selected Papers. Lecture Notes in Computer Science</source>
          , Corfu, Greece (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huiskes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>New strategies for image annotation: Overview of the photo annotation task at imageclef 2010</article-title>
          .
          <source>Working notes of CLEF</source>
          <year>2010</year>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mark</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Huiskes</surname>
            ,
            <given-names>B.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lew</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          :
          <article-title>New trends and ideas in visual concept detection: The mir ickr retrieval evaluation initiative</article-title>
          .
          <source>In: MIR '10: Proceedings of the 2010 ACM International Conference on Multimedia Information Retrieval</source>
          , New York, NY, USA, ACM (
          <year>2010</year>
          )
          <volume>527</volume>
          {
          <fpage>536</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>A Consumer Photo Tagging Ontology: Concepts and Annotations</article-title>
          . In: THESEUS/ImageCLEF Pre-Workshop
          <year>2009</year>
          ,
          <article-title>Co-located with the Cross-Language Evaluation Forum (CLEF) Workshop and</article-title>
          13th European Conference on Digital Libraries ECDL, Corfu, Greece,
          <year>2009</year>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Semantic context transfer across heterogeneous sources for domain adaptive video search</article-title>
          .
          <source>In: Proceedings of the seventeen ACM international conference on Multimedia, ACM</source>
          (
          <year>2009</year>
          )
          <volume>155</volume>
          {
          <fpage>164</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A circumplex model of a ect</article-title>
          .
          <source>Journal of personality and social psychology 39(6)</source>
          (
          <year>1980</year>
          )
          <fpage>1161</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kludas</surname>
          </string-name>
          , J.:
          <article-title>Overview of the wikipediamm task at imageclef 2009</article-title>
          . In: Working notes of CLEF. (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kludas</surname>
          </string-name>
          , J.:
          <article-title>Overview of the Wikipedia Image Retrieval Task at ImageCLEF 2011</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kludas</surname>
          </string-name>
          , J.:
          <article-title>Overview of the Wikipedia Retrieval Task at ImageCLEF 2010</article-title>
          . In: Working notes of CLEF. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Daroczy</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pethes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benczur</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          : SZTAKI @
          <article-title>ImageCLEF 2011</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurie</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic Contexts and Fisher Vectors for the ImageCLEF 2011 Photo Annotation Task</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Znaidia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le Borgne</surname>
          </string-name>
          , H.:
          <article-title>CEA LISTs participation to Visual Concept Detection Task of ImageCLEF 2011</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mbanya</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hentschel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ndjiki-Nya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Sample Selection,
          <article-title>Category Speci c Features and Reasoning</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Nagel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Kuhhirt, U.,
          <string-name>
            <surname>Wolter</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The Fraunhofer IDMT at ImageCLEF 2011 Photo Annotation Task</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Van De Sande</surname>
            ,
            <given-names>K.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snoek</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          :
          <article-title>The University of Amsterdam's Concept Detection System at ImageCLEF 2011</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Rasche</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vertan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Testing a Method for Statistical Image Classi cation in Image Retrieval</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dellandrea</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bres</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>LIRIS-Imagine at ImageCLEF 2011 Photo Annotation task</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Izawa</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motohashi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takagi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Annotation and Retrieval System Using Confabulation Model for ImageCLEF2011 Photo Annotation</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Spyromitros-Xiou s</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Sechidis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsoumakas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vlahavas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>MLKD's Participation at the CLEF 2011 Photo Annotation and Concept-Based Retrieval Tasks</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Albatal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Safadi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quenot</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>LIG-MRIM at Image Photo Annotation task in ImageCLEF 2011</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2011</year>
          , Amsterdam, The Netherlands.
          <article-title>(</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>