<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>New Strategies for Image Annotation: Overview of the Photo Annotation Task at ImageCLEF 2010</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefanie Nowak</string-name>
          <email>stefanie.nowak@idmt.fraunhofer.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Huiskes</string-name>
          <email>mark.huiskes@liacs.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fraunhofer IDMT</institution>
          ,
          <addr-line>Ilmenau</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leiden Institute of Advanced Computer Science, Leiden University</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ImageCLEF 2010 Photo Annotation Task poses the challenge of automated annotation of 93 visual concepts in Flickr photos. The participants were provided with a training set of 8,000 Flickr images including annotations, EXIF data and Flickr user tags. Testing was performed on 10,000 Flickr images, differentiated between approaches considering solely visual information, approaches relying on textual information and multi-modal approaches. Half of the ground truth was acquired with a crowdsourcing approach. The evaluation followed two evaluation paradigms: per concept and per example. In total, 17 research teams participated in the multi-label classification challenge with 63 submissions. Summarizing the results, the task could be solved with a MAP of 0.455 in the multi-modal configuration, with a MAP of 0.407 in the visual-only configuration and with a MAP of 0.234 in the textual configuration. For the evaluation per example, 0.66 F-ex and 0.66 OS-FCS could be achieved for the multi-modal configuration, 0.68 F-ex and 0.65 OS-FCS for the visual configuration and 0.26 F-ex and 0.37 OS-FCS for the textual configuration.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The steadily increasing amount of multimedia data poses challenging questions
on how to index, visualize, organize, navigate or structure multimedia
information. Many different approaches are proposed in the research community, but
often their benefit is not clear as they were evaluated on different datasets with
different evaluation measures. Evaluation campaigns aim to establish an
objective comparison between the performance of different approaches by posing
well-defined tasks including datasets, topics and measures. This paper presents
an overview of the ImageCLEF 2010 Photo Annotation Task. The task aims at
the automated detection of visual concepts in consumer photos. Section 2
introduces the task and describes the database, the annotation process, the ontology
and the evaluation measures applied. Section 3 summarizes the approaches of
the participants to solve the task. Next, the results for all configurations are
presented and discussed in Section 4 and Section 5, respectively. Finally, Section 6
summarizes and concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>Task Description</title>
      <p>The ImageCLEF Visual Concept Detection and Annotation Task poses a
multilabel classification challenge. It aims at the automatic annotation of a large
number of consumer photos with multiple annotations. The task can be solved
by following three different approaches:
1. Automatic annotation with content-based visual information of the images.
2. Automatic annotation with Flickr user tags and EXIF metadata in a purely
text-based scenario.
3. Multi-modal approaches that consider both visual and textual information
like Flickr user tags or EXIF information.</p>
      <p>
        In all cases the participants of the task were asked to annotate the photos of
the test set with a predefined set of keywords (the concepts), allowing for an
automated evaluation and comparison of the different approaches. Concepts are
for example abstract categories such as Family&amp;Friends or Partylife, the Time of
Day (Day, Night, sunny, ), Persons (no person, single person, small group or big
group), Quality (blurred, underexposed ) and Aesthetics; 52 from the 53 concepts
that were used in the ImageCLEF 2009 benchmark are used again [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In total the
number of concepts was extended to 93 concepts. In contrast to the annotations
from 2009, the new annotations were obtained with a crowdsourcing approach
that utilizes Amazon Mechanical Turk. The task uses a subset of the MIR Flickr
25,000 image dataset [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for the annotation challenge. The MIR Flickr collection
supplies all original tag data provided by the Flickr users (noted as Flickr user
tags). In the collection there are 1386 tags which occur in at least 20 images,
with an average total number of 8.94 tags per image. These Flickr user tags
are made available for the textual and multi-modal approaches. For most of the
photos the EXIF data is included and may be used.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Evaluation Objectives</title>
        <p>This year the focus of the task lies on the comparison of the strengths and
limitations of the different approaches:
– Do multi-modal approaches outperform text only or visual only approaches?
– Which approaches are best for which kind of concepts?
– Can image classifiers scale to the large number of concepts and data?</p>
        <p>Furthermore, the task challenges the participants to deal with an unbalanced
number of annotations per photo, an unbalanced number of photos per concept,
the subjectivity of concepts like boring, cute or fancy and the diversity of photos
belonging to the same concept. Further, the textual runs have to cope with a
small number of images without EXIF data and/or Flickr user tags.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Annotation Process</title>
        <p>
          The complete dataset consists of 18,000 images annotated with 93 visual
concepts. The manual annotations for 52 concepts were acquired by Fraunhofer
IDMT in 2009. (The concept Canvas from 2009 was discarded.) Details on the
manual annotation process and concepts, including statistics on concept
frequencies can be found in [
          <xref ref-type="bibr" rid="ref1 ref3">3, 1</xref>
          ]. In 2010, 41 new concepts were annotated with a
crowdsourcing approach using the Amazon Mechanical Turk. In the following,
we just focus on the annotation process of these new concepts.
        </p>
        <p>
          Amazon Mechanical Turk (MTurk, www.mturk.com) is an online marketplace
in which mini-jobs can be distributed to a crowd of people. At MTurk these
minijobs are called HITs (Human Intelligence Tasks). They represent a small piece
of work with an allocated price and completion time. The workers at MTurk,
called turkers, can choose the HITs they would like to perform and submit the
results to MTurk. The requester of the work collects all results from MTurk after
they are completed. The workflow of a requester can be described as follows: 1)
design a HIT template, 2) distribute the work and fetch results and 3) approve
or reject work from turkers. For the design of the HITs, MTurk offers support by
providing a web interface, command line tools and developer APIs. The requester
can define how many assignments per HIT are needed, how much time is allotted
to each HIT and how much to pay per HIT. MTurk offers several ways of assuring
quality. Optionally, the turkers can be asked to pass a qualification test before
working on HITs, multiple workers can be assigned the same HIT and requesters
can reject work in case the HITs were not finished correctly. The HIT approval
rate each turker achieves by completing HITs can be used as a threshold for
authorization to work. Before the annotations of the ImageCLEF 2010 tasks
were acquired, we performed a pre-study to investigate if annotations from
nonexperts are reliable enough to be used in an evaluation benchmark. The results
were very promising and encouraged us to adapt this service for the 2010 task.
Details of the pre-study can be found in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Design of HIT templates: In total, we generated four different HIT
templates at MTurk. For all concepts, the annotations per photo were obtained
three times. Later the final annotations are built from the majority vote of these
three opinions. For the annotation of the 41 new concepts we made use of the
pre-knowledge that we have from the old annotations. Therefore the 41 concepts
were structured into four groups:</p>
        <sec id="sec-2-2-1">
          <title>1. Vehicles</title>
          <p>The ImageCLEF 2009 dataset contains a number of photos annotated with
the concept Vehicle. These photos were further annotated with the concepts
car, bicycle, ship, train, airplane and skateboard. A textbox offered the
possibility to input further categories. The turkers could select a checkbox saying
that no vehicle is depicted in the photo to cope with the case of false
annotations. The corresponding survey with guidelines can be found in Figure 1.
Each HIT was rewarded with 0.01$.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2. Animals</title>
          <p>The ImageCLEF 2009 photo collection already contains several photos that
were annotated with the concept animals. The turkers at Amazon were asked
to further classify these photos in the categories dog, cat, bird, horse, fish and
insect. Again, a textbox offered additional input possibilities. For each HIT
a reward of 0.01$ was paid.
3. Persons</p>
          <p>The dataset contains photos that were annotated with a person concept
(single person, small group or big group of persons). These photos were further
classified with human attributes like female, male, Baby, Child, Teenager,
Adult and old person. Each HIT was rewarded with 0.01$.
2.3</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Ontology</title>
        <p>
          The concepts were organised in an ontology. For this purpose the Consumer
Photo Tagging Ontology [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] of 2009 was extended with the new concepts. The
hierarchy allows making assumptions about the assignment of concepts to
documents. For instance, if a photo is classified to contain trees, it also contains
plants. Then, next to the is-a relationship of the hierarchical organization of
concepts, also other relationships between concepts can determine label
assignments. The ontology requires for example that for a certain sub-node only one
concept can be assigned at a time (disjoint items) or that a special concept (e.g.
portrait ) postulates other concepts like persons or animals. The ontology allows
the participants to incorporate knowledge in their classification algorithms, and
to make assumptions about which concepts are probable in combination with
certain labels. Further, it is used in the evaluation of the submissions.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Evaluation Measures</title>
        <p>
          The evaluation follows the concept-based and example-based evaluation paradigms.
For the concept-based evaluation the Average Precision (AP) is utilized. This
measure showed better characteristics than the Equal Error Rate (EER) and
Area under Curve (AUC) in a recent study [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. For the example-based
evaluation we apply the example-based F-Measure (F-ex). The Ontology Score of last
year was extended with a different cost map that is based on Flickr metadata
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and serves as additional evaluation measure. It is called Ontology Score with
Flickr Context Similarity (OS-FCS) in the following.
2.5
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Submission</title>
        <p>The participants submitted their results for all photos in a single text file that
contains the photo ID as first entry per row followed by 93 floating point values
between 0 and 1 (one value per concept). The floating point values are regarded
as confidence while computing the AP. After the confidence values for all photos,
the text file contains binary values for each photo (so again each line contains
the photo ID followed by 93 binary values). The measures F-ex and OS-FCS
need a binary decision about the presence or absence of the concepts. Instead of
applying a strict threshold at 0.5 of the confidence values, the participants have
the possibility to threshold each concept for each image individually. All groups
had to submit a short description of their runs and state which configuration
they chose (annotation with visual information only, annotation with textual
information only or annotation with multi-modal information). In the following
the visual configuration is abbreviated with ”V”, the textual with ”T” and the
multi-modal one with ”M”.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Participation</title>
      <p>In total 54 groups registered for the visual concept detection and annotation task,
41 groups signed the license agreement and were provided with the training and
test sets, 17 of them submitted results in altogether 63 runs. The number of runs
was restricted to a maximum of 5 runs per group. There were 45 runs submitted
in the visual only configuration, 2 in the textual only configuration and 16 in
the multi-modal configuration.</p>
      <p>
        BPACAD|SZTAKI [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]: The team of the Computer and Automation
Research Institute of the Hungarian Academy of Science submitted one run in the
visual configuration. Their approach is based on Histogram of Oriented
Gradients descriptors which were clustered with a 128 dimensional Gaussian Mixture
Model. Classification was performed with a linear logistic regression model with
a χ2 kernel per category.
      </p>
      <p>CEA-LIST: The team from CEA-LIST, France submitted one run in the
visual configuration. They extract various global (colour, texture) and local
(SURF) features. The visual concepts are learned with a fast shared boosting
approach and normalized with a logistic function.</p>
      <p>
        CNRS|Telecom ParisTech [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]: The CNRS group of Telecom ParisTech,
Paris, France participated with five multi-modal runs. Their approach is based
on SIFT features represented by multi-level spatial pyramid bag-of-words. For
classification a one-vs-all trained SVM is utilized.
      </p>
      <p>
        DCU [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]: The team of Dublin City University, Ireland submitted one run in
the textual configuration. They followed a document expansion approach based
on the Okapi feedback method to expand the image metadata and concepts
and applied DBpedia as external information source in this step. To deal with
images without any metadata, the relationships between concepts in the training
set is investigated. The date and time information of the EXIF metadata was
extracted to predict concepts like Day.
      </p>
      <p>
        HHI [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: The team of Fraunhofer HHI, Berlin, Germany submitted five runs
in the visual-only configuration. Their approach is based on the bag of words
approach and introduces category specific features and classifiers including
quality related features. They use opponent SIFT features with dense sampling and
a sharpness feature and base their classification on a multi-kernel SVM
classifier with χ2 distance. Second, they incorporate a post-processing approach that
considers relations and exclusions between concepts. Both extensions resulted in
an increase in performance compared to the standard bag-of-words approach.
      </p>
      <p>
        IJS [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]: The team of Joˇzef Stefan Institute, Slovenia and Department of
Computer Science, Macedonia submitted four runs in the visual configuration.
They use various global and local image features (GIST, colour histograms,
SIFT) and learn predictive clustering trees classifiers. For each descriptor a
separate classifier is learned and the probabilities output of all classifiers is combined
for the final prediction. Further, they investigate ensembles of predictive
clustering tree classifiers. The combination of global and local features leads to better
results than using local features alone.
      </p>
      <p>
        INSUNHIT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: The group of the Harbin Institute of Technology, China
participated with five runs in the visual configuration. They use dense SIFT
features as image descriptors and classify with a na¨ıve-bayes nearest neighbour
approach. The classifier is extended with a random sampling image to class
distance to cope with imbalanced classes.
      </p>
      <p>
        ISIS [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]: The Intelligent Systems Lab of the University of Amsterdam, The
Netherlands submitted five runs in the visual configuration. They use a dense
sampling strategy that combines a spatial pyramid approach and saliency points
detection, extract different SIFT features, perform a codebook transformation
and classify with a SVM approach. The focus lies on the improvement of the
scores in the evaluation per image. They use the distance to the decision plane
in the SVM as probability and determine the threshold for binary annotation
from this distance.
      </p>
      <p>
        LEAR and XRCE [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]: The team of LEAR and XEROX, France made a
joint contribution with a total of ten runs, five submitted in the visual and five
in the multi-modal configuration. They use SIFT and colour features on several
spatial scales and represent them as improved Fisher vectors in a codebook of
256 words. The textual information is represented as a binary presence/absence
vector of the most common 698 Flickr user tags. For classification a linear SVM is
compared to a k -NN classifier with learned neighbourhood weights. Both
classification models are computed with the same visual and textual features and late
and early fusion approaches are investigated. All runs considering multi-modal
information outperformed the runs in the visual configuration.
      </p>
      <p>
        LIG [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]: The team of Grenoble University, France submitted one run in the
visual configuration of the Photo Annotation task. They extract colour SIFT
features and cluster them with a k -means clustering procedure in 4000 clusters. For
classification a SVM with RBF kernel is learned in an one-against-all approach
and based on the 4000 dimensional histogram of word occurrences.
      </p>
      <p>LSIS [16]: The Laboratory of Information Science and Systems, France
submitted two runs in the visual configuration. They propose features based on
extended local binary patterns extracted with spatial pyramids. For classification
they use a linear max-margin SVM classifier.</p>
      <p>MEIJI [17]: The group of Meiji University, Kanagawa, Japan submitted in
total five runs. They followed a conceptual fuzzy set approach applied to visual
words, a visual words baseline with SIFT descriptors and a combination with
a Flickr User Tag system using TF-IDF. Classification is based on a matching
of visual word combinations between the training casebase and the test image.
For the visual word approach the cosine distance is applied for similarity
determination. In total, two runs were submitted in the visual configuration and
three in the multi-modal one. Their multi-modal runs outperform the visual
configurations.</p>
      <p>MLKD: The team of the Aristotle University of Thessaloniki, Greece
participated with three runs; one in each configuration. For the visual and the textual
runs ensemble classifier chains are used as classifiers. The visual configuration
applies C-SIFT features with a Harris-Laplace salient point detector and clusters
them in a 4000 word codebook. As textual features, the 250 most frequent Flickr
user tags of the collection are represented in a binary feature vector per image.
The multi-modal configuration chooses the confidence score of the model (textual
or visual) per concept for which a better AP was determined in the evaluation
phase. As a result, the multi-modal approach outperforms the visual and the
textual models.</p>
      <p>Romania [18]: The team of the University Bucharest, Romania participated
with five runs in the visual configuration. Their approach considers the extraction
of colour histograms and combine them with a method of structural description.
The classification is performed using a Linear Discriminant Analysis (LDA) and
a weighted average retrieval rank (ARR) method. The annotations resulting from
the LDA classifier were refined considering the joint probabilities of concepts.
As a result the ARR classification outperforms the LDA classification.</p>
      <p>UPMC/LIP6 [19]: The team of University Pierre et Marie Curie, Paris,
France participated in the visual and the multi-modal configuration. They
submitted a total of five runs (3V, 2M). Their approach investigates the fusion of
results from different classifiers with supervised and semi-supervised
classification methods. The first model is based on fusing outputs from several
RankingSVM classifiers that classified the images based on visual features (SIFT,
HSV, Mixed+PCA). The second model further incorporates unlabeled data from
the test set for which the initial classifiers are confident to assign a certain
label and retrains the classifiers based on the augmented set. Both models were
tested with the additional inclusion of Flickr user tags using the Porter stemming
algorithm. For both models the inclusion of user tags improved the results.</p>
      <p>WROCLAW [20]: The group of Wroclaw University, Poland submitted
five runs in the visual configuration. They focus on global colour and texture
features and adapt an approach which annotates photos through the search for
similar images and the propagation of their tags. In their configurations several
similarity measures (Minkowski distance, Cosine distance, Manhattan distance,
Correlation distance and Jensen-Shannon divergence) are investigated. Further,
an approach based on a Penalized Discriminant Analysis classifier was applied.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>This section presents the results of the Photo Annotation Task 2010. First, the
overall results of all teams independent of the configuration are presented. In
the following subsections the results per configuration are highlighted.</p>
      <p>In Table 1 the results for the evaluation per concept independent of the
applied configuration are illustrated for the best run of each group. The results
for all runs can be found at the Photo Annotation Task website1. The task
could be solved best with a MAP of 0.455 (XRCE) followed by a MAP of 0.437
(LEAR). Both runs make use of multi-modal information. Table 2 illustrates
the overall ranking for the results of the evaluation per example. The table is
sorted descending for the F-ex measure. The best results were achieved in a
visual configuration with 0.68 F-ex (ISIS) and in a multi-modal configuration
with 0.66 OS-FCS (XRCE).
The results for the two textual runs are presented in Table 4. Both groups
achieve close results in the concept-based evaluation. However, the
examplebased evaluation measures show a significant difference between the results of
both teams.
The following section discusses some of the results in more detail. The best results
for each concept are summarized in Table 6. On average the concepts could be
detected with a MAP of 0.48 considering the best results per concept from all
configurations and submissions. From 93 concepts, 61 could be annotated best
with a multi-modal approach, 30 with a visual approach and two with a textual
one. Most of the concepts were classified best by one configuration of the XRCE,
ISIS or LEAR group.
1 http://www.imageclef.org/2010/PhotoAnnotation</p>
      <p>MAP RANK F-ex RANK OS-FCS</p>
      <p>ISIS
XRCE
LEAR</p>
      <p>HHI</p>
      <p>IJS
BPACAD</p>
      <p>Romania
INSUNHIT</p>
      <p>LSIS</p>
      <p>LIG</p>
      <p>MEIJI
WROCLAW</p>
      <p>MLKD</p>
      <p>UPMC
CEALIST</p>
      <p>The best classified concepts are the ones from the mutually exclusive
categories: Neutral-Illumination (98.2% AP, 94% F), No-Visual-Season (96,5% AP,
88% F), No-Persons (91.9% AP, 68% F), No-Blur (91.5% AP, 68% F).
Following, the concepts Outdoor (90.9% AP, 50% F), Sky, (89.5% AP, 27% F) Day
(88.1% AP, 51% F) and Clouds (85.9% AP, 14% F) were annotated with a high
AP. The concepts with the worst annotation quality were abstract (4.6% AP, 1%
F), old-person (11.6% AP, 2% F), work (13.1% AP, 3% F), technical (14.2% AP,
4% F), Graffiti (14.5% AP, 1% F), and boring (16.2% AP, 6% F). The
percentages in parentheses denote the detection performance in AP and the frequency
(F) of the concept occurrence in the images of the test set. Although there is a
trend that concepts that occur more frequently in the image collection can be
detected better, this does not hold for all concepts. Figure 2 shows the frequency
of concepts in the test collection plotted against the best AP achieved by any
submission.</p>
      <p>Although the performance of the textual runs is much lower in average than
in the visual and textual runs, there are two concepts that can be annotated best
in a textual configuration: skateboard and abstract. The concept skateboard was
just annotated in six images of the test set and twelve of the training set. In the
user tags of three images the word “skateboard” was present, while two images
have no user tags and the sixth image does not contain words like “skateboard”
or “skateboarding”. It seems as if there is not enough visual information available
to learn this concept while the textual and multi-modal approaches can make
use of the tags and extract the correct concept from the tags for at least half
of the images. The concept abstract was annotated more often (1,2% in the test
set and 4,7% in the training set).</p>
      <p>
        Further, one can see a great difference in annotation quality between the
old concepts from 2009 that were carefully annotated by experts (number 1-52)
and the new concepts (number 53-93) annotated with the service of Mechanical
Turk. The average annotation quality in terms of MAP for the old concepts is
0.57 while it is 0.37 for the new concepts. The reason for this is unclear. One
reason may lie in the quality of the annotations of the non-experts. However,
recent studies found that the quality of crowdsourced annotations is similar to
the annotation quality of experts [
        <xref ref-type="bibr" rid="ref4">21, 4, 22</xref>
        ]. Another reason could be the choice
and difficulty of the new concepts, as some of them are not as obvious and
objective as the old ones. Further, some of the new concepts are special and
their occurrence in the dataset is lower ( 7% in average) than the occurrence of
the old concepts ( 17% in average).
      </p>
      <p>One possibility to determine the reliability of a test collection is to
calculate Cronbach’s alpha value [23]. It defines a holistic measure of reliability and
analyses the variance of individual test items and total test scores. The measure
returns a value ranging between zero and one, for which bigger scores indicate
a higher reliability. The Cronbach’s alpha values show a high reliability for the
whole test collection with 0.991, 0.991 for the queries assessed by experts and
0.956 for the queries assessed by MTurk. Therefore the scores point to a reliable
test collection for both the manual expert annotations and the crowdsourced
annotations and cannot explain the differences in MAP by the annotating systems.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The ImageCLEF 2010 Photo Annotation Task posed a multi-label annotation
challenge for visual concept detection in three general configurations (textual,
visual and multi-modal). The task attracted a considerable number of
international teams with a final participation of 17 teams that submitted a total of
63 runs. In summary, the challenge could be solved with a MAP of 0.455 in
the multi-modal configuration, with a MAP of 0.407 in the visual only
configuration and with a MAP of 0.234 in the text configuration. For the evaluation
per example 0.66 F-ex and 0.66 OS-FCS could be achieved for the multi-modal
configuration, 0.68 F-ex and 0.65 OS-FCS for the visual configuration and 0.26
F-ex and 0.37 OS-FCS for the textual configuration. All in all, the multi-modal
approaches got the best scores for 61 out of 93 concepts, followed by 30 concepts
that could be detected best with the visual approach and two that won with a
textual approach. As just two runs were submitted in the textual configuration,
it is not possible to determine the abilities of purely textual classifiers reliably.
In general, the multi-modal approaches outperformed visual and textual
configurations for all teams that submitted results for more than one configuration.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>We would like to thank the CLEF campaign for supporting the ImageCLEF
initiative. This work was partly supported by grant 01MQ07017 of the German
research program THESEUS funded by the Ministry of Economics.
16. Paris, S., Glotin, H.: Linear SVM for LSIS Pyramidal Multi-Level Visual only
Concept Detection in CLEF 2010 Challenge. In: Working Notes of CLEF 2010,
Padova, Italy. (2010)
17. Motohashi, N., Izawa, R., Takagi, T.: Meiji University at the ImageCLEF2010
Visual Concept Detection and Annotation Task: Working notes. In: Working Notes
of CLEF 2010, Padova, Italy. (2010)
18. Rasche, C., Vertan, C.: A Novel Structural-Description Approach for Image
Retrieval. In: Working Notes of CLEF 2010, Padova, Italy. (2010)
19. Fakeri-Tabrizi, A., Tollari, S., Usunier, N., Amini, M.R., Gallinari, P.: UPMC/LIP6
at ImageCLEFannotation 2010. In: Working Notes of CLEF 2010, Padova, Italy.
(2010)
20. Stanek, M., Maier, O.: The Wroclaw University of Technology Participation at
ImageCLEF 2010 Photo Annotation Track. In: Working Notes of CLEF 2010,
Padova, Italy. (2010)
21. Alonso, O., Mizzaro, S.: Can we get rid of TREC assessors? Using Mechanical
Turk for relevance assessment. In: SIGIR 2009 Workshop on the Future of IR
Evaluation. (2009)
22. Hsueh, P., Melville, P., Sindhwani, V.: Data Quality from Crowdsourcing: a Study
of Annotation Selection Criteria. In: Proceedings of the NAACL HLT 2009
Workshop on Active Learning for Natural Language Processing. (2009)
23. Bodoff, D.: Test theory for evaluating reliability of IR test collections. Information
Processing &amp; Management 44(3) (2008) 1117–1145</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF 2009 Large-Scale Visual Concept Detection and Annotation Task</article-title>
          . In Peters,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Tsikrika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            , Mu¨ller, H., KalpathyCramer, J.,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Caputo</surname>
          </string-name>
          , B., eds.
          <source>: Multilingual Information Access Evaluation Vol. II Multimedia Experiments: Proceedings of the 10th Workshop of the Cross-Language Evaluation Forum (CLEF</source>
          <year>2009</year>
          ),
          <source>Revised Selected Papers. Lecture Notes in Computer Science</source>
          , Corfu, Greece (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Huiskes</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lew</surname>
            ,
            <given-names>M.S.:</given-names>
          </string-name>
          <article-title>The MIR Flickr Retrieval Evaluation</article-title>
          .
          <source>In: Proc. of the ACM International Conference on Multimedia Information Retrieval</source>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>A Consumer Photo Tagging Ontology: Concepts and Annotations</article-title>
          . In: THESEUS/ImageCLEF Pre-Workshop
          <year>2009</year>
          ,
          <article-title>Co-located with the Cross-Language Evaluation Forum (CLEF) Workshop and</article-title>
          13th European Conference on Digital Libraries ECDL, Corfu, Greece,
          <year>2009</year>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Ru¨ger, S.:
          <article-title>How reliable are Annotations via Crowdsourcing: a Study about Inter-annotator Agreement for Multi-label Image Annotation</article-title>
          .
          <source>In: MIR '10: Proceedings of the International Conference on Multimedia Information Retrieval</source>
          , New York, NY, USA, ACM (
          <year>2010</year>
          )
          <fpage>557</fpage>
          -
          <lpage>566</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lukashevich</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dunker</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Ru¨ger, S.:
          <article-title>Performance Measures for Multilabel Evaluation: a Case Study in the Area of Image Classification</article-title>
          .
          <source>In: MIR '10: Proceedings of the International Conference on Multimedia Information Retrieval</source>
          , New York, NY, USA, ACM (
          <year>2010</year>
          )
          <fpage>35</fpage>
          -
          <lpage>44</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Llorente</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Ru¨ger, S.:
          <article-title>The Effect of Semantic Relatedness Measures on Multi-label Classification Evaluation</article-title>
          .
          <source>In: ACM International Conference on Image and Video Retrieval</source>
          ,
          <string-name>
            <surname>CIVR.</surname>
          </string-name>
          (
          <year>July 2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Dar´oczy,
          <string-name>
            <surname>B.</surname>
          </string-name>
          , Petr´as, I., Benczu´r,
          <string-name>
            <given-names>A.A.</given-names>
            ,
            <surname>Nemeskey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Pethes</surname>
          </string-name>
          , R.: SZTAKI @
          <article-title>ImageCLEF 2010</article-title>
          . In: Working Notes of CLEF 2010, Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Sahbi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <string-name>
            <surname>TELECOM ParisTech at ImageCLEF 2010 Photo Annotation</surname>
          </string-name>
          <article-title>Task: Combining Tags and Visual Features for Learning-Based Image Annotation</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Min</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>G.J.F.</given-names>
          </string-name>
          :
          <article-title>A Text-Based Approach to the ImageCLEF 2010 Photo Annotation Task</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Mbanya</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hentschel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <surname>M.</surname>
          </string-name>
          , Nu¨rnberger,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ndjiki-Nya</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Augmenting Bag-of-Words - Category Specific Features and Concept Reasoning</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Dimitrovski</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kocev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loskovska</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , D˜zeroski, S.:
          <article-title>Detection of Visual Concepts and Annotation of Images using Predictive Clustering Trees</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Random Sampling Image to Class Distance for Photo Annotation</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. van de Sande,
          <string-name>
            <given-names>K.E.A.</given-names>
            ,
            <surname>Gevers</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>The University of Amsterdam's Concept Detection System at ImageCLEF 2010</article-title>
          . In: Working Notes of CLEF 2010, Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mensink</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Csurka</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perronnin</surname>
          </string-name>
          , F., S´anchez, J.,
          <string-name>
            <surname>Verbeek</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>LEAR and XRCE's participation to Visual Concept Detection Task - ImageCLEF 2010</article-title>
          . In: Working Notes of CLEF 2010, Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Batal</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mulhem</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <string-name>
            <surname>MRIM-LIG at ImageCLEF 2010 Visual Concept</surname>
          </string-name>
          <article-title>Detection and Annotation task</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2010</year>
          , Padova,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>