<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the ImageCLEF 2014 Scalable Concept Image Annotation Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mauricio Villegas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Paredes</string-name>
          <email>rparedesg@prhlt.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>PRHLT, Universitat Politecnica de Valencia Cam de Vera</institution>
          <addr-line>s/n, 46022 Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>308</fpage>
      <lpage>328</lpage>
      <abstract>
        <p>The ImageCLEF 2014 Scalable Concept Image Annotation task was the third edition of a challenge aimed at developing more scalable image annotation systems. Unlike traditional image annotation challenges, which rely on a set of manually annotated images as training data, the participants were only allowed to use data and/or resources that as new concepts to detect are introduced do not require signi cant human e ort (such as hand labeling). The participants were provided with web data consisting of 500,000 images, which included textual features obtained from the web pages on which the images appeared, as well as various visual features extracted from the images themselves. To optimize their systems, the participants were provided with a development set of 1,940 samples and its corresponding hand labeled ground truth for 107 concepts. The performance of the submissions was measured using a test set of 7,291 samples which was hand labeled for 207 concepts among which 100 were new concepts unseen during development. In total 11 teams participated in the task submitting overall 58 system runs. Thanks to the larger amount of unseen concepts in the results the generalization of the systems has been more clearly observed and thus demonstrating the potential for scalability.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Automatic concept detection within images is a challenging and as of yet
unsolved research problem. Over the past decades impressive improvements have
been achieved, albeit admittedly not yet successfully solving the problem. Yet,
these improvements have been typically obtained on datasets for which all
images have been manually, and thus reliably, labeled. For instance, it has become
common in past image annotation benchmark campaigns [
        <xref ref-type="bibr" rid="ref15 ref3 ref9">9,15,3</xref>
        ] to use
crowdsourcing approaches, such as the Amazon Mechanical Turk1, in order to let
multiple annotators label a large collection of images. Nevertheless, crowdsourcing
is expensive and di cult to scale to a very large amount of concepts. The image
annotation datasets furthermore usually include exactly the same concepts in
the training and test sets, which may mean that the evaluated visual concept
detection algorithms are not necessarily able to cope with detecting additional
concepts beyond what they were trained on. To address these shortcomings a
novel image annotation task [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] was proposed in 2012 for which automatically
gathered web data had to be used for concept detection, where the concepts
varied between the evaluation sets. The aim of that task was to reduce the reliance
of cleanly annotated data for concept detection and rather focus on uncovering
structure from noisy data, emphasizing the importance of the need for scalable
annotation algorithms able to determine for any given concept whether or not it
is present in an image. The rationale behind the scalable image annotation task
was that there are billions of images available online appearing on webpages,
where the text surrounding the image may be directly or indirectly related to
its content, thus providing clues as to what is actually depicted in the image.
Moreover, images and the webpages on which they appear can be easily obtained
for virtually any topic using a web crawler. In existing work such noisy data has
indeed proven useful, e.g. [
        <xref ref-type="bibr" rid="ref16 ref21 ref22">16,22,21</xref>
        ].
      </p>
      <p>
        This paper presents the overview of the third edition of the Scalable Concept
Image Annotation task [
        <xref ref-type="bibr" rid="ref19 ref20">19,20</xref>
        ], one of the four benchmark campaigns organized
by ImageCLEF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] in 2014 under the CLEF initiative2. Section 2 describes the
task in detail, including the participation rules and the provided data and
resources. Followed by this, Section 3 presents and discusses the results of the
submissions. Finally, Section 4 concludes the paper with nal remarks and
future outlooks.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Overview of the Task</title>
      <sec id="sec-2-1">
        <title>Motivation and Objectives</title>
        <p>Image concept detection research generally has relied on training data that has
been manually, and thus reliably annotated, an expensive and laborious endeavor
that cannot easily scale as the number of concepts is increased. As an alternative
to clean labeled data, a very large amount of images can be easily gathered from
the web, and furthermore, from the webpages that contain the images, text
associated with them can be obtained. However, the degree of relationship between
the surrounding text and the image varies greatly. Moreover, the webpages can
be of any language or even a mixture of languages, and they tend to have many
writing mistakes. Overall the data can be considered to be very noisy. Motivated
by this need for scalability and the possibility of cheaply obtaining useful data,
the ImageCLEF 2014 Scalable Concept Image Annotation task concentrated
exclusively on developing annotation systems that rely only on automatically
obtained data.</p>
        <p>To illustrate the objective of the evaluation, consider for example that
someone searches for the word \rainbow" in a popular image search engine. It would
be expected that many results be of landscapes in which in the sky a rainbow
is visible. However, other types of images will also appear, see Figure 1a. The
images will be related to the query in di erent senses, and there might even be
2 http://www.clef-initiative.eu
(a) Images from a search query of \rainbow".</p>
        <p>(b) Images from a search query of \sun".
images that do not have any apparent relationship. In the example of Figure 1a,
one image is a text page of a poem about a rainbow, and another is a
photograph of an old cave painting of a rainbow serpent. See Figure 1b for a similar
example on the query \sun". As can be observed, the data is noisy, although
it does have the advantage that this data can also handle the possible di erent
senses that a word can have, or the di erent types of images that exist, such as
natural photographs, paintings and computer-generated imagery.</p>
        <p>In order to handle the web data, there are several resources that could be
employed in the development of scalable annotation systems. Many resources
can be used to help match general text to given concepts, amongst which some
examples are stemmers, word disambiguators, de nition dictionaries, ontologies
and encyclopedia articles. There are also tools that can help to deal with noisy
text commonly found on webpages, such as language models, stop word lists
and spell checkers. And last but not least, language detectors and statistical
machine translation systems are able to process webpage data written in various
languages.</p>
        <p>In summary, the goal of the scalable image annotation task was to evaluate
di erent strategies to deal with noisy data, so that the unsupervised web data
can be reliably used for annotating images for practically any topic.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Challenge Description</title>
        <p>The challenge3 consisted of the development of an image annotation system
given training data that only included images crawled from the Internet, the
corresponding webpages on which they appeared, as well as precomputed visual
3 Challenge website at http://imageclef.org/2014/annotation
and textual features. As mentioned in the previous section, the aim of the task
was for the annotation systems to be able to easily change or scale the list of
concepts used for image annotation. Apart from the image and webpage data,
the participants were also permitted and encouraged to use similar datasets and
any other automatically obtainable resources to help in the processing and usage
of the training data. However, the most important rule was that the systems were
not permitted to use any kind of data that had been explicitly and manually
labeled for concept detection learning.</p>
        <p>For the development of the annotation systems, the participants were
provided with the following:
{ A training dataset of images and corresponding webpages compiled speci cally
for the task, including precomputed visual and textual features (see Section
2.3).
{ Source code of a simple baseline annotation system (see Section 2.4).
{ Tools for computing the appropriate performance measures (see Section 2.5).
{ A development set of images with ground truth annotations (including
precomputed visual features) for estimating the system performance.</p>
        <p>After a period of three and a half months to work on the development set,
a test set of images was released which did not include any ground truth labels.
The participants had to use their developed systems to predict the concepts for
each of the input images and submit these results to the task organizers. About
one month was given to work on the test data and a maximum of 10 submissions
(also referred to as runs) were allowed per participating group. Since one of the
objectives was that the annotation systems be able to scale or change the list
of concepts for annotation, the list of concepts for the test set was not exactly
the same as those for the development set. Moreover, each test image had its
own list of concepts to detect, so not all images had to be annotated for all the
possible concepts. The development set consisted of 1,940 samples labeled for
107 concepts, and the test set consisted of 7,291 samples labeled for 207 concepts
(the same 107 concepts from development and 100 additional ones).</p>
        <p>
          The concepts to be used for annotation were de ned as one or more WordNet
synsets [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. So, for each concept there was a concept name, the type (either noun
or adjective), and the sense number(s). De ning the concepts this way, made
it straightforward to obtain the concept de nition, synonyms, hyponyms, etc.
Additionally, for most of the concepts, a link to a Wikipedia article about the
respective concept was provided. The complete list of concepts, as well as the
number of samples in the test sets, is included in Appendix A.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Dataset</title>
        <p>
          The dataset4 used was very similar to the one of the rst two editions of the
task [
          <xref ref-type="bibr" rid="ref19 ref20">19,20</xref>
          ]. To create the dataset, initially a database of over 31 million images
was created by querying Google, Bing and Yahoo! using words from the Aspell
4 Dataset available at http://risenet.prhlt.upv.es/webupv-datasets
English dictionary [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. The images and corresponding webpages were
downloaded, taking care to avoid data duplication. Then, a subset of 500,000 images
(to be used as the training set) was selected from this database by choosing the
top images from a ranked list. Half of this data was exactly the same as the
training set from last year, the additional data was merely intended to supply
images for the new concepts that were introduced. The motivation for selecting
a subset was to provide smaller data les that would not be so prohibitive for
the participants to download and handle. The ranked list was generated by
retrieving images from our database using the list of concepts, in essence more or
less as if the search engines had only been queried for these. From the ranked
list, some types of problematic images were removed, and it was guaranteed that
each image had at least one webpage in which they appeared. Unlike the training
set, the development and test sets were manually selected and labeled for the
concepts being evaluated. For further details on how the dataset was created,
please refer to [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>Textual Data: Since the textual data was to be used only during training, it
was only provided for the training set. Four sets of data were made available to
the participants. The rst one was the list of words used to nd the image when
querying the search engines, along with the rank position of the image in the
respective query and search engine it was found on. The second set of textual
data contained the image URLs as referenced in the webpages they appeared
in. In many cases the image URLs tend to be formed with words that relate
to the content of the image, which is why they can also be useful as textual
features. The third set of data were the webpages in which the images appeared,
for which the only preprocessing was a conversion to valid XML just to make
any subsequent processing simpler. The nal set of data were features obtained
from the text extracted near the position(s) of the image in each webpage it
appeared in.</p>
        <p>To extract the text near the image, after conversion to valid XML, the script
and style elements were removed. The extracted text were the webpage title and
all the terms closer than 600 in word distance to the image, not including the
HTML tags and attributes. Then a weight s(tn) was assigned to each of the
words near the image, de ned as
s(tn) = P</p>
        <p>1
8t2T s(t)</p>
        <p>X
8tn;m2T</p>
        <p>
          Fn;m sigm(dn;m) ;
(1)
where tn;m are each of the appearances of the term tn in the document T , Fn;m
is a factor depending on the DOM (e.g. title, alt, etc.) similar to what is done
in the work of La Cascia et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and dn;m is the word distance from tn;m to
the image. The sigmoid function was centered at 35, had a slope of 0.15 and
minimum and maximum values of 1 and 10 respectively. The resulting features
include for each image at most the 100 word-score pairs with the highest scores.
Visual Features: Before visual feature extraction, images were ltered and
resized so that the width and height had at most 240 pixels while preserving
the original aspect ratio. These raw resized images were provided to the
participants but also seven types of precomputed visual features. The rst feature
set Colorhist consisted of 576-dimensional color histograms extracted using our
own implementation. These features correspond to dividing the image in 3 3
regions and for each region obtaining a color histogram quanti ed to 6 bits. The
second feature set GETLF contained 256-dimensional histogram based features.
First, local color-histograms were extracted in a dense grid every 21 pixels for
windows of size 41 41. Then, these local color-histograms were randomly
projected to a binary space using 8 random vectors and considering the sign of the
resulting projection to produce the bit. Thus, obtaining a 8-bit representation of
each local color-histogram that can be considered as a word. Finally, the image
is represented as a bag-of-words, leading to a 256-dimensional histogram
representation. The third set of features consisted of GIST [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] descriptors. The
other four feature types were obtained using the colorDescriptors software [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ],
namely SIFT, C-SIFT, RGB-SIFT and OPPONENT-SIFT. The con guration
was dense sampling with default parameters and a hard assignment 1,000
codebook using a spatial pyramid of 1 1 and 2 2 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Since the vectors of the spatial
pyramid were concatenated, this resulted in 5,000-dimensional feature vectors.
The codebooks were generated using 1.25 million randomly selected features and
the k-means algorithm.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Baseline Systems</title>
        <p>A toolkit was supplied to the participants as a performance reference for the
evaluation, as well as to serve as a starting point. This toolkit included software
that computed the evaluation measures (see Section 2.5) and the
implementations of two baselines. The rst baseline was a simple random, which is
important since any system that gets worse performance than random is useless. The
other baseline, referred to as Co-occurrence Baseline, was a basic technique that
gives better performance than random, although it was simple enough to give
the participants a wide margin for improvement. In the latter technique, when
given an input image, obtains its nearest k = 32 images from the training set.
Then, the textual features corresponding to these k nearest images are used to
derive a score for each of the concepts. This is done by using a concept-word
co-occurrence matrix estimated from all of the training set textual features. In
order to make the vocabulary size more manageable, the textual features are
rst processed keeping only English words. Finally, the amount of concepts
assigned to the image is variable, the concepts selected are the ones with a score
higher than the sum of the mean and the standard deviation for all the concept
scores of that image. Since there were seven visual features provided, each one
was considered separately for a baseline.
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Performance Measures</title>
        <p>Ultimately the goal of an image annotation system is to make decisions about
which concepts to assign to given image from a prede ned list of concepts. Thus
to measure annotation performance what should be considered is how good are
those decisions. On the other hand, in practice many annotations systems are
based on estimating a score for each of the concepts and then a second technique
uses these scores to nally decide which concepts are chosen. For systems of this
type a measure of performance can be based only on the concept scores, which
considers all aspects of the system except for the technique used for concept
decisions, making it an interesting characteristic to measure.</p>
        <p>For this task, two basic performance measures have been used for comparing
the results of the di erent submissions. The rst one is the F-measure (F1),
which takes into account the nal annotation decisions, and the other is the
Average Precision (AP), which considers the concept scores.</p>
        <p>The F1 is de ned as</p>
        <p>2P R</p>
        <p>F1 = P + R ; (2)
where P is the precision and R is the recall. In the context of image annotation,
the F1 can be estimated from two di erent perspectives, one being concept-based
and the other sample-based. In the former, one F1 is computed for each concept,
and in the latter one F1 is computed for each image to annotate. In both cases,
the arithmetic mean is used as a global measure of performance, and will be
referenced as MF1-concepts and MF1-samples, respectively.</p>
        <p>The AP is algebraically de ned as</p>
        <p>AP =
1 XjKj k
jKj k=1 rank(k)
;
(3)
where K is the ordered set of the ground truth annotations, being the order
induced by the annotation scores, and rank(k) is the order position of the k-th
ground truth annotation. The fraction k= rank(k) is actually the precision at the
k-th ground truth annotation, and has been written like this to be explicit on
the way it is computed. In the cases that there are ties in the scores, a random
permutation is applied within the ties. The AP can also be estimated for both
the concept-based and sample-based perspectives, however, the concept-based
AP is not a suitable measure of annotation performance (it is more adequate
for a retrieval scenario), so only the sample-based AP has been considered in
this evaluation. As a global measure of performance, also the arithmetic mean
is used, which will be referred to as MAP-samples.</p>
        <p>A bit of care must be taken when comparing systems using the MAP-samples
measure. What the MAP-samples turns out saying is that if for a given image the
scores are used to sort the concepts, how good would it rank the true concepts
for the image. Depending on the system, its scores could or could not be optimal
for ranking the concepts. Thus a system with a relatively low MAP-samples,
could still have a good annotation performance if the method used to select the
concepts is adequate for its concept scores. Because of this, as well as the fact
that there can be systems that do not rely on scores, it was optional for the
participants of the task to provide scores.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Results</title>
      <sec id="sec-3-1">
        <title>Participation</title>
        <p>
          The participation was excellent, although there was a slight decrease in
participation with respect to last year. In total, 11 groups took part in the task
and submitted overall 58 system runs. Among the 11 participating groups, only
7 of them submitted a corresponding paper describing their system, thus only
for these there were speci c details available. Last year the participation was
13 groups, 58 runs and 9 papers. The following 11 teams were the ones that
participated:
{ DISA: [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] The team from the Laboratory of Data Intensive Systems and
Applications of the Masaryk University (Brno, Czech Republic) was represented
by Petra Budikova, Jan Botorek, Michal Batko and Pavel Zezula.
{ IPL: [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] The team from the Information Processing Laboraroty of the Athens
University of Economics and Business (Athens, Greece) was represented by
Spyridon Stathopoulos and Theodore Kalamboukis.
{ KDEVIR: [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] The team from the Computer Science and Engineering
department of the Toyohashi University of Technology (Aichi, Japan), was
represented by Ismat Ara Reshma, Md Zia Ullah and Masaki Aono.
{ MIL: [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] The team from the Machine Intelligence Lab of the University of
Tokyo (Tokyo, Japan) was represented by Atsushi Kanehira, Masatoshi
Hidaka, Yusuke Mukuta, Yuichiro Tsuchiya, Tetsuaki Mano and Tatsuya Harada.
{ MindLab: [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] The team from the Machine learning, perception and
discovery Lab from the Universidad Nacional de Colombia (Bogota, Colombia) was
represented by Jorge A. Vanegas, John Arevalo, Sebastian Otalora, Fabian
Paez, Santiago A. Perez-Rubiano and Fabio A. Gonzalez.
{ MLIA: [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] The team from the Department of Advanced Information
Technology of the Kyushu University (Fukuoka, Japan) was represented by Xing
Xu, Atsushi Shimada and Rin-ichiro Taniguchi.
{ RUC: [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] The team from the School of Information of the Renmin University
of China (Beijing, China) was represented by Xirong Li, Xixi He, Gang Yang,
Qin Jin and Jieping Xu.
{ FINKI:5 The team from the Faculty of Computer Science and Engineering
of the Ss. Cyril and Methodius University (Skopje, Republic of Macedonia)
was represented by Ivica Dimitrovski.
{ IMC:5 The team from the Institute of Media Computing of the Fudan
University (Shanghai, China) was represented by Yong Cheng.
5 No paper describing their system submitted.
{ INAOE:5 The team from the Instituto Nacional de Astrof sica, Optica y
Electronica (Puebla, Mexico) was represented by Hugo Jair Escalante and
Luis Pellegrin.
{ NII:5 The team from the National Institute of Informatics (Tokyo, Japan)
was represented by Duy Dinh Le.
Since the objective of this task was to compare annotation systems that are
scalable, a very important aspect to evaluate is precisely their scalability. However,
unlike the annotation performance, it is di cult to quantify the scalability of
a system so that the submissions can be compared in this respect. Therefore,
instead of attempting to give a measure for scalability, in this section we make
a few comments about possible aspects of the proposed systems in which the
scalability could be compromised.
        </p>
        <p>One characteristic observed in this year's task was that three teams based
their system on Convolutional Neural Networks (CNN) pre-trained using
ImageNet, a dataset which was manually hand labeled for 1,000 WordNet synsets.
Two of the teams, MIL and MindLab, used the CNN output of an intermediate
layer as a visual feature. There are works in which it has been observed that CNN
features perform well in new problems di erent from the one that was trained
for, so in some sense their use does not violate the competition rule of no hand
labeled data usage. However, a minor detail is that the ImageNet synsets overlap
considerably with the current task's concepts, so the annotation performance for
these systems might be a bit optimistic in comparison to the others. The third
team that used CNN was MLIA, which employed the synsets predicted by the
CNN to clean the concepts automatically assigned using the webpage data. In
this case the performance of the system could be greatly a ected if the concepts
for annotation di er signi cantly from the ones of ImageNet.</p>
        <p>Also this year most of the teams proposed approaches based on classi ers
that need to be learned. In the case of the MIL team, the classi er is multilabel.
A multilabel classi er could be problematic since each time the list of concepts to
detect changes, the classi er would have to be relearned. However, the PAAPL
algorithm of MIL is designed with special consideration of scalability, so in their
case it does not seem an issue. The alternative of multilabel is having one
classi er per concept, which are learned one concept at a time using positive and
negative samples. For scalability, the learning should be based on a selection
of negative images so that this process is independent of how many concepts
there are. It seems that all of the teams consider this adequately. However, with
respect to a multilabel classi er this might not be the optimal approach. When
new concepts are introduced it could be advisable to learn new classi ers to
consider the relationships between the concepts. However, this relationship could
be taken into account in a step after classi cation which is what the KDEVIR
team has done with their constructed ontologies.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Annotation Performance Results</title>
        <p>The test set for this year was composed of 4 subsets of samples, each of which
had a di erent list of concepts for annotation. The rst subset contained 3,000
images which were exactly the same as last year's development and test sets, and
the list of concepts for annotation were also the same 116. The second subset had
1,747 images and the list of concepts were 52 related to the topic animals. The
third subset had 479 images and the list of concepts were 41 related to the topic
foods. For both the animals and foods subsets all of the concepts for annotation
were not among the ones seen in development. The nal subset had 2065 images
and the list of concepts for annotation were all the 207.</p>
        <p>Table 2 presents the performance measures (mentioned in 2.5) for the
baseline techniques and all of the submitted runs by the participants. The table
includes the results for the complete test set (referenced as all ) and for three of
the mentioned subsets: animals, foods and the one annotated for all the 207
concepts, referenced as ani., food and 207, respectively. Also for the MF1-concepts
measure the unseen column presents the results for the complete test set, but
only considering the 100 concepts that did not appear in the development set.
The systems are ordered by performance, beginning at the top with the best
performing one. This order was derived by considering for the test set the
average rank when comparing all of the systems, using the complete test set for the
three performance measures and also MF1-concepts unseen. Ties were broken by
the average of the same measures. Considering only the performance measures,
this ranking indicates that the best system this year was the one developed by
KDEVIR.</p>
        <p>For an easier comparison and a more intuitive visualization, the results for the
complete test set and all the submissions are presented as graphs in Figure 2. In
the graphs the error bars correspond to the 95% con dence intervals estimated by
Wilson's method, employing the standard deviation for the individual measures
(for the samples or concepts, and for the average precisions (AP) or F-measures
(F1), depending on the case). For the MF1-concepts measure two results for
each submission is presented, one that includes all concepts and another that
considers only the unseen concepts. Similarly Figure 3 presents the results for
all the submissions, but in each case depicting the performance for three of the
subsets of the test set: animals, foods and the 207 concepts subset. The fourth
subset of the test is intended for comparison of the systems with respect to the
previous edition of the task, so this is presented in a separate graph, Figure 4,
although for space reasons and make the comparison more illustrative, only the
best submission of each group is included.</p>
        <p>Finally, in Figure 5 there is for each of the 207 test set concepts, a boxplot
for the F1 when combining all runs. In order to t all of the concepts in the
all concepts
unseen concepts
same graph, for multiple outliers with the same value, only one is shown. The
concepts have been sorted by the median performance of all submissions, which
in a way orders them by di culty.
As can be observed in Figure 4, the performance of the systems has improved
somewhat with respect to what was obtained in the previous edition of the
task. This year 5 teams obtained all of the performance measures over 30% in
contrast to just 3 from last year. An interesting detail is that it seams that
the improvements for the MF1 measures are greater than for the MAP-samples.
Thus it can be observed that this year better approaches have been developed
for making the nal concept annotations decisions.
foods
animals
207</p>
        <p>Observing the results for the complete test set in Figure 2, the best
MAPsamples and MF1-samples values are somewhat lower than for last year's test
set. This could be related to the fact that this year the list of concepts was
larger, thus making the problem a bit more di cult. With respect to the
MF1concepts measure, last year the results were characterized by having relatively
large con dence intervals, which made drawing conclusions a bit di cult
specially for the unseen concepts. The increase in the number of unseen concepts
has made the results clearer. For the three measures the performance for two of
the systems submitted by KDEVIR signi cantly outperforms all of the others.
Most impressive is the advantage obtained for the MF1-concepts measure and
even more if only the unseen concepts are considered, obtaining a performance
over 65%. Note that the good performance for the unseen concepts is due to the
fact that many of the new concepts were used in the animals and foods subsets
MAP-samples
MF1-samples
MF1-concepts
2013
TPT</p>
        <p>MUINLIM20O1R3E20RU1U3NCED20&amp;1CU3EVAU2R0L1JI3SCT&amp;U20N1E3D2M01IC3KCD2E01V3IR2013
2014</p>
        <p>IR
KDEV</p>
        <p>
          IL
M
which had smaller concept lists, making the problem a bit easier. Analyzing the
key details of the systems presented in Table 1, it can be noted that the success
of the KDEVIR system is most probably due to the usage of concept ontologies
both in the training phase for better selecting the images used for optimizing the
classi ers and in the testing phase for taking into account the relationships
between the concepts. Moreover, the KDEVIR system also employed the technique
of last year's winner TPT [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], which uses a learning technique that takes into
account context, e ectively nding a way to exploit the information available in
the noisy webpage data.
        </p>
        <p>Even though the foods and animals subsets consisted purely of unseen
concepts, in the results in Figure 3 it can be seen that the performance for these
is in general much better, mostly because of the known relationship between
the size of the list of concepts for annotation and the performance. The smaller
the list of concepts, the easier the problem becomes. Moreover, similar to last
year in Figure 5 it can be observed that the unseen concepts do not tend to
perform worse. The di culty of each particular concept a ects more the
performance than the fact that these have not been seen during development, or
from another perspective the systems are able to generalize rather well to new
concepts.</p>
        <p>Considering both the annotation performance measures and the scalability
analysis, it can be declared that this year's winner is the KDEVIR system. The
fact that KDEVIR only used the provided visual features shows the
characteristic of this evaluation, which in contrast to usual image annotation tasks with
labeled training data, this challenge requires work in more fronts in order to get
important improvements.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>This paper presented an overview of the ImageCLEF 2014 Scalable Concept
Image Annotation task, the third edition of a challenge aimed at developing
more scalable image annotation systems. The goal was to develop annotation
systems that for training, only rely on unsupervised web data and other cheaply
obtainable resources, thus making it easy to add or change the concepts for
annotation.</p>
      <p>
        The participation was similar to last year although with a slight decrease, 11
teams submitted in total 58 system runs. The performance of the submitted
systems was somewhat superior to last year's results, in particular improving more
for the MF1 measures, which indicate a greater success in the developed
techniques for choosing the nal annotated concepts. Thanks to the larger amount
of concepts in the test set that were not seen during development, the results
for the MF1-concepts measure had narrower con dence intervals, so it made the
comparison the systems more conclusive. Moreover, by having subsets in the test
set which had to be annotated using only unseen concepts, it has been observed
that the systems are able to generalize well. The clear winner of this year's
evaluation was the KDEVIR [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] team, which after analyzing the key components
of the system it can be observed that most of the success is due to the usage of a
classi er learning technique that takes into account context, e ectively nding a
way to exploit the information available in the noisy webpage data; and the
usage of automatically generated concept ontologies both in the training phase for
better selecting the images used for optimizing the classi ers and in the testing
phase for taking into account the relationships between the concepts.
      </p>
      <p>The results of the task have been very interesting and show that useful
annotation systems can be built using noisy web crawled data. Since the problem
requires to cover many fronts, there is still a lot of work that can be done, so
it would be interesting to continue this line of research. Papers on this topic
should be published, demonstration systems based on these ideas be built and
more evaluation of this sort be organized. Also it remains to see how this can
be used to complement systems that are based on clean hand labeled data and
nd ways to take advantage of both the supervised and unsupervised data.
{|ri||e||||blood
{|su|bm|ar||ine|||felidae
| | |yam
{cl|oseup|||captive
| |
{|sp|id|er |||shadow
{{cmlooundulme||sesn||t|spectacles
|||
||| |poster
{overcas|t||cervidae
|||
{trunk
||
{outd||oo||r||||chaantidae
||
{|coast
{hu|nt||ing||||||tmuableer</p>
      <p>|
|
{|fo|x||
{embroi|der||y|unpaved
||| |branch
{|pa|rk| ection
{equi|dae||||||freemale
||
{{seol|diler||||daytime
||
|
{{||rreopdte|||inlet |||||||||tsaupnrotterllteope
|
|
{|ta|ble|
{|sculptu||re ||||cahrtilhdropod</p>
      <p>||
{{mmuond||key||||||chaigthway
||
||
{ph|on|e |||bird
|
{|dr|in|k|||butter y
{brid|ge |||baby
||
{bus
||
{gard||en||||||dwoanrtkheoyg
||
{dog
{pota||to||||||tsemeonkaeger
||
||
{wago|n |||penguin
||
{|bo|tt|le |||lettuce
{cucu|mb|er||bear
||
{|ri|ce |||painting
{diag||ram|||leaf
||
{|rh|in|o|||knife
{|ho|rs|e|||squirrel
{|roas|ted|||onion</p>
      <p>|
{|rain||||cityscape</p>
      <p>|
{|ch|air|
{fo|otw|ea||r ||||rcoacrk
|
{|zo|o ||||rabbit
{chee|se |||violin
||
{{|forriaend||ge||||||cmaarrrsoutpial
||</p>
      <p>|
{|m|am|m|al ||cloud
{|croc|odile ||spoon</p>
      <p>| |
{|ca|ul|io|we|r| sh
{eleph|an|t||tricycle
||
{|po|ol||||river
{|st|raw|b|err|y|harbor
{|ap|pl|e|||tra c
{|lio|n ||||guitar
{port|rai|t||gira e
||
{{twr|ialidn||||indoor
||
| ||||toy
{|o|we|r | |wolf
{|coun|trysid||e |banana</p>
      <p>| |
{protest|||tomato
|||
{|ra|inb|ow|||grape
{|ai|rpl|an|e||kangaroo
{in|str|um|en|t|berry
|
{gala|xy |||human
||
{bo|ok||||pinniped
|
{|pig
{truck||||silhouette</p>
      <p>|
|||
{furnitu||re ||||fsoigrnest
|||
{walr|us |||camel
||
{|road||||helicopter</p>
      <p>|
{|be|ac|h|||motorcycle
{mushroom
||||||grass
{|dr|um||||sausage
{|as|pa|rag|us||logo
{{ltai||gkeer||||bicycle
|
{avoc||ad||o||||bbruoccoli
| alo
||
{|co|rn||||raccoon
{build|in|g||pasta
||
{|br|ea|d|||egg
{|ca|stl|e|||cheetah
{pineapp|le||eggplant
|||
{ostrich|||watermelon
|||
{||
{{cphaeurmsr||ocinhng|||o||||||lsmeaoonpudanrtdain
|||
||
{|sk|y ||||vehicle
{goril|la |||koala
||
{dese|rt |||sea
||
{moo|n|||plant
||
{pan ||||newspaper
||
{|su|n ||||aerial
{boat
||
{mea|t|||underwater
|||
{pum|pk||in ||||nhiepbpuolapotamus
||
{|r|e
{tree||||snow
|||| |pear
{|chim|pa|nz||ee |space</p>
      <p>|
{|fo|g||||water
{|soup||||nighttime</p>
      <p>|
{|ze|br|a|||sunrise/sunset
{|ca|rto|on|||food
{|r|ew|or|k||vegetable
{|lig|ht|nin|g ||fruit
h s
t t
fro cep</p>
      <p>n
ts o
o C
l
p .e
x c
o n
B a
: m
5 r
. fo
g r
i e
F p</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The authors are very grateful with the CLEF initiative for supporting ImageCLEF.
The research leading to these results has received funding from the European Union's
Seventh Framework Programme (FP7/2007-2013) under the tranScriptorium project
(#600707) and from the Spanish MEC under the STraDA project
(TIN2012-37475C02-01).</p>
    </sec>
    <sec id="sec-6">
      <title>Concept List 2014</title>
      <p>The following tables present the 207 concepts used in the ImageCLEF 2014 Scalable
Concept Image Annotation task. In the electronic version of this document, each
concept name and Wikipedia article name are hyperlinks to webpages of the corresponding
WordNet synset and the Wikipedia article, respectively.</p>
      <p>Concepts seen in both development and test sets:</p>
      <p>Concept WordNet 3</p>
      <p>type sense#
antelope noun 1
apple noun 1
arthropod noun 1
asparagus noun 2
avocado noun 1
banana noun 2
bear noun 1
berry noun 1, 2
blood noun 1
branch noun 2
bread noun 1
broccoli noun 1
buffalo noun 1
butterfly noun 1
camel noun 1
canidae noun 1
captive noun 2
carrot noun 1
cauliflower noun 1
cervidae noun 1
cheese noun 1
cheetah noun 1
chimpanzee noun 1</p>
      <p>corn noun 3, 1
crocodile noun 1
cucumber noun 2
donkey noun 2</p>
      <p>egg noun 2
eggplant noun 1
elephant noun 1
equidae noun 1
felidae noun 1
flamingo noun 1
fox noun 1
fried adj. 1
fruit noun 1
galaxy noun 3
giraffe noun 1
gorilla noun 1
grape noun 1
hippopotamus noun 1
human noun 1
hunting noun 1
kangaroo noun 1
knife noun 1
koala noun 1
leaf noun 1
leopard noun 2
lettuce noun 3
lion noun 1</p>
      <p>#test</p>
      <p>Concept WordNet 3</p>
      <p>type sense#
mammal noun 1
marsupial noun 1
meat noun 1
monkey noun 1
mud noun 1
mushroom noun 5, 1
nebula noun 3
onion noun 1, 3
orange noun 1
ostrich noun 2
pan noun 1
pasta noun 2
pear noun 1
penguin noun 1</p>
      <p>pig noun 1
pineapple noun 2
pinniped noun 1
pool noun 1
potato noun 1
pumpkin noun 2
rabbit noun 1
raccoon noun 2
reptile noun 1
rhino noun 1
rice noun 1
rifle noun 1
roasted adj. 1
rock noun 1, 2
rodent noun 1
sausage noun 1
soup noun 1
spider noun 1
spoon noun 1
squirrel noun 1
strawberry noun 1
submarine noun 1
tiger noun 2
tomato noun 1
trunk noun 1
tuber noun 1
turtle noun 2
vegetable noun 1
walrus noun 1
warthog noun 1
watermelon noun 2
wild noun 2
wolf noun 1
yam noun 1, 4
zebra noun 1
zoo noun 1</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Budikova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Botorek</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zezula</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : DISA at ImageCLEF 2014:
          <article-title>The search-based solution for scalable image annotation</article-title>
          .
          <source>In: CLEF 2014 Evaluation Labs and Workshop</source>
          , Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , Muller, H.,
          <string-name>
            <surname>Martinez-Gomez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Acar</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patricia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marvasti</surname>
            , N., Uskudarl ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cazorla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Varea</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morell</surname>
          </string-name>
          , V.:
          <article-title>ImageCLEF 2014: Overview and analysis of the results</article-title>
          .
          <source>In: CLEF proceedings. Lecture Notes in Computer Science</source>
          , Springer Berlin Heidelberg (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Imagenet: A largescale hierarchical image database</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          . IEEE Conference on. pp.
          <volume>248</volume>
          {
          <issue>255</issue>
          (june
          <year>2009</year>
          ), doi:10.1109/CVPR.
          <year>2009</year>
          .5206848
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.):
          <article-title>WordNet An Electronic Lexical Database</article-title>
          . The MIT Press, Cambridge, MA; London (May
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kanehira</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hidaka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukuta</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsuchiya</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mano</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harada</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>MIL at ImageCLEF 2014: Scalable System for Image Annotation</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>La</given-names>
            <surname>Cascia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sethi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sclaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>Combining textual and visual cues for contentbased image retrieval on the World Wide Web</article-title>
          .
          <source>In: Content-Based Access of Image and Video Libraries</source>
          ,
          <year>1998</year>
          . Proceedings. IEEE Workshop on. pp.
          <volume>24</volume>
          {
          <issue>28</issue>
          (
          <year>1998</year>
          ), doi:10.1109/IVL.
          <year>1998</year>
          .694480
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lazebnik</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponce</surname>
          </string-name>
          , J.:
          <article-title>Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories</article-title>
          .
          <source>In: Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision</source>
          and Pattern Recognition - Volume
          <volume>2</volume>
          . pp.
          <volume>2169</volume>
          {
          <fpage>2178</fpage>
          . CVPR '06, IEEE Computer Society, Washington, DC, USA (
          <year>2006</year>
          ), doi:10.1109/CVPR.
          <year>2006</year>
          .68
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          : Renmin University of China at
          <article-title>ImageCLEF 2014 Scalable Concept Image Annotation</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liebetrau</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The CLEF 2011 Photo Annotation and Concept-based Retrieval Tasks</article-title>
          . In: Petras,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Forner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Clough</surname>
          </string-name>
          , P.D. (eds.)
          <article-title>CLEF 2011 Labs</article-title>
          and Workshop, Notebook Papers,
          <fpage>19</fpage>
          -22
          <source>September</source>
          <year>2011</year>
          , Amsterdam, The Netherlands (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Oliva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope</article-title>
          .
          <source>Int. J. Comput. Vision</source>
          <volume>42</volume>
          (
          <issue>3</issue>
          ),
          <volume>145</volume>
          {175 (May
          <year>2001</year>
          ), doi:10.1023/A:1011139631724
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Reshma</surname>
            ,
            <given-names>I.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ullah</surname>
            ,
            <given-names>M.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aono</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>KDEVIR at ImageCLEF 2014 Scalable Concept Image Annotation Task: Ontology based Automatic Image Annotation</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Sahbi</surname>
          </string-name>
          , H.:
          <string-name>
            <surname>CNRS - TELECOM ParisTech at ImageCLEF 2013 Scalable Concept</surname>
          </string-name>
          <article-title>Image Annotation Task: Winning Annotations with Context Dependent SVMs</article-title>
          . In:
          <article-title>CLEF 2013 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. Valencia,
          <source>Spain (September</source>
          <volume>23</volume>
          -26
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. van de Sande,
          <string-name>
            <given-names>K.E.</given-names>
            ,
            <surname>Gevers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Snoek</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.G.</surname>
          </string-name>
          :
          <article-title>Evaluating Color Descriptors for Object and Scene Recognition</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>32</volume>
          ,
          <fpage>1582</fpage>
          {
          <fpage>1596</fpage>
          (
          <year>2010</year>
          ), doi:10.1109/TPAMI.
          <year>2009</year>
          .154
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Stathopoulos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalamboukis</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          : IPL at ImageCLEF 2014:
          <article-title>Scalable Concept Image Annotation</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2012 Flickr Photo Annotation and Retrieval Task</article-title>
          . In:
          <article-title>CLEF 2012 working notes</article-title>
          . Rome, Italy (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fergus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freeman</surname>
            , W.: 80
            <given-names>Million</given-names>
          </string-name>
          <string-name>
            <surname>Tiny</surname>
          </string-name>
          <article-title>Images: A Large Data Set for Nonparametric Object and Scene Recognition</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on
          <volume>30</volume>
          (
          <issue>11</issue>
          ),
          <year>1958</year>
          {1970 (nov
          <year>2008</year>
          ), doi:10.1109/TPAMI.
          <year>2008</year>
          .128
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Vanegas</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arevalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Otalora</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez-Rubiano</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>F.A.</given-names>
          </string-name>
          : MindLab at ImageCLEF 2014:
          <article-title>Scalable Concept Image Annotation</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Image-Text Dataset Generation for Image Annotation and Retrieval</article-title>
          . In: Berlanga,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          , P. (eds.) II Congreso Espan~ol de Recuperacion de Informacion,
          <string-name>
            <surname>CERI</surname>
          </string-name>
          <year>2012</year>
          . pp.
          <volume>115</volume>
          {
          <fpage>120</fpage>
          . Universidad Politecnica de Valencia, Valencia, Spain (June 18-19
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
          </string-name>
          , R.:
          <article-title>Overview of the ImageCLEF 2012 Scalable Web Image Annotation Task</article-title>
          . In: Forner,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Womser-Hacker</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <article-title>CLEF 2012 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. Rome,
          <source>Italy (September</source>
          <volume>17</volume>
          -20
          <year>2012</year>
          ), http://mvillegas.info/pub/Villegas12_CLEF_
          <article-title>Annotation-Overview</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2013 Scalable Concept Image Annotation Subtask</article-title>
          . In:
          <article-title>CLEF 2013 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. Valencia,
          <source>Spain (September</source>
          <volume>23</volume>
          -26
          <year>2013</year>
          ), http: //mvillegas.info/pub/Villegas13_CLEF_
          <article-title>Annotation-Overview</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Ma</surname>
          </string-name>
          , W.Y.:
          <article-title>ARISTA - image search to annotation on billions of web photos</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <source>2010 IEEE Conference on</source>
          . pp.
          <volume>2987</volume>
          {
          <issue>2994</issue>
          (
          <year>June 2010</year>
          ), doi:10.1109/CVPR.
          <year>2010</year>
          .5540046
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Usunier</surname>
          </string-name>
          , N.:
          <article-title>Large scale image annotation: learning to rank with joint word-image embeddings</article-title>
          .
          <source>Machine Learning</source>
          <volume>81</volume>
          ,
          <volume>21</volume>
          {
          <fpage>35</fpage>
          (
          <year>2010</year>
          ), doi:10.1007/s10994-010-5198-3
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shimada</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>ichiro</surname>
            <given-names>Taniguchi</given-names>
          </string-name>
          , R.: MLIA at
          <article-title>ImageCLFE 2014 Scalable Concept Image Annotation Challenge</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop, Online Working Notes. She eld,
          <source>UK (September</source>
          <volume>15</volume>
          -18
          <year>2014</year>
          )
          <article-title>Concepts seen only in the test set:</article-title>
          <source>Wikipedia article Antelope 28 Apple 70 Arthropod 78 Asparagus 24 Avocado 24 Banana 46 Bear 34 Berry 38 Blood 22 Branch 994 Bread 70 Broccoli 32 African buffalo 62 Butterfly 16 Camel 46 Canidae 188 Captivity (animal) 332 Carrot 52 Cauliflower 28 Cervidae 114 Cheese 76 Cheetah 32 Chimpanzee 38 Maize 50 Crocodile 32 Cucumber 36 Donkey 20 Egg (food) 54 Eggplant 34 Elephant 40 Equidae 156 Felidae 208 Flamingo 22 Fox 26 Frying 58 Fruit 390 Galaxy 21 Giraffe 46 Gorilla 32 Grape 78 Hippopotamus 60 Human 998 Hunting 40 Kangaroo 24 Knife 30 Koala 30 Leaf 2012 Leopard 44 Lettuce 56 Lion 40</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>