<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the ImageCLEF 2016 Scalable Concept Image Annotation Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrew Gilbert</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Piras</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josiah Wang</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fei Yan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arnau Ramisa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emmanuel Dellandrea</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Gaizauskas</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mauricio Villegas</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krystian Mikolajczyk</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Since 2010, ImageCLEF has run a scalable image annotation task, to promote research into the annotation of images using noisy web page data. It aims to develop techniques to allow computers to describe images reliably, localise di erent concepts depicted and generate descriptions of the scenes. The primary goal of the challenge is to encourage creative ideas of using web page data to improve image annotation. Three subtasks and two pilot teaser tasks were available to participants; all tasks use a single mixed modality data source of 510,123 web page items for both training and test. The dataset included raw images, textual features obtained from the web pages on which the images appeared, as well as extracted visual features. Extracted from the Web by querying popular image search engines, the dataset was formed. For the main subtasks, the development and test sets were both taken from the \training set". For the teaser tasks, 200,000 web page items were reserved for testing, and a separate development set was provided. The 251 concepts were chosen to be visual objects that are localizable and that are useful for generating textual descriptions of the visual content of images and were mined from the texts of our extensive database of image-webpage pairs. This year seven groups participated in the task, submitting over 50 runs across all subtasks, and all participants also provided working notes papers. In general, the groups' performance is impressive across the tasks, and there are interesting insights into these very relevant challenges.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>How can you use large-scale noisy data to improve image classi cation,
caption generation and text illustration? This challenging question is the basis of
this year's image annotation challenge. Every day, users struggle with the
everincreasing quantity of data available to them. Trying to nd \that" photo they
took on holiday last year, the image on Google of their favourite actress or band,
or the images of the news article someone mentioned at work. There are a huge
number of images that can be cheaply found and gathered from the Internet.
However, more valuable is mixed-modality data, for example, web pages
containing both images and text. A signi cant amount of information about the image
is present on these web pages and vice-versa. However, the relationship between
the surrounding text and images varies greatly, with much of the text being
redundant and unrelated. Despite the obvious bene ts of using such information
in automatic learning, the weak supervision it provides means that it remains a
challenging problem. Fig. 1 illustrates the expected results of the task.</p>
      <p>The Scalable Concept Image Annotation task is a continuation of the general
image annotation and retrieval task that has been part of ImageCLEF since its
very rst edition in 2003. In the early years the focus was on retrieving relevant
images from a web collection given (multilingual) queries, from 2006 onwards
annotation tasks were also held, initially aimed at object detection, but more
recently also covering semantic concepts. In its current form, the 2016 Scalable
Concept Image Annotation task is its fth edition, having been organized in
2012 [24], 2013 [26], 2014 [25], and 2015 [8]. In the 2015 edition [8], the image
annotation task was expanded to concept localization and also natural language
sentential description of images. In this year's edition, we further introduced a
text illustration `teaser' task, to evaluate systems that analyse a text document
and select the best illustration for the text from a large collection of images
provided. As there is an increased interest in recent years in research combining
text and vision, the new tasks introduced in both the 2015 and 2016 editions
aim at further stimulating and encouraging multimodal research that uses both
text and visual data for image annotation and retrieval.</p>
      <p>This paper presents the overview of the fth edition of the Scalable
Concept Image Annotation task [24,26,25,8], one of the three benchmark campaigns
organized by ImageCLEF [22] in 2016 under the CLEF initiative1. Section 2
describes the task in detail, including the participation rules and the provided data
and resources. Section 3 presents and discusses the results of the submissions
received for the task. Finally, Section 4 concludes the paper with nal remarks
and future outlooks.</p>
    </sec>
    <sec id="sec-2">
      <title>Overview of the Task</title>
      <sec id="sec-2-1">
        <title>Motivation and Objectives</title>
        <p>Image annotation has relied on training data that has been manually, and thus
reliably annotated. Annotating training data is an expensive and laborious
endeavour that cannot be easily scaled, particularly as the number of concepts</p>
        <sec id="sec-2-1-1">
          <title>1 http://www.clef-initiative.eu</title>
          <p>(a) Images from a search query of \rainbow".</p>
          <p>(b) Images from a search query of \sun".
grows. However, images for any topic can be cheaply gathered from the Web,
along with associated text from the web pages that contain the images. The
degree of relationship between these web images and the surrounding text varies
considerably, i.e., the data are very noisy, but overall these data contain useful
information that can be exploited to develop annotation systems. Figure 2 shows
examples of typical images found by querying search engines. As can be seen, the
data obtained are useful and furthermore a wider variety of images is expected,
not only photographs but also drawings and computer generated graphics. This
diversity has the advantage that this data can also handle the di erent possible
senses that a word can have or the various types of images that exist. Likewise,
there are other resources available that can help to determine the relationships
between text and semantic concepts, such as dictionaries or ontologies. There
are also tools that can contribute to deal with noisy text commonly found on
web pages, such as language models, stop word lists and spell checkers.</p>
          <p>Motivated by the need for exploiting this useful (albeit noisy) data, the
ImageCLEF 2016 Scalable Concept Image Annotation task aims to develop
techniques to allow computers to describe images reliably, localise the di erent
concepts depicted in the images, generate a description of the scene and select
images to illustrate texts. The primary objective of the 2016 edition is to encourage
creative ideas of using noisy, web page data so that it can be used to improve
various image annotation tasks { concept annotation and localization, selecting
important concepts to be described, generating natural language descriptions,
and retrieving images to illustrate a text document.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Challenge Description</title>
        <p>This year the challenge2 consisted of 3 subtasks and a teaser task.3
1. Subtask 1 (Image Annotation and Localization): The image annotation task
remains the same as the 2015 edition. Participants are required to develop a
system that receives as input an image and produces as output a prediction
of which concepts are present in that image, selected from a prede ned list
of concepts. Like the 2015 edition, they should also output bounding boxes
indicating where the concepts are located within the image.
2. Subtask 2 (Natural Language Caption Generation): This subtask was geared
towards participants interested in developing systems that generate textual
descriptions directly with an image as input. For example, by using visual
detectors to identify concepts and generating textual descriptions from the
detected concepts, or by learning neural sequence models in a joint fashion
to create descriptions conditioned directly on the image. Participants used
their own image analysis methods, for example by using the output of their
image annotation systems developed for Subtask 1. They are also encouraged
to augment their training data with the noisy content of the web page.
3. Subtask 3 (Content Selection): This subtask was primarily designed for
those interested in the Natural Language Generation aspects of Subtask 2
while avoiding visual processing of images. It concentrated on the content
selection phase when generating image descriptions, i.e. which concepts (from
all possible concepts depicted) should be selected, and mentioned in the
corresponding description? Gold standard input, bounding boxes labelled with
concepts for each test image was provided, and participants were expected
to develop systems that predict the bounding box instances most likely to be
mentioned in the corresponding image descriptions. Unlike the 2015 edition,
participants were not required to generate complete sentences but were only
requested to provide a list of bounding box instances per image.
4. Teaser task (Text Illustration): This pilot task is designed to evaluate the
performance of methods for text-to-image matching. Participants were asked
to develop a system to analyse a given text document and nd the best
illustration for it from a set of all available images. At test time, participants
were provided as input a selection of text documents as queries, and the goal
was to select the best illustration for each text from a collection of 200,000
images.</p>
        <p>As a common dataset, participants were provided with 510,123 web images,
the corresponding web pages on which they appeared, as well as precomputed
visual and textual features (see Sect. 2.4). As in the 2015 task, external training
2 Challenge website at http://imageclef.org/2016/annotation
3 A Second teaser task was also introduced, aimed at evaluating systems that identify
the GPS coordinates of a text document's topic based on its text and image data.
However, we had no participants for this task, and thus will not discuss this second
teaser task in this paper.
data such as ImageNet ILSVRC2015 and MSCOCO is also allowed, and
participants were also encouraged to use other resources such as ontologies, word
disambiguators, language models, language detectors, spell checkers, and
automatic translation systems.</p>
        <p>We observed in the 2015 edition that this large-scale noisy web data was
not used as much as we anticipated { participants mainly used external training
data. To encourage participants to utilise the provided data for training, in this
edition participants were expected to produce two sets of related results:
1. using only external training data;
2. using both external data and the noisy web data of 510,123 web pages.
The aim is for participants to improve the performance of externally trained
systems, using the provided noisy web data. However none of the participants
submitted results, only on the supplied noisy training; this is probably due to
the fact the groups are chasing the optimal image annotation results, and not
actively attempting to research into using the noisy training data.
Development datasets: In addition to the training dataset and visual/textual
features mentioned above, the participants were provided with the following for
the development of their systems:
{ A development set of images (a small subset of the training data) with ground
truth labelled bounding box annotations and precomputed visual features for
estimating the system performance for Subtask 1.
{ A development set of images with at least ve textual descriptions per image
for Subtask 2.
{ A subset of the development set above for Subtask 3, with gold standard
inputs (bounding boxes labelled with concepts) and correspondence annotation
between bounding box inputs and terms in textual descriptions.
{ A development set for the Teaser task, with approximately 3,000 image-web
page pairs. This set is disjoint from the 510,123 noisy dataset.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Concepts</title>
        <p>For the three subtasks, the 251 concepts were retained from the 2015 edition.
They were chosen to be visual objects that are localizable and that are useful
for generating textual descriptions of the visual content of images. They include
animate objects such as people, dogs and cats, inanimate objects such as houses,
cars and balls, and scenes such as city, sea and mountains. With the concepts
mined from the texts of our database of 31 million image-webpage pairs [23].
Nouns that are subjects or objects of sentences are extracted and mapped onto
WordNet synsets [7]. In addition, ltered to `natural', basic-level categories (dog
rather than a Yorkshire terrier ), based on the WordNet hierarchy and heuristics
from a large-scale text corpora [28]. The organisers manually shortlisted the
nal list of concepts such that they were (i) visually concrete and localizable;
(ii) suitable for use in image descriptions; (iii) at an appropriate `every day'
level of speci city that was neither too general nor too speci c. The complete
list of concepts, as well as the number of samples in the test sets, is included in
Appendix A.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Dataset</title>
        <p>The dataset this year4 was built on the 500,000 image-webpage pairs from the
2015 edition. The 2015 dataset used was very similar to previous three editions
of the task [24,26,25]. To create the dataset, a database of over 31 million images
was created by querying Google, Bing and Yahoo! using words from the Aspell
English dictionary [23]. The images and corresponding web pages were
downloaded, taking care to avoid data duplication. Then, a subset of 500,000 images
was selected from this database by choosing the top images from a ranked list.
For further details on the dataset creation, please refer to [24]. By retrieving
images from our database using the list of concepts, the ranked list was generated,
in essence, more or less as if the search engines was queried. From the ranked
list, some types of problematic images were removed, and each image had at
least one web page in which they appeared.</p>
        <p>To incorporate the teaser task this year, the 500,000 image dataset from
2015 was augmented with 10,123 new image-webpage pairs, taken from a
subset of the BreakingNews dataset [16] which we developed, expanding the size of
the dataset to 510,123 image-webpage pairs. The aim of generating the
BreakingNews dataset was to further research into image and text annotation, where
the textual descriptions are loosely related to their corresponding images. Unlike
the main subtasks, the textual descriptions in this dataset do not describe real
image content but provide connotative and ambiguous relations that may not be
directly inferred from images. More speci cally, the documents and images were
obtained from various online news sources such as BBC News and The Guardian.
For the teaser task, a random subset of image-article pairs was selected from the
original dataset, and we ensured that each image corresponds to only one text
article. The reports were converted to a `web page' via a generic template.</p>
        <p>Like last year, for the main subtasks the development and test sets were both
taken from the \training set". Both sets were retained from last year, making the
evaluation for the three subtasks comparable across both 2015 and 2016 editions.
To generate these sets, a set of 5,520 images was selected using a CNN trained
to identify images suitable for sentence generation. Crowd-sourcing, annotated
the images in three stages: (i) image level annotation for the 251 concepts; (ii)
bounding box annotation; (iii) textual description annotation. A subset of these
samples was then selected for subtask 3 and further annotated by the organisers
with correspondence annotations between bounding box instances and terms in
textual descriptions.</p>
        <p>The development set for the main subtask contained 2,000 samples, out of
which 500 samples were further annotated and used as the development set for
4 Dataset available at http://risenet.prhlt.upv.es/webupv-datasets
subtask 3. Only 1,979 samples from the development set include at least one
bounding box annotation. The number of textual descriptions for the
development set ranged from 5 to 51 per image (with a mean of 9.5 and a median of
8 descriptions). The test set for subtasks 1 and 2 contains 3,070 samples, while
the test set for subtask 3 comprises 450 samples which are disjoint from the test
set of subtasks 1 and 2.</p>
        <p>For the teaser task, 3,337 random image-article pairs were selected from the
BreakingNews dataset as the development set; these are disjoint from the 10,123
selected in the main dataset. Again, each image corresponds to only one article.</p>
        <p>Like last year, the training and the test images were all contained within
the 510,123 images. In the case of the teaser task, we divided the dataset into
310,123 for training and 200,000 for testing, where all 10,123 documents from the
BreakingNews dataset were contained within the 200,000 test set. Participants
of the teaser task were thus not allowed to explore the data for these 200,000
test documents.</p>
        <p>The training and development sets for all tasks were released approximately
three months before the submission deadline. For subtasks 1 and 2, participants
were expected to provide classi cation/generate a description for all 510,123
images. The test data for subtask 3 was released one week before the submission
deadline. While the train/test split for the teaser tasks was provided right from
the beginning, the test input was only released 1.5 months before the
deadline. The test data were 180,000 text documents extracted from a subset of the
web pages in the 200,000 test split. Text extraction was performed using the
get text() method of the Beautiful Soup library5, after removal of unwanted
elements (and their content) such as script or style. A maximum of 10 submissions
per subtask (also referred to as runs) was allowed per participating group.
Textual Data: Four sets of data were made available to the participants. The
rst one was the list of words used to nd the image when querying the search
engines, along with the rank position of the image in the respective query and
search engine used. The second set of textual data contained the image URLs as
referenced in the web pages they appeared in. In many cases, the image URLs
tend to be formed with words that relate to the content of the image, which
is why they can also be useful as textual features. The third set of data was
the web pages in which the images appeared, for which the only preprocessing
was a conversion to valid XML just to make any subsequent processing simpler.
The nal set of data were features obtained from the text extracted near the
position(s) of the image in each web page it appeared in.</p>
        <p>To extract the text near the image, after conversion to valid XML, the script
and style elements were removed. The extracted texts were the web page title,
and all the terms closer than 600 in word distance to the image, not including
the HTML tags and attributes. Then a weight s(tn) was assigned to each of the</p>
        <sec id="sec-2-4-1">
          <title>5 https://www.crummy.com/software/BeautifulSoup/</title>
          <p>words near the image, de ned as
s(tn) = P</p>
          <p>1
8t2T
s(t)</p>
          <p>X
8tn;m2T</p>
          <p>
            Fn;m sigm(dn;m) ;
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
where tn;m are each of the appearances of the term tn in the document T , Fn;m
is a factor depending on the DOM (e.g. title, alt, etc.) similar to what is done
in the work of La Cascia et al. [10], and dn;m is the word distance from tn;m
to the image. The sigmoid function was centered at 35, had a slope of 0.15 and
minimum and maximum values of 1 and 10 respectively. The resulting features
include for each image at most the 100 word-score pairs with the highest scores.
Visual Features: Before visual feature extraction, images were ltered and
resized so that the width and height had at most 240 pixels while preserving
the original aspect ratio. These raw resized images were provided to the
participants but also eight types of precomputed visual features. The rst feature
set Colorhist consisted of 576-dimensional colour histograms extracted using our
implementation. These features correspond to dividing the image in 3 3 regions
and for each region obtaining a colour histogram quanti ed to 6 bits. The second
feature set GETLF contained 256-dimensional histogram based features. First,
local color-histograms were extracted in a dense grid every 21 pixels for windows
of size 41 41. Then, these local color-histograms were randomly projected to
a binary space using eight random vectors and considering the sign of the
resulting projection to produce the bit. Thus, obtaining an 8-bit representation of
each local color-histogram that can be regarded as a word. Finally, the image
is represented as a bag-of-words, leading to a 256-dimensional histogram
representation. The third set of features consisted of GIST [13] descriptors. The
following four feature types were obtained using the colorDescriptors software [19],
namely SIFT, C-SIFT, RGB-SIFT and OPPONENT-SIFT. The con guration
was dense sampling with default parameters and a hard assignment 1,000
dimension codebook using a spatial pyramid of 1 1 and 2 2 [11]. Concatenation of
the vectors of the spatial pyramid resulted in 5,000-dimensional feature vectors.
The codebooks were generated using 1.25 million randomly selected features and
the k-means algorithm. Moreover, nally, CNN feature vectors have been
provided computed as the seventh layer feature representations extracted from a
deep CNN model pre-trained with the ImageNet dataset [17] using the Berkeley
Ca e library6.
2.5
          </p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>Performance Measures</title>
        <p>Subtask 1 Ultimately the goal of an image annotation system is to make
decisions about which concepts to assign and localise to a given image from a
prede ned list of concepts. Consideration on how to measure annotation
performance should be how good and accurate are those decisions. Ideally, a recall
6 More details can be found at https://github.com/BVLC/caffe/wiki/Model-Zoo
measure would also be used to penalise a system that has additional false
positive output. However given di culties and unreliability of the hand labelling of
the concepts for the test images it was not possible to guarantee all concepts
were labelled. However, the labels present are assumed to be accurate and of a
high quality.</p>
        <p>
          The annotation and localization of Subtask 1 were evaluated using the
PASCAL VOC [6] style metric of intersection over union (IoU), IoU is de ned as
IoU = jBBfg \ BBgtj
jBBfg [ BBgtj
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
Where BB is a rectangle bounding box, f g is a foreground proposed
annotation label, gt is the ground truth label of the concept. It calculates the area
of intersection between the foreground in the proposed output localization and
the ground-truth bounding box localization, divided by the area of their union.
IoU is superior to a more simple measure of the percentage of correctly labelled
pixels as IoU is normalised by the size of the object automatically and penalises
segmentation's that include the background. Causing small changes in the
percentage of correctly labelled pixels to correspond to large di erences in IoU, and
as the dataset has a wide variation in object size, the performance increases
from our approach are more reliably measured. The evaluation of the ground
truth and proposed output overlap was recorded from 0% to 90%. At 0%, this
is equivalent to an image level annotation output, and 50% is the standard
PASCAL VOC style metric used. The localised IoU is then used to compute
the mean average precision (MAP) of each concept independently. The MAP
is reported both per concept and averaged over all concepts. In comparison to
previous years, the MAP was averaged over all possible concept labels in the
test data, instead of just the concepts the participant used. This was to penalise
correctly approaches that only contained a subset of a full approach such as
a face detector, as these were producing unrepresentative performances overall
MAP, however, registering on only a few concepts.
        </p>
        <p>Subtask 2 Subtask 2 was evaluated using the Meteor evaluation metric [4],
which is an F -measure of word overlaps taking into account stemmed words,
synonyms, and paraphrases, with a fragmentation penalty to penalise gaps and
word order di erences. This measure was chosen as it was shown to correlate
well with human judgments in evaluating image descriptions [5]. Please refer to
Denkowski and Lavie [4] for details about this measure.</p>
        <p>Subtask 3 Subtask 3 was evaluated with the ne-grained metric for content
selection which we introduced in last year's edition. Please see [8] or [27] for a
detailed description. The content selection metric is the F1 score averaged across
all 450 test images, where each F1 score is computed from the precision and
recall averaged over all gold standard descriptions for the image. Intuitively, this
measure evaluates how well the sentence generation system selects the correct
concepts to be described against gold standard image descriptions. Formally, let
I = fI1; I2; :::IN g be the set of test images. Let GIi = fGI1i ; GI2i ; :::; GIMi g be the
set of gold standard descriptions for image Ii, where each GImi represents the set
of unique bounding box instances referenced in gold standard description m of
image Ii. Let SIi be the set of unique bounding box instances referenced by the
participant's generated sentence for image Ii. The precision P Ii for test image
Ii is computed as:
where jGImi \ SIi j is the number of unique bounding box instances referenced in
both the gold standard description and the generated sentence, and M is the
number of gold standard descriptions for image Ii.</p>
        <p>Similarly, the recall RIi for test image Ii is computed as:</p>
        <p>P Ii =
1 XM jGImi \ SIi j</p>
        <p>M m jSIi j
RIi =
1 XM jGImi \ SIi j</p>
        <p>M m jGImi j
F Ii = 2</p>
        <p>P Ii
P Ii + RIi</p>
        <p>
          RIi
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
The content selection score for image Ii, F Ii , is computed as the harmonic mean
of P Ii and RIi :
The nal P , R and F scores are computed as the mean P , R and F scores across
all test images.
        </p>
        <p>
          The advantage of the macro-averaging process in equations (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) and (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) is that
it implicitly captures the relative importance of the bounding box instances
based on how frequently to which they are referred across the gold standard
descriptions.
        </p>
        <p>Teaser task For the teaser task, participants are requested to rank the 200,000
test images according to their distance to each input text document. Recall
at the k-th rank position (R@k) of the ground truth image were used as the
performance metrics. The testing of several values of k was performed, and
participants were asked to submit the top 100 ranked images. Please refer to
Hodosh et al. [9] for more details about the metrics.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Results</title>
      <sec id="sec-3-1">
        <title>Participation</title>
        <p>This year the participation was not so good as 2015 where it increased
considerably in previous years. In total seven groups took part in the task and submitted
overall 50 system runs. All seven participating groups submitted a working paper
describing their system, thus for these there were speci c details available:
{ CEA LIST: [2] The team from CEA, LIST, Laboratory of Vision and Content
Engineering, France, represented by Herve Le Borgne, Etienne Gadeski, Ines
Chami, Thi Quynh Nhi Tran, Youssef Tamaazousti, Alexandru Lucian G^nsca
and Adrian Popescu.
{ CNRS TPT: [18] The team from CNRS TELECOM ParisTech, France,
represented by Hichem Sahbi.
{ DUTh: [1] The team from Democritus University of Thrace, DUTh, Greece,
was represented by Georgios Barlas, Maria Ntonti and Avi Arampatzis.
{ ICTisia: [29] The team from Key Laboratory of Intelligent Information
Processing, Institute of Computing Technology Chinese Academy of Sciences,
China, represented by Yongqing Zhu, Xiangyang Li, Xue Li, Jian Sun,
Xinhang Song and Shuqiang Jiang.
{ INAOE: [14] The team from Instituto Nacional de Astrof sica, Optica y
Electronica (INAOE), Mexico was represented by Luis Pellegrin, A. Pastor
LopezMonroy, Hugo Jair Escalante and Manuel Montes-Y-Gomez.
{ MRIM-LIG: [15] The team from LIG - Laboratoire d'Informatique de
Grenoble, and CNRS Grenoble, France, was represented by Maxime Portaz, Mateusz
Budnik, Philippe Mulhem and Johann Poignant.
{ UAIC: [3] The team from UAIC: Faculty of Computer Science,
\Alexandru Ioan Cuza" University, Romania, represented by Alexandru Cristea and
Adrian Iftene.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results for Subtask 1: Image Annotation and Localization</title>
        <p>Unfortunately subtask 1 had a lower participation than last year, however there
were some excellent results showing improvements over previous years. All
submissions were able to provide results on all 510,123 images, indicating that
all groups have developed systems that are scalable enough to annotate large
amounts of images. However one group only processed 180K image
(MRIMLIG [15]) due to computational constraints. Final results are presented in
Table 1 in terms of mean average precision (MAP) over all images of all concepts,
with both 0% overlap (i.e. no localization) and 50% overlap.</p>
        <p>Three of the four groups have achieved good performance across the dataset,
in particular, the approach of CEA LIST. An excellent result given the
challenging nature of the images used and the wide range of concepts provided. The
graph in Figure 3 shows the performance of each submission for an increasing
amount of overlap of the ground truth labels. All the approaches show a steady
drop o in performance which is encouraging, illustrating that the approaches
do not fail to detect some concepts correctly even with a high degree of accuracy.
Even 90% overlap with the ground truth the MAP for CEA LIST was 0.20,</p>
        <p>Group
CEA LIST
MRIM-LIG</p>
        <p>CNRS
UAIC
which is impressive. The results from the groups seem encouraging, and the
approaches use a now standard CNN as their foundation. Improved neural network
structures such as the proposed approach from VGG [20], have provided much
of this improvement.</p>
        <p>CEA LIST used a recent deep learning framework [20], however, focused on
improving the localisation of the concepts. They attempted to use a face body
part detector, boosted by last year's results. However, the use of a face detector
was oversold in the previous years results and didn't improve the performance.
They used EdgeBoxes a generic objectness object detector, however the
performance also didn't increase as expected in the test runs. They hypothesise that
this could be due to the generation of many more candidate bounding boxes, and
a signi cant number estimate the concept incorrectly. MRIM-LIG also used
a classical deep learning framework and the object localisation of [21], where
an apriori set of bounding boxes are de ned which are expected to contain a
single concept each. They also investigated the false lead on performance
improvement through face detection, with a similar lack of performance increase.
Finally CNRS focused on concept detection and used label enrichment to
increase the training data quantity in conjunction with an SVM and VGG [20]
deep network. As each group could submit ten di erent approaches, in general,
the best-submitted approaches contained a fusion of all the various components
of their proposed approaches.</p>
        <p>Some of the test images have nearly 100 ground truth labelled concepts, and
due to limited resources, some of the submitted groups might not have labelled
all possible concepts in each image. However, Fig. 4 shows a similar performance
between groups as previously in Fig. 3. An improvement over previous years
where groups struggled to annotate the 500K images fully.</p>
        <p>Much of the di erence between the groups, is their ability to localise the
concepts e ectively. The ten concepts with the highest average MAP across the
groups, with 0% overlap with the bounding box are in general human-centric:
face, hair, arm, woman, tree, man, car, ship, dress and airplane. These are
nonrigid classes that are being detected on the image, however not yet successfully
localised in the picture as well. With the constraint of 50% overlap with the
ground truth bounding box, the list becomes more based around de ned
objects: car, aeroplane, hair, park, oor, boot, sea, street, face and tree. These are
objects that have a rigid shape that can be learnt. Table 2 shows numerical
examples of the most successfully localised concepts, together with the percentage
of concept occurrence per image in the test data. No method managed to
localise 38 concepts, these include the concepts: nut, mushroom, banana, ribbon,
planet, milk, orange fruit and strawberry. These are smaller and less represented
concepts, in both the test and validation data, in generally occurring in less that
2% of the test images. In fact, many of these concepts were poorly localised in
the previous years challenge too, making this an area to direct the challenge
objectives in future years.</p>
        <p>Discussion for subtask 1 From a computer vision perspective, we would argue
that the ImageCLEF challenge has two key di erences in its dataset
construction to that of the other popular data sets ImageNet [17] and MSCOCO [12].
All three are working on detection and classi cation of concepts within images.
However, the ImageCLEF dataset is created from Internet web pages,
providing a fundamental di erence to the other popular datasets. The web pages are
unsorted and unconstrained meaning the relationship or quality of the text and
image about a concept can be very variable. Therefore, instead of a high-quality
Flickr style photo of a car from ImageNet, the image in the ImageCLEF dataset
could be a fuzzy abstract car shape in the corner of the image. Allowing the
ImageCLEF image annotation challenge to provide additional opportunities to
test proposed approaches on. Another important di erence is that in addition
to the image, text data from web pages can be used to train and generate the
output description of the image in a natural language form.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Results for Subtask 2: Natural Language Caption Generation</title>
        <p>For subtask 2, participants were asked to generate sentence-level textual
descriptions for all 510,123 training images. Two teams, ICTisia and UAIC,
participated in this subtask. Table 3 shows the Meteor scores, for all submitted
runs by both participants. The Meteor score for the human upper-bound was
estimated to be 0.3385 via leave-one-out cross validation, i.e. by evaluating one
description against the other descriptions for the same image and repeating the
process for all descriptions.</p>
        <p>ICTisia achieved the better Meteor score of 0.1837, by building on the
stateof-the-art joint CNN-LSTM image captioning system, but ne-tuning the
parameters of the image CNN as well as the LSTM. On the other hand, UAIC, who
also participated last year, improved on their Meteor score with 0.0934
compared to their best performance from last year (0.0813). They generated image
descriptions using a template-based approach and leveraged external ontologies
and CNNs to improve their results compared to their submissions from last year.</p>
        <p>Neither teams have managed to bridge the gap between system performance
and the human upper-bound this year, showing that there is still scope for further
improvement on the task of generating image descriptions.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Results for Subtask 3: Content Selection</title>
        <p>For subtask 3 on content selection, participants were provided with gold standard
labelled bounding box inputs for 450 test images, released one week before the
submission deadline. Participants were expected to develop systems capable of
predicting, for each image, the bounding box instances (among the gold standard
input) that will be mentioned in the gold standard human-authored textual
descriptions.</p>
        <p>Two teams, DUTh and UAIC, participated in this task. Table 4 shows the
F -score, Precision and Recall across 450 test images for each participant, both
of whom submitted only a single run. The generation of a random per image
baseline by selecting at most three bounding boxes from the gold standard input
at random was perfomed. Like subtask 2, a human upper-bound was computed
via leave-one-out cross validation. The results for these are also shown in Table 4.
As observed, both participants performed signi cantly better than the random
baseline. Compared against the human upper-bound, like subtask 2, much work
can still be done to improve further the performance on the task.</p>
        <p>Unlike the previous two subtasks, neither team used neural networks directly
for content selection. DUTh achieved a higher F -score compared to the best
performing team from last year (0.5459 vs. 0.5310), by training SVM classi ers
to predict whether a bounding box instance is important or not, using various
image descriptors. UAIC used the same system as subtask 2, and while they
did not signi cantly improve on their F -score from last year, their recall score
showed a slight increase. An interesting note is that both teams this year seem to
have concentrated on recall R at the expense of a lower precision P , in contrast
to last year's best performing team who used an LSTM to achieve high precision
but with a much lower recall.
Two teams, CEA LIST and INAOE, participated in the teaser task on text
illustration. Participants were provided with 180,000 text documents as input,
and for each document were asked to provide the top 100 ranked images that
correspond to the document (from a collection of 200,000 images). Table 5 shows
the recall at di erent ranks k (R@k), for a selected subset of 10,112 input
documents comprised of news articles from the BreakingNews dataset (see Sect. 2.4).
Table 6 shows the same results, but on the full 180,000 test documents. Because
the full set of test documents were extracted from generic web pages, the domain
of the text varies. As such, they may consist of noisy documents such as text
from navigational links or advertisements.
INAOE</p>
        <p>Run
INAOE</p>
        <p>Run</p>
        <p>
          This task yielded some interesting results. Bearing in mind the di culty of
the task (selecting one correct image from 200,000 images), CEA LIST yielded
a respectable score that is clearly better than chance performance. The recall also
increased as the rank k is increased. CEA LIST's approach involves mapping
visual and textual modalities onto a common space and combining this method
with a semantic signature. INAOE on the other hand produced excellent results
with run 1, which is a retrieval approach based on a bag-of-words representation
weighted with tf-idf, achieving a recall of 37% even at rank 1 and almost 80% at
rank 100 (in Table 5). In contrast, their runs based on a neural network trained
word2vec representation achieved a much lower recall, although it did increase
to 29.59% at rank 100. Comparing Tables 5 and 6, both teams performed better
on the larger test set of 180,000 generic (and noisy) web text than the smaller
test set of 10,112 restricted to news articles. Although interestingly INAOE's
bag-of-words approach performed worse at smaller ranks (
          <xref ref-type="bibr" rid="ref1 ref10 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-10</xref>
          ) for the full test
set compared to the news article test set, although still signi cantly better than
their word2vec representation. This increase in overall scores, despite the
significant increase in the size of the test set, suggests that there may be some slight
over tting to the training data with most of the methods.
        </p>
        <p>It should be noted that the results of both teams are not directly comparable,
as INAOE based their submission on the assumption that the webpages for test
images are available at test time while CEA LIST did not. This assumption
made the text illustration problem signi cantly less challenging since the test
documents were extracted directly from these webpages, hence the superior
performance by INAOE. On hindsight, this should have been speci ed more clearly
in our task description for a level playing eld. As such we do not consider one
method being superior over the other, but instead concentrate on the technical
contributions of each team.
There are two major limitations that we have identi ed with the challenge this
year. Very few of the groups used the provided data set and features, we found
this surprising, considering the state of the art CNN features and many others
were included. However, this is likely to be due to the complexity and challenge of
the 510,123 web page based images. Given they were from the Internet with little,
a large number of the images are poor representations of the concept. In fact,
some participants annotated a signi cant amount of their more comprehensive
training data, as their learning process assumes perfect or near perfect training
examples, it will fail. As the number of classes increases and become more varied
annotating all comprehensive data will be made more di cult.</p>
        <p>Another shortcoming of the overall challenge is the di culty of ensuring the
ground truth has 100% of concepts labelled, thus allowing a recall measure to
be used. Especially problematic as the concepts selected include ne-grained
categories such as eyes and hands that are small but frequently occur in the
dataset. Also, it was di cult for annotators to reach a consensus in annotating
bounding boxes for less well-de ned categories such as trees and eld. Given
the current crowd-source based hand-labelling of the ground truth, the concepts
have missed annotations. Thus, in this edition, a recall measure is not evaluated
for subtask 1.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>This paper presented an overview of the ImageCLEF 2016 Scalable Concept
Image Annotation task, the fth edition of a challenge aimed at developing more
scalable image annotation systems. The focus of the three subtasks and teaser
task available to participants had the goal to develop techniques to allow
computers to annotate the images reliably, localise the di erent concepts depicted
in the images, select important concepts to be described, generate a description
of the scene, and retrieve a relevant image to illustrate a text document.</p>
      <p>The participation was lower than the previous year, however, in general,
the performance of the submitted systems was somewhat superior to last year's
results for subtask 1. In part probably due to the increased CNN usage as the
feature representation had improved localisation techniques. The clear winner of
this year's subtask 1 evaluation was the CEA LIST [2] team, which focused on
using a state of the art CNN architecture and then also investigated improved
localisation of the concepts which helped provide a good performance increase. In
contrast to subtask 1, the participants for subtask 2 did not signi cantly improve
the results from last year. The approaches used were very similar to those of last
year. For subtask 3, both participating teams concentrated on achieving high
recall with traditional approaches like SVM's, compared to last year's winning
team which focused on obtaining high precision with a neural network approach.
For the pilot teaser task of text illustration, both participating teams performed
respectably, with di erent techniques proposed with varied results. Because of
the ambiguity surrounding one aspect of the task description, the results of the
teams are not directly comparable.
are based on clean hand-labelled data and
the supervised and unsupervised data.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>
        The Scalable Concept Image Annotation Task was co-organized by the VisualSense
(ViSen) consortium under the ERA-NET CHIST-ERA D2K 2011 Programme, jointly
supported by UK EPSRC Grants EP/K01904X/1 and EP/K019082/1, French ANR
Grant ANR-12-CHRI-0002-04 and Spanish MINECO Grant PCIN-2013-047. The task
was also supported by the European Union (EU) Horizon 2020 grant READ
(Recognition and Enrichment of Archival Documents) (Ref: 674943).
13. Oliva, A., Torralba, A.: Modeling the Shape of the Scene: A Holistic
Representation of the Spatial Envelope. Int. J. Comput. Vision 42(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ), 145{175 (May 2001),
doi:10.1023/A:1011139631724
14. Pellegrin, L., Lopez-Monroy, A.P., Escalante, H.J., Montes-Y-Gomez, M.: INAOE's
participation at ImageCLEF 2016: Text Illustration Task. In: CLEF2016 Working
Notes. CEUR Workshop Proceedings, CEUR-WS.org, Evora, Portugal (September
2016)
15. Portaz, M., Budnik, M., Mulhem, P., Poignant, J.: MRIM-LIG at ImageCLEF 2016
Scalable Concept Image Annotation Task. In: CLEF2016 Working Notes. CEUR
Workshop Proceedings, CEUR-WS.org, Evora, Portugal (September 2016)
16. Ramisa, A., Yan, F., Moreno-Noguer, F., Mikolajczyk, K.: Breakingnews: Article
annotation by image and text processing. CoRR abs/1603.07141 (2016), http:
//arxiv.org/abs/1603.07141
17. Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z.,
Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet Large
Scale Visual Recognition Challenge. International Journal of Computer Vision
(IJCV) pp. 1{42 (April 2015)
18. Sahbi, H.: CNRS TELECOM ParisTech at ImageCLEF 2016 Scalable Concept
Image Annotation Task: Overcoming the Scarcity of Training Data. In: CLEF2016
Working Notes. CEUR Workshop Proceedings, CEUR-WS.org, Evora, Portugal
(September 2016)
19. van de Sande, K.E., Gevers, T., Snoek, C.G.: Evaluating Color Descriptors for
Object and Scene Recognition. IEEE Transactions on Pattern Analysis and Machine
Intelligence 32, 1582{1596 (2010), doi:10.1109/TPAMI.2009.154
20. Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale
image recognition. In: arXiv preprint arXiv:1409.1556 (2014)
21. Uijlings, J.R., van de Sande, K.E., Gevers, T., Smeulders, A.W.: Selective search
for object recognition. International journal of computer vision 104(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), 154{171
(2013)
22. Villegas, M., Muller, H., Garc a Seco de Herrera, A., Schaer, R., Bromuri, S.,
Gilbert, A., Piras, L., Wang, J., Yan, F., Ramisa, A., Dellandrea, E., Gaizauskas,
R., Mikolajczyk, K., Puigcerver, J., Toselli, A.H., Sanchez, J.A., Vidal, E.: General
Overview of ImageCLEF at the CLEF 2016 Labs. Lecture Notes in Computer
Science, Springer International Publishing (2016)
23. Villegas, M., Paredes, R.: Image-Text Dataset Generation for Image Annotation
and Retrieval. In: Berlanga, R., Rosso, P. (eds.) II Congreso Espan~ol de
Recuperacion de Informacion, CERI 2012. pp. 115{120. Universidad Politecnica de
Valencia, Valencia, Spain (June 18-19 2012)
24. Villegas, M., Paredes, R.: Overview of the ImageCLEF 2012 Scalable Web
Image Annotation Task. In: Forner, P., Karlgren, J., Womser-Hacker, C. (eds.)
CLEF 2012 Evaluation Labs and Workshop, Online Working Notes. Rome,
Italy (September 17-20 2012), http://mvillegas.info/pub/Villegas12_CLEF_
Annotation-Overview.pdf
25. Villegas, M., Paredes, R.: Overview of the ImageCLEF 2014 Scalable Concept
Image Annotation Task. In: CLEF2014 Working Notes. CEUR Workshop
Proceedings, vol. 1180, pp. 308{328. CEUR-WS.org, She eld, UK (September 15-18
2014), http://ceur-ws.org/Vol-1180/CLEF2014wn-Image-VillegasEt2014.pdf
26. Villegas, M., Paredes, R., Thomee, B.: Overview of the ImageCLEF 2013
Scalable Concept Image Annotation Subtask. In: CLEF 2013 Evaluation Labs and
Workshop, Online Working Notes. Valencia, Spain (September 23-26 2013), http:
//mvillegas.info/pub/Villegas13_CLEF_Annotation-Overview.pdf
27. Wang, J., Gaizauskas, R.: Generating image descriptions with gold standard visual
inputs: Motivation, evaluation and baselines. In: Proceedings of the 15th European
Workshop on Natural Language Generation (ENLG). pp. 117{126. Association for
Computational Linguistics, Brighton, UK (September 2015), http://www.aclweb.
org/anthology/W15-4722
28. Wang, J.K., Yan, F., Aker, A., Gaizauskas, R.: A poodle or a dog? Evaluating
automatic image annotation using human descriptions at di erent levels of
granularity. In: Proceedings of the Third Workshop on Vision and Language. pp. 38{45.
Dublin City University and the Association for Computational Linguistics, Dublin,
Ireland (August 2014)
29. Zhu, Y., Li, X., Li, X., Sun, J., Song, X., Jiang, S.: Joint Learning of CNN and
LSTM for Image Captioning. In: CLEF2016 Working Notes. CEUR Workshop
Proceedings, CEUR-WS.org, Evora, Portugal (September 2016)
      </p>
    </sec>
    <sec id="sec-6">
      <title>A Concept List 2016</title>
      <p>The following tables present the 251 concepts used in the ImageCLEF 2016 Scalable
Concept Image Annotation task. In the electronic version of this document, each
concept name is a hyperlink to the corresponding WordNet synset webpage.
#dev. #test
hog
hole
hook
horse
hospital
house
jacket
jean
key
keyboard
kitchen
knife
ladder
lake
leaf
leg
letter
library
lighter
lion
lotion
magazine
male child
man
mask
mat
mattress
microphone
milk
mirror
monkey
motorcycle
mountain
mouse
mouth
mushroom</p>
      <p>neck
necklace
necktie
nest
newspaper
nose
nut
office
onion
orange
oven
painting
pan
park
pen
pencil
piano
picture
pillow
planet
pool
pot
potato
prison
pumpkin
rabbit
rack
radio
ramp
ribbon</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barlas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ntonti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arampatzis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>DUTh at the ImageCLEF 2016 Image Annotation Task: Content Selection</article-title>
          .
          <source>In: CLEF2016 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Evora, Portugal (
          <year>September 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Borgne</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gadeski</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chami</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>T.Q.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamaazousti</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , G^nsca,
          <string-name>
            <given-names>A.L.</given-names>
            ,
            <surname>Popescu</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Image annotation and two paths to text illustration</article-title>
          .
          <source>In: CLEF2016 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Evora, Portugal (
          <year>September 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cristea</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iftene</surname>
            ,
            <given-names>A.:</given-names>
          </string-name>
          <article-title>Using Machine Learning Techniques, Textual and Visual Processing in Scalable Concept Image Annotation Challenge</article-title>
          .
          <source>In: CLEF2016 Working Notes. CEUR Workshop Proceedings</source>
          , CEUR-WS.org, Evora, Portugal (
          <year>September 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Denkowski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavie</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Meteor universal: Language speci c translation evaluation for any target language</article-title>
          .
          <source>In: Proceedings of the EACL 2014 Workshop on Statistical Machine Translation</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Elliott</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Keller, F.:
          <article-title>Comparing automatic evaluation measures for image description</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          . pp.
          <volume>452</volume>
          {
          <fpage>457</fpage>
          . Association for Computational Linguistics, Baltimore, Maryland (
          <year>June 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Everingham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eslami</surname>
            ,
            <given-names>S.M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Gool</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>C.K.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The pascal visual object classes challenge: A retrospective</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>111</volume>
          (
          <issue>1</issue>
          ),
          <volume>98</volume>
          {136 (Jan
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.):
          <article-title>WordNet An Electronic Lexical Database</article-title>
          . The MIT Press, Cambridge, MA; London (May
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gilbert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piras</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dellandrea</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaizauskas</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolajczyk</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Overview of the imageclef 2015 scalable image annotation, localization and sentence generation task</article-title>
          . In: Working Notes of CLEF 2015 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , Toulouse, France, September 8-
          <issue>11</issue>
          ,
          <year>2015</year>
          . (
          <year>2015</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1391</volume>
          /inv-pap6
          <source>-CR</source>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hodosh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Young</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hockenmaier</surname>
          </string-name>
          , J.:
          <article-title>Framing image description as a ranking task: Data, models and evaluation metrics</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          (JAIR)
          <volume>47</volume>
          (
          <issue>1</issue>
          ),
          <volume>853</volume>
          {899 (May
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>La</given-names>
            <surname>Cascia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Sethi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sclaro</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.:</surname>
          </string-name>
          <article-title>Combining textual and visual cues for contentbased image retrieval on the World Wide Web</article-title>
          .
          <source>In: Content-Based Access of Image and Video Libraries</source>
          ,
          <year>1998</year>
          . Proceedings. IEEE Workshop on. pp.
          <volume>24</volume>
          {
          <issue>28</issue>
          (
          <year>1998</year>
          ), doi:10.1109/IVL.
          <year>1998</year>
          .694480
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lazebnik</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponce</surname>
          </string-name>
          , J.:
          <article-title>Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene Categories</article-title>
          .
          <source>In: Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision</source>
          and Pattern Recognition - Volume
          <volume>2</volume>
          . pp.
          <volume>2169</volume>
          {
          <fpage>2178</fpage>
          . CVPR '06, IEEE Computer Society, Washington, DC, USA (
          <year>2006</year>
          ), doi:10.1109/CVPR.
          <year>2006</year>
          .68
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belongie</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hays</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perona</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramanan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dollar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zitnick</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>Microsoft COCO: common objects in context</article-title>
          .
          <source>CoRR abs/1405</source>
          .0312 (
          <year>2014</year>
          ), http://arxiv.org/abs/1405.0312
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>