<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Visual Concept Detection Task in ImageCLEF 2008</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Deselaers</string-name>
          <email>deselaers@cs.rwth-aachen.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allan Hanbury</string-name>
          <email>hanbury@prip.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>PRIP, Institute of Computer-Aided Automation, Vienna University of Technology</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>RWTH Aachen University, Computer Science Department</institution>
          ,
          <addr-line>Aachen</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Visual Concept Detection Task (VCDT) of ImageCLEF 2008 is described. A database of 2,827 images were manually annotated with 17 concepts. Of these, 1,827 were used for training and 1,000 for testing the automated assignment of categories. In total 11 groups participated and submitted 53 runs. The runs were evaluated using ROC curves, from which the Area Under the Curve (AUC) and Equal Error Rate (EER) were calculated. For each concept, the best runs obtained an AUC of 80% or above.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Searching for images is, despite intensive research on alternative methods in the last 20 years, still
a task which is mainly done based on textual information. For a long time, searching for images
based on text was the most feasible method because on the one hand, the number of images to
be searched was rather restricted, and on the other hand, only few people needed to access huge
repositories of images. Both of these conditions have changed. The number of available images
is growing more rapidly than ever due to the falling prices of high-end imaging equipment for
professional use and of digital cameras for consumer use. Publicly available image databases such
as Google picassa and Flickr have become major sites of interest on the Internet.</p>
      <p>Nevertheless, accessing images is still a tedious task because sites such as Flickr do not allow
images to be accessed based on their content but only based on the annotations that users create.
These annotations are commonly disorganised, not very precise, and multilingual. Access problems
can be addressed by improving the textual access methods, but none of these improvements can
ever be perfect as long as the users do not annotate their images perfectly, which is very unlikely.
Therefore, content-based methods have to be employed to improve access methods to digitally
stored images.</p>
      <p>A problem with content-based methods is that they are often costly and cannot be applied
in real-time. An intermediate step is to automatically create textual labels based on the images'
content. To make these labels as useful as possible, frequently occurring visual concepts should
be annotated in a standard manner.</p>
      <p>In the visual concept detection task (VCDT) of ImageCLEF 2008, the aim was to apply labels
of frequent categories in the photo retrieval task to the images and evaluate how well automated
visual concept annotation algorithms function. Additionally, participants of the VCDT could
create annotations for all images used in the photo retrieval task, which were provided to the
participants of this task.
Person</p>
      <p>Animal</p>
      <p>Water
(i.e. river,
lake, etc)</p>
      <p>Sky</p>
      <p>Sunny
Partly
Cloudy
Overcast</p>
      <p>Day
Night</p>
      <p>Road or
pathway</p>
      <p>Buildings</p>
      <p>Beach
Mountains</p>
      <p>Vegetation</p>
      <p>Tree</p>
      <p>In the following, we describe the visual concept detection task of ImageCLEF 2008, the
database used, the methods of the participating groups, and the results.
2</p>
      <p>
        Database and Task Description
As database for the ImageCLEF 2008 visual concept detection task, a total of 2,827 images were
used. These are taken from the same pool of images used to create the IAPR-TC12 database [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
but are not included in the IAPR-TC12 database used in the ImageCLEF photo retrieval task.
      </p>
      <p>The visual concepts were chosen based on concepts used in work on visual concept annotation.
They were organised in a hierarchy, shown in Figure 1.</p>
      <p>
        As for the object detection task in 2007 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a web interface was created for manual annotation
of the images by the concepts. Annotation was mainly carried out by undergraduate students
at the RWTH Aachen University and by the track coordinators. A general opinion expressed by
the annotators was that the concept annotation required more time than the object annotation
of 2007. The number of images that were voluntarily annotated this year was also signi cantly
smaller than the 20,000 images annotated by object labels in 2007.
      </p>
      <p>Of the 2,827 manually annotated images, 1,827 were distributed with annotations to the
participants as training data. The remaining 1,000 images were provided without labels as test data.
The participants' task was to apply labels to these 1,000 images.</p>
      <p>Table 1 gives an overview of the frequency of the 17 visual concepts in the training data and
in the test data and Figure 2 gives an example image for each of the categories.</p>
      <p>For each run, results for each concept were evaluated by plotting ROC curves. The results
for each concept were summarised by two values: the area under the ROC curve (AUC) and the
Equal Error Rate (EER). The latter is the error rate at which the false positive rate is equal to
the false negative rate. Furthermore, for each run, the average AUC and average EER over all
concepts were calculated.</p>
      <p>indoor
night
tree
sky</p>
      <p>Results from the Evaluation
In total 11 groups participated and submitted 53 runs. Below, we brie y describe the methods
employed by each group:
CEA-LIST. The Lab of applied research on software-intensive technologies of the CEA, France
submitted 3 runs using di erent image features accounting for color and spatial layout with
nearest neighbour and support vector machine classi ers.</p>
      <p>HJ FA. The Microsoft Key Laboratory of Multimedia Computing and Communication of the
University of Science and Technology, China submitted one run using color and SIFT
descriptors which are combined and classi ed using a nearest neighbour classi er.
IPAL I2R. The IPAL French-Singaporean Joint Lab of the Institute for Infocomm Research in</p>
      <p>Singapore submitted 8 runs using a variety of di erent image descriptors.</p>
      <p>LSIS. The Laboratory of Information Science and Systems, France submitted 7 runs using a
structural feature combined with several other features using multi-layer perceptrons.
MMIS. The Multimedia and Information Systems Group of the Open University, UK submitted
4 runs using CIELAB and Tamura features and combinations of these.</p>
      <p>Makerere. The Faculty of Computing and Information Technology, Makerere University, Uganda
submitted one run using luminance, dominant colors, and di erent texture and shape features
which are classi ed using a nearest neighbour classi er.</p>
      <p>RWTH. The Human Language Technology and Pattern Recognition Group from RWTH Aachen
University, Germany submitted one run using a patch-based bag-of-visual words approach
using a log-linear classi er.</p>
      <p>TIA. The Group for Machine Learning for Image Processing and Information Retrieval from the
National Institute of Astrophysics, Optics and Electronics, Mexico submitted 7 runs using
global and local features with support vector machines and random forest classi ers.
UPMC. The University Pierre et Marie Curie in Paris, France submitted 5 runs using fuzzy
decision forests.</p>
      <p>XRCE. The Textual and Visual Pattern Analysis group from the Xerox Research Center Europe
in France submitted two runs using multi-scale, regular grid, patch-based image features and
a Fisher-Kernel Vector classi er.
budapest. The Datamining and Websearch Research Group, Hungarian Academy of Sciences,
Hungary submitted 13 runs using a wide variety of di erent features, classi ers, and
combinations.</p>
      <p>The average EER and average AUC for each submitted run are given in Table 2. From this
table, it can be seen that the best overall runs were submitted by XRCE.</p>
      <p>Table 3 shows a breakdown of the results per concept. For each concept, the best and worst
EER and AUC are shown, along with the average EER and AUC over all runs submitted. The
best results were obtained for all concepts by XRCE, with budapest doing equally well on the night
concept. The AUC per concept for all the best runs is 80.0% or above. Among the best results,
the concepts having the highest scores are indoor and night . The concept with the worst score
among the best results is road or pathway, most likely due to the high variability in the appearance
of this concept. The concept with the highest average score, in other words, the concept that was
detected best in most runs is sky. Again, the concept with the worst average score is road or
pathway.</p>
      <p>Only one group participating in the photo retrieval task of ImageCLEF made use of the concept
annotation created by participants of the VCDT. The group from Universit Pierre et Marie Curie
in Paris, France used the annotation in 12 of their 19 runs. However, the runs of that group were
not ranked very highly and thus it is hard to judge how big the impact was. Their best two runs
used the annotation and have a P (20) value of slightly over 0.26 while their third best run, which
did not use these data has a P (20) value of 0.25.
4</p>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>This paper summarises the ImageCLEF 2008 Visual Concept Detection Task. The aim was to
automatically annotate images with concepts, with a list of 17 hierarchically organised concepts
provided. The results demonstrate that this task can be solved reasonably well, with the best run
having an average AUC over all concepts of 90.66%. Six further runs obtained AUCs between
80% and 90%. When evaluating the runs on a per concept basis, the best run also obtained an
AUC of 80% or above for every concept. Concepts for which automatic detection was particularly
successful are: indoor/outdoor , night , and sky. The worst results were obtained for the concept
road or pathway.</p>
      <p>group
CEA_LIST
CEA_LIST
CEA_LIST
HJ_FA
IPAL_I2R
IPAL_I2R
IPAL_I2R
IPAL_I2R
IPAL_I2R
IPAL_I2R
IPAL_I2R
IPAL_I2R
LSIS
LSIS
LSIS
LSIS
LSIS
LSIS
LSIS
MMIS
MMIS
MMIS
MMIS
Makerere
RWTH
TIA
TIA
TIA
TIA
TIA
TIA
TIA
UPMC
UPMC
UPMC
UPMC
UPMC
UPMC
XRCE
XRCE
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
budapest
run
#
concept
EER
group
average</p>
      <p>worst
97.4
96.6
89.7
85.7
97.4
84.6
80.0
89.9
88.3
93.8
86.8
89.7
95.7
96.4
92.1
93.7
85.7</p>
      <p>XRCE
XRCE
XRCE
XRCE
XRCE/budapest
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE
XRCE</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Deselaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Viitaniemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Benczur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brendel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Daroczy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Escalante</given-names>
            <surname>Balderas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gevers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hernandez Gracidas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C. H.</given-names>
            <surname>Hoi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Laaksonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. Marin</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Rui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sebe</surname>
          </string-name>
          , J. Stottinger, and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Overview of the imageclef 2007 object retrieval task</article-title>
          .
          <source>In Proceedings of the CLEF 2007 Workshop</source>
          ,
          <year>2007</year>
          . to appear.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Grubinger</surname>
          </string-name>
          , Paul Clough, Henning Muller, and Thomas Deselaers.
          <source>The IAPR TC-12</source>
          benchmark
          <article-title>- a new evaluation resource for visual information systems</article-title>
          .
          <source>In Proceedings of the International Workshop OntoImage'2006</source>
          , pages
          <fpage>13</fpage>
          {
          <fpage>23</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>