<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MUFIN at ImageCLEF 2011: Success or Failure?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petra Budikova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Batko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Zezula</string-name>
          <email>zezulag@fi.muni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Masaryk University</institution>
          ,
          <addr-line>Brno</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>In all elds of research it is important to discuss and compare various methods that are being proposed to solve given problems. In image retrieval, the ImageCLEF competitions provide such comparison platform. We have participated in the Photo Annotation Task of the ImageCLEF 2011 competition with a system based on the MUFIN Annotation Tool. Our approach is, in contrast to typical classi er solutions, based on a general annotation system for web images that provides general keywords for arbitrary image. However, the free-text annotation needs to be transformed into the 99 concepts given by the competition task. The transformation process is described in detail in the rst part of this paper. In the second part, we discuss the results achieved by our solution. Even though the free-text annotation approach was not as successful as the classi er-based approaches, the results are competitive especially for the concepts involving real-world objects. On the other hand, our approach does not require training and is scalable to any number of concepts.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        In the course of the past few decades, multimedia data processing has become
an integral part of many application elds, including medicine, art, security, etc.
This poses a number of challenges to the computer science { we need techniques
for e cient data representation, storing and retrieval. In many applications, it is
also necessary to understand the semantic meaning of a multimedia object, i.e. to
know what is represented in a picture or what a video is about. Such information
is usually expressed in a textual form, e.g. as a text description that accompanies
the multimedia object. The semantic description of an object can be obtained
manually or (semi)-automatically. Since the manual annotation is an extremely
labor-intensive and time-consuming task for larger data collections, automatic
annotation or classi cation of multimedia objects is of high importance. Perhaps
the most intensive is the research on automatic annotation of images, which is
essential for semantic image retrieval [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>In all elds of research it is important to discuss and compare various
methods that are being proposed to solve given problems. In image retrieval, the
ImageCLEF competitions provide such comparison platform. Each year, a set
of speci c tasks is de ned that re ects the most challenging problems of current
research. In 2011 as well as in the two previous years, one of the challenges was
the Photo Annotation Task.</p>
      <p>
        This paper presents the techniques used by the MUFIN group to handle the
Annotation Task. We believe that our methods will be interesting for the
community since our approach is di erent from the solutions presented in previous
years [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Rather than creating a solution tailored for this speci c task we
employed a general purpose annotation system and used the returned keywords to
identify the relevant concepts. The results show that while this approach logically
lags behind the precision of the more nely tuned solutions, it is still capable of
solving some instances pretty well.
      </p>
      <p>The paper is organized as follows. First, we brie y review the related work on
image annotation and distinguish two important classes of annotation methods.
Next, the ImageCLEF Annotation Task is described in more detail and the
quality of the training data is discussed. Finally, we present the methods used
by the MUFIN group, analyze their performance and discuss the results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The reason why we are interested in annotation is to simplify access to the
multimedia data. Depending on a situation, di erent types of metadata may be
needed, both in content and form. In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], three forms of annotation are discussed:
free text, keywords chosen from a dictionary, and concepts from some ontology.
While the free text annotation does not require any structure, the other two
options pose some restrictions on the terminology used, and in particular make
the selection of keywords smaller. On certain conditions, we then begin to call
the task classi cation or categorization rather than annotation.
      </p>
      <p>
        Even though the conditions are not strictly de ned, the common
understanding is that classi cation task works with a relatively small number of concepts
and typically uses machine learning to create specialized classi ers for the given
concepts. To train the classi ers, a su ciently large training dataset with labeled
data needs to be available. The techniques that can be engaged in the learning
are numerous, including SVMs, kNN classi ers, or probabilistic approaches [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] describes a number of classi cation setups using di erent types of
concept learning.
      </p>
      <p>
        On the contrary, annotation usually denotes a task where a very large or
unlimited number of concepts is available and typically no training data is given.
The solution to such task needs to exploit some type of data mining, in case of
image annotation it is often based on content-based image retrieval over
collections of images with rich metadata. Such system is described for example in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Typically, the retrieval-based annotation systems exploit tagged images from
photo-sharing sites.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>ImageCLEF Photo Annotation Task</title>
      <p>
        In the ImageCLEF Photo Annotation Task, the participants were asked to
assign relevant keywords to a number of test images. The full setup of the contest
is described in the Task overview [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. From our perspective, two facts are
important: (1) the keywords to be assigned were chosen from a xed set of 99
concepts, and (2) a set of labeled training images was available. Following our
de nition of terms from the previous section, the task thus quali es as a classi
cation problem. As such, it is most straightforwardly solved by machine learning
approaches, using the training data to tune the parameters of the model. The
quality of the training data is then crucial for the correctness of the classi cation.
      </p>
      <p>
        As explained in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], it is di cult to obtain a large number of labeled images
both for training and contest evaluation. It is virtually impossible to gather such
data only with the help of a few domain experts. Therefore, only a part of the
data was labeled by domain experts. The rest was annotated in a crowdsourcing
way, using workers from the Amazon Mechanical Turk portal. Even though the
organizers of the contest did their best to ensure that only sane results would be
accepted, the gathered data still contain some errors. In the following, we would
like to comment on some of them that we have noticed, so that they could be
corrected in the future. Also, we will discuss later how these errors may have
in uenced the performance of the annotation methods.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Training data de ciencies</title>
        <p>
          During the preparation of our solution for the Photo Annotation Task, we have
identi ed to following types of errors in the labeled training data:
{ Logical nonsense: Some annotations in the training dataset contradict the
laws of the real world. The most signi cant nonsense we found was a
number of images with the following triplet of concepts: single person, man,
woman. Such combination appeared for 784 images. Similarly, still life
and active are concepts that do not match together. Though the emotional
annotations are more subjective and cannot be so easily discarded as
nonsense, we also believe that annotating an image by both cute and scary
concepts is an oxymoron.
{ Annotation inconsistence: Since the contest participants were not provided
by any explanation of the concepts, it was not always clear to us what counts
as relevant for a given concept and what does not. One such unclear concept
was single person. Should a part of a body be counted as a person or
not? It seems that this question was also unclear to the trainset annotators
as the concept bodypart sometimes co-occurred with single person and
sometimes with no person.
{ Concept overuse: The emotional and abstract concepts are de nitely di
cult to assign. Even within our research group, we could not decide what
determines the \cuteness" of an image or what is the de nition of a
\natural" image. Again, the labeled training data were not of much help as we
could not discover any inner logic in them. We suspect that the Amazon
Turk workers solved this problem by using the emotional and abstract terms
as often as possible. For illustration, from among the 8000 training images
3910 were labeled cute, 3346 were considered to be visual arts and 4594
images were perceived as natural.
Each error type is illustrated by a few examples in Figure 1.
The MUFIN Annotation Tool came into existence in the beginning of this year as
an extension of the MUFIN Image Search1, a content-based search engine that
we have been developing for several years. Our aim was to provide an online
annotation system for arbitrary web images. The rst prototype version of the
system is available online2 and was presented in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. As the rst experiments with
the tool provided promising results, we decided to try our tool in the ImageCLEF
Annotation Task.
        </p>
        <p>The fundamental di erence in the basic orientation of MUFIN Annotation
Tool and the Annotation Task is that our system provides annotations while the
task asks for classi cation. Our system provides free-text annotation of images,</p>
        <sec id="sec-3-1-1">
          <title>1 http://mu n. .muni.cz/imgsearch/ 2 http://mu n. .muni.cz/annotation/</title>
          <p>using any keywords that seem relevant using the content-based searching. To be
able to use our tool for the task, we needed to transform the provided keywords
into the restricted set of concepts given by the task. Moreover, even though
the MUFIN tool is quite good at describing the image content it does not give
much information about emotions and technical-related concepts (we will discuss
the reasons later). Therefore, we also had to extend our system and add new
components for specialized processing of some concepts. The overall architecture
of the processing engaged in the Annotation Task is depicted in Figure 2. In the
following sections, we will describe the individual components in more detail.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>MUFIN Annotation Tool</title>
        <p>
          The MUFIN Annotation Tool is based on the MUFIN Image Search engine,
which retrieves the nearest neighbors of a given image based on visual and text
similarity. The image search system, described in full detail in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], enables fast
retrieval of similar images from very large collections. The visual similarity is
de ned by ve MPEG-7 global descriptors { Scalable Color, Color Structure,
Color Layout, Edge Histogram, and Region Shape { and their respective
distance functions [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. If some text descriptions of images are available, their tf-idf
similarity score is also taken into consideration. In the ImageCLEF contest,
freetext image descriptions and EXIF tags were available for some images. Together
with the image, these were used as the input of the retrieval engine.
        </p>
        <p>
          To obtain an annotation of some input image, we rst evaluate the
nearest neighbor query over a large collection of high-quality images with rich and
trustworthy text metadata. In particular, we are currently using the Pro media
dataset, which contains 20M images from a microstock site [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. When the query
is processed by the MUFIN search, we obtain a set of images with their
respective keywords. In the Pro media dataset, each image is accompanied by a set
of title words (typically 3 to 10 words) and keywords (about 20 keywords per
image in average). Both the title words and the keywords of all images in the
result set are merged together (the title words receiving a higher weight) and
the frequencies of individual lemmas are identi ed. A list of stopwords and the
WordNet lexical database [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] are used to remove irrelevant word types, names,
etc. The Annotation Tool then returns the list of the most frequent keywords
with the respective frequencies, which express the con dence of the annotation.
4.2
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Annotation to concept transformation</title>
        <p>To transform the free-text annotation into the ImageCLEF concepts it was
necessary to nd the semantic relations between the individual keywords and
concepts. For this purpose, we used the WordNet lexical database, which provides
structured semantic information for English nouns, verbs, adjectives, and
adverbs. The individual words are grouped into sets of cognitive synonyms (called
synsets), each expressing a distinct concept. These synsets are interlinked by
di erent semantic relations, such as hypernym/hyponym, synonym, meronym,
etc. It is thus possible to nd out whether any two synsets are related and how.
In our case, we were interested in the relationships between the synsets of our
annotation keywords and the synsets of the ImageCLEF concepts. To obtain
them, we rst needed to nd the relevant synsets for our keywords and the task
concepts.</p>
        <p>The WordNet synset is de ned as a set of words with the same meaning,
accompanied by a short description of the respective semantic concept. Very
often, a word has several meanings and therefore is contained in several synsets.
For instance, the word \cat" appears both in a synset describing the domestic
animal and a synset describing an attractive woman. If we have only a keyword
and no other information about the word sense or context, we need to consider
all synsets that contain this keyword. This is the case of keywords returned by
the MUFIN Annotation Tool. We do not try here to determine whether the
synsets are really relevant but rely on the majority voting of a large number of
keywords that are processed.</p>
        <p>The situation is however di erent in case of the ImageCLEF concepts where
it is much more important to know the correct synsets. Fortunately, we have
actually two possible ways of determining the relevant synsets. First, we can
sort them out manually since the number of concepts is relatively small. The
other, more systematic solution, will run the whole annotation process with all
the candidate synsets of the concepts, log the contributions of the respective
individual synsets, evaluate their performance, and rule out those with a low
success rate.</p>
        <p>Once the synsets are determined, we can look for the relationships. Again,
there are di erent types of relations and some of them are relevant for one
concept but irrelevant for another. Again, we have the same two possibilities of
choosing the relevant relationships as in the case of synsets. In our
implementation, we have used the manual cleaning approach for synsets and automatic
selection approach for relationships.</p>
        <p>With the relevant synsets and relationships, we count a relevance score of
each ImageCLEF concept during the processing of each image. The score is
increased each time a keyword-synset is related to concept-synset. The increase
is proportional to the con dence score of the keyword as produced by the MUFIN
Annotation Tool.</p>
        <p>Finally, the concepts are checked against the OWL ontology provided within
the Annotation Task. The concepts are visited in a decreasing order of their
scores and whenever a con ict between two concepts is detected, the concept
with a lower score is discarded.
4.3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Additional image processing</title>
        <p>The mining in keywords of similar images allows us to obtain such information
as is usually contained in the image descriptions. This is most often related to
image content, so the concepts related to nature, buildings, vehicles, etc. can be
identi ed quite well. However, the Annotation Task considers also concepts that
are less often described in the text. To get some more information about these,
we employed the following three additional information sources:
{ Face recognition: The face recognition algorithms are well-known in the
image processing. We employ face recognition to determine the number of
persons in an image.
{ EXIF tag processing: Some of the input photos are accompanied by EXIF
tags that provide information about various image properties. When
available, we use these tags to decide the relevance of concepts related to
illumination, focus, and time of the day.
{ MUFIN visual search: Apart from the large Pro media collection, we also
have the training dataset that can be searched with respect to the visual
similarity of images. However, since the trainset is rather small and some
concepts are represented by only a few images, there is quite a high
probability that the nearest neighbors will not be relevant. Therefore, we only
consider neighbors within a small range of distances (determined by
experiments).
4.4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Trainset statistics input</title>
        <p>De nitely the most di cult concepts to assign are the ones related to user's
emotions and also the abstract concepts such as technical, overall quality,
etc. As discussed in Section 3.1, it is not quite clear either to us or to the people
who annotated the trainset what these concepts precisely mean. Therefore, it is
very di cult to determine their relevance using the image visual content. The
text provided with the images is also not helpful in most cases.</p>
        <p>We nally decided to rely on the correlations between image content and
the emotions it most probably evokes. For example, images of babies or nature
are usually deemed cute. A set of such correlation rules was derived from the
trainset and used to choose the emotional and abstract concepts.
4.5</p>
      </sec>
      <sec id="sec-3-6">
        <title>Our submissions at ImageCLEF</title>
        <p>In our primary run, all the above-described components were integrated as
depicted in Figure 2. In addition, we submitted three more runs where we tried
various other settings to nd out whether the proposed extensions really
provided some added quality to the search results. The components that were left
out in some experiments are depicted by dashed lines in Figure 2. We also
experimented with the transformation of our concept scores into the con dence
values expressed as percentage, which was done either as concept-speci c or
concept-independent. The individual run settings were as follows:
{ Mu nSubmission100: In this run, all the available information was exploited
including photo tags, EXIF tags, visual image descriptors, and trainset
statistics. Concept-speci c mapping of annotation scores to con dence values
was applied.
{ Mu nSubmission101: The same settings were used for this run as in the
previous case but concept-independent mapping of annotation scores was
applied.
{ Mu nSubmission110: In this run, we did not use the EXIF tags for the
processing of concepts related to daytime and illumination as described in
Section 4.3. The MUFIN visual search in the trainset was omitted as well.
However, the textual EXIF tags were used as a part of the input for the MUFIN
Annotation Tool. Concept-independent mapping of annotation scores was
applied.
{ Mu nSubmission120: In this run, the EXIF tags were not applied at all,
neither as a part of the text-and-visual query in the basic annotation step nor
in the additional processing. Again, the MUFIN visual search was skipped.
Concept-independent mapping of annotation scores was applied.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion of results</title>
      <p>
        As detailed in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], the following three quality metrics were evaluated to compare
the submitted results: Mean interpolated Average Precision (MAP), F-measure
(F-ex), and Semantic R-Precision (SR-Precision). As we expected, our best
submission was Mu nSubmission100 which achieved 0.299 MAP, 0.462 F-ex, and
0.628 SR-precision. The other submissions that we have tried received slightly
worse scores. After the task evaluation and the release of the algorithm for
computing the MAP metric, we also re-evaluated our system with better settings of
the MUFIN Annotation Tool that we have improved since the ImageCLEF task
submission. Using these, we were able to gain one or two percent increase of the
MAP score. With respect to the MAP measure, our solution ranked at position
13 among the 18 participating groups.
      </p>
      <p>Apart from the overall results, it is also interesting to take a closer look at the
performance of the various solutions for individual concepts. The complete list of
concept results for each group is available on the ImageCLEF web pages3. Here
we focus on the categories and particular examples of concepts where MUFIN
annotation performed either well or poorly and discuss the possible reasons.</p>
      <p>First of all, we need to de ne what we consider a good result. Basically, there
are two ways: either we only look at the performance, e.g. following the MAP
measure, or we consider the performance in relation to the di culty of assigning
the given concept. The assigning di culty can be naturally derived from the
competition results { when no group was able to achieve high precision with
some concept, then the concept is problematic. Since we believe that the second
way is more suitable, we express our results as a percentage of the best MAP
achieved for the given concept. Table 1 and Figure 3 summarize the results of
Mu nSubmission100 expressed in this way.</p>
      <sec id="sec-4-1">
        <title>3 http://imageclef.org/2011/Photo</title>
        <p>Content element
Impression
Quality
Representation
Scene description</p>
        <p>Landscape elements
Pictured objects
Urban elements
Expressed impression
Felt impression
Aesthetics
Blurring
Art
Impression
Macro
Portrait
Still life
Abstract categories
Activity
Events
Place
Seasons
Time of day</p>
        <p>Table 1 shows the results averages in groups of semantically close categories
as speci ed by the ontology provided for ImageCLEF. We can observe that the
MUFIN approach is most successful in categories that are (1) related to visual
image content rather than higher semantics, and (2) probable to be re ected
in image tags. These are, in particular, the categories describing the depicted
elements, landscape, seasons, etc. Categories related to impressions, events, etc.
represent the other end of the spectrum; they are di cult to decide using only
the visual information and (especially the impressions) are rarely described via
tags.</p>
        <p>However, the average MAP values do not di er that much between categories.
The reason for this is revealed if we take a closer look at the results for individual
concepts, as depicted in Figure 3. Here we can notice low peaks in otherwise
well performing categories and vice versa. For instance, the clouds concept in
the landscapes category performs rather poorly. This is caused by the fact that
clouds appear in many images but only as a part of a background, which is
not important enough to appear in the annotation. On the contrary, airplanes
are more interesting and thus regularly appear in the annotations. In fact, we
again encounter the di erence between the annotation and classi cation tasks {
in annotation we are usually interested in the most important/interesting tags
while in classi cation all relevant tags are wanted.</p>
        <p>Several more extremes are pointed out in Figure 3. For instance, the concept
cute performs well because of its high frequency in the dataset. On the other
hand, for the concept overexposed a specialized classi er is much more suitable
than the annotation mining. The detailed discussion of the best tting methods
for individual categories is beyond the scope of this paper. However, we believe
that it is worth further studies to sort out the di erent types of concepts as well
as annotation approaches and try to establish some relationships between them.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>
        In this study, we have described the MUFIN solution of the ImageCLEF Photo
Annotation Task, which is based on free-text annotation mining, and compared
it to more specialized, classi er-based approaches. The method we presented has
its pros and cons. Mining information from annotated web collections is
complicated by a number of features related to the way the annotations are created.
As discussed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], we need to expect errors, typing mistakes, synonyms, etc.
However, there are also ways of overcoming these di culties. In our approach,
we have exploited a well-annotated collection, the semantical information
provided by WordNet, and a specialized ontology. Using these techniques, we have
been able to create an annotation system that shows precision comparable to
average classi ers, which are usually trained for speci c purposes only. The main
advantage of our solution lies in the fact that it requires minimum training (and
is therefore less dependent on the availability of high-quality training data) and
is scalable to any number of concepts.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by Brno PhD Talent Financial Aid and by
the national research projects GAP 103/10/0886 and VF 20102014004. The hardware
infrastructure was provided by the METACentrum under the programme LM 2010005.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Batko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucchese</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perego</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rabitti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sedmidubsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zezula</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Building a web-scale image similarity search system</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          <volume>47</volume>
          , 599{
          <fpage>629</fpage>
          (
          <year>2010</year>
          ), http://dx.doi.org/10.1007/s11042-009- 0339-z,
          <volume>10</volume>
          .1007/s11042-009-0339-z
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Budikova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zezula</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Evaluation platform for content-based image retrieval systems</article-title>
          . In: To appear
          <source>in Theory and Practice of Digital Libraries (TPDL</source>
          <year>2011</year>
          ) (
          <fpage>26</fpage>
          -
          <issue>28</issue>
          <year>September 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Budikova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zezula</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Online image annotation</article-title>
          .
          <source>In: 4th International Conference on Similarity Search and Applications (SISAP</source>
          <year>2011</year>
          ). pp.
          <volume>109</volume>
          {
          <issue>110</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Datta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.Z.</given-names>
          </string-name>
          :
          <article-title>Image retrieval: Ideas, in uences, and trends of the new age</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>40</volume>
          (
          <issue>2</issue>
          ) (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fellbaum</surname>
          </string-name>
          , C. (ed.):
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . The MIT Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A survey of methods for image annotation</article-title>
          .
          <source>J. Vis. Lang. Comput</source>
          .
          <volume>19</volume>
          (
          <issue>5</issue>
          ),
          <volume>617</volume>
          {
          <fpage>627</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kwasnicka</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paradowski</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Machine learning methods in automatic image annotation</article-title>
          .
          <source>In: Advances in Machine Learning II</source>
          , pp.
          <volume>387</volume>
          {
          <issue>411</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , 0001,
          <string-name>
            <given-names>L.Z.</given-names>
            ,
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Ma</surname>
          </string-name>
          , W.Y.:
          <article-title>Image annotation by large-scale content-based image retrieval</article-title>
          . In: Nahrstedt,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Turk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Rui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Klas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Mayer-Patel</surname>
          </string-name>
          ,
          <string-name>
            <surname>K</surname>
          </string-name>
          . (eds.) ACM Multimedia. pp.
          <volume>607</volume>
          {
          <fpage>610</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. MPEG-
          <article-title>7: Multimedia content description interfaces</article-title>
          .
          <source>Part</source>
          <volume>3</volume>
          : Visual. ISO/IEC 15938-3:
          <year>2002</year>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huiskes</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          :
          <article-title>New strategies for image annotation: Overview of the photo annotation task at imageclef 2010</article-title>
          . In: Braschler,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Harman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Pianta</surname>
          </string-name>
          , E. (eds.) CLEF (Notebook Papers/LABs/Workshops) (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nowak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liebetrau</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The CLEF 2011 Photo Annotation and Concept-based Retrieval Tasks</article-title>
          . In: CLEF 2011 working notes (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Trant</surname>
          </string-name>
          , J.:
          <article-title>Studying social tagging and folksonomy: A review and framework</article-title>
          .
          <source>J. Digit. Inf</source>
          .
          <volume>10</volume>
          (
          <issue>1</issue>
          ) (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>