<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UAIC participation at ImageCLEF 2012 Photo Annotation Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mihai Pîțu</string-name>
          <email>mihai.pitu@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniela Grijincu</string-name>
          <email>daniela.grijincu@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrian Iftene</string-name>
          <email>adiftene@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>UAIC: Faculty of Computer Science, “Alexandru Ioan Cuza” University</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the participation of our group in the ImageCLEF 2012 Photo Annotation Task. Our approach is based on visual and textual features as we experiment with different strategies in order to extract the semantics inside an image. First, we construct a textual dictionary of tags using the most frequent words present in the user tag annotated images from the training data sets. A linear kernel is then developed based on this dictionary. To gather more information from the images we further extract local and global visual features using TopSurf and Profile Entropy Features as well as Color Moments technique. We then aggregate these features with Support Vector Machines classification algorithm and train separate SVM models for each concept. In the end, to improve our system's performance, we add a postprocessing step that verifies the consistency of the predicted concepts and also applies a face detection algorithm in order to increase the recognition accuracy of the person related concepts. Our submission consists of one visual-only and four multi-modal runs. We further give a more detailed perspective of our system and discuss our results and conclusions.</p>
      </abstract>
      <kwd-group>
        <kwd>ImageCLEF</kwd>
        <kwd>Image classification</kwd>
        <kwd>Photo annotation</kwd>
        <kwd>SVMs</kwd>
        <kwd>TopSurf</kwd>
        <kwd>Bag-of-Words model</kwd>
        <kwd>kernel methods</kwd>
        <kwd>PEF</kwd>
        <kwd>Color moments</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        ImageCLEF 20121 Photo Annotation Task2 represents a competition that aims to
improve the state of the art of the Computer Vision field by addressing the problem of
automated image annotation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The participants are asked to create systems that can
automatically assign an image a subset of concepts from a list of 94 possible visual
concepts.
      </p>
      <p>In 2012, the organizers offered a database consisting of 15,000 training images
annotated with the corresponding 94 binary labels and a set of 10,000 test images
which were to be automatically annotated (see Figure 1). The images were extracted
from Flickr3 online photo sharing application and so each image had the associated
EXIF data and Flickr user tags. Among the reasons that make this task of image
annotation a difficult one are the diversity of the concepts simultaneously present in
an image, the subjectivity of the existing annotations in the training set (especially
regarding the feelings related concepts) and the fact that the training samples are
unbalanced and so there may be more examples for a concept then for another.</p>
      <p>
        The system we propose combines different state of the art image processing
techniques (TopSurf, PEF, Color Moments) with Support Vector Machines and
Kernel functions we defined in an attempt to obtain good overall performances. This
was our second participation in Photo Annotation task, after our contribution from
2009 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The rest of the article is organized as follows: Section 2 presents the visual and
textual features we extracted to describe the images, Section 3 covers the
classification and post processing modules of our system, Section 4 details our
submitted runs and Section 5 outlines our conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Visual and Textual Features</title>
      <sec id="sec-2-1">
        <title>2.1 Local Visual Features – TopSurf</title>
        <p>
          TopSurf4 [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] is a visual library that combines SURF interest points [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ] with visual
words based on a large pre-computed codebook [
          <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
          ] and returns the most important
visual information in the image (based on assigned Tf-Idf scores5). SURF interest
points and the associated descriptors provide (partial) invariance to affine
transformations of objects in images, but the number of interest points may vary
between 0 and a few thousands, depending of the size and details of a photo. Because
to every SURF interest point corresponds a descriptor (a 64 dimensional array), the
4 TopSurf: http://press.liacs.nl/researchdownloads/topsurf/
5 Tf-Idf: http://tfidf.com/
problem of matching such descriptors arises. As matching thousands of descriptors of
a given image against a large database is highly time consuming and practical
infeasible, TopSurf library assigns every SURF descriptor a visual word from the
precomputed codebook and associates a limited number (the most important) of such
visual words to the image. The time of the extraction process slightly increases
(experiments [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] shows that for SURF interest point extraction is required on average
0.37s and 0.07s for the assignment of the visual words), but matching TopSurf
descriptors improves the time complexity and quality of the overall process.
        </p>
        <p>
          The TopSurf library assigns Tf-Idf scores [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] to every visual word in the image and
returns the most important ones. In our system we use the cosine similarity to measure
the distance (angle) between two given images described by their corresponding
TopSurf descriptor:
1, 2 = cos
=
        </p>
        <p>1 ∗ 2
| 1| | 2|
=
∑
∑
1
1</p>
        <p>2
∑
2</p>
        <p>The similarity score will be between 0 and 1 (because the angle of the vectors d1
and d2 is smaller than 90 degrees), with 1 for identical descriptors and 0 for
absolutely different ones. The time needed to compare these descriptors is, on
average, 0.2 ms (with a database of 100,000 images).</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Profile Entropy Features</title>
        <p>
          Profile Entropy Features (PEF) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is a technique of extracting global visual features
which combines the texture characteristics with the shapes present in a given image
by computing the simple arithmetic mean in horizontal or vertical direction.
        </p>
        <p>
          The PEF features are computed on an image I by using the normalized RGB
channels: = , = , ! = 1 − − , where # = $&amp;$%. The profiles of the
orthogonal projections of the pixels to the horizontal X axis is noted ' ( and to the
)
vertical Y axis (' (), where op is the projection operator (arithmetic or harmonic
*
mean). The length of a profile is + = , or + = - , (where , denotes I’s
columns and - , denotes I’s rows) and we estimate its probability distribution
function (( . ) on / = 01 √+ bins [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Then for each channel and operator,
we compute: Ф )( , = ( . ' )( and we set PEF components to the normalized
entropy of this distribution:
        </p>
        <p>The algorithm repeats for each of the 3 equal horizontal sub-images (see Figure 2)
and on the whole image. The PEF descriptor is denoted by the concatenation of 4567,
456?, 456% the mean and variance of the 3 channels, thus we have 4 regions × 5
features × 3 channels = 60 dimensions that describe the image I.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3 Color moments</title>
        <p>
          Color moments represent a method that can be used to differentiate images based on
their features of color. The main idea behind color moments is the assumption that the
distribution of color can be interpreted as a probability distribution, which can be
characterized by a number of moments (mean, variance, etc.). Stricker and Orengo
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] used three central moments of an image’s color distribution: mean, standard
deviation and skewness. The same authors showed that traditional methods like color
histograms are sensitive to minor modifications in illumination or affine
transformations.
        </p>
        <p>A color can be abstractly represented by using color models like RGB (Red, Green,
and Blue) or HSV (Hue, Saturation and Value). Thus, each of the three dimensions of
the chosen color model is characterized by three moments of a color distribution,
resulting in a nine dimension vector which will describe the color distribution in a
given image.</p>
        <p>5 =
∑B
( B,
5 is the mean or the average color value in the image, ( B is the value of the jth
pixel in the ith dimension of the color model and N is the number of pixels in the
image.</p>
        <p>C = D
∑B
C is the standard deviation (the square root of the variance) of the distribution.</p>
        <p>E
= D
∑B</p>
        <p>( B − 5 &amp;,
is the skewness of the distribution which is a measure of the degree of its
asymmetry.</p>
        <p>The similarity function F9F can be used to adjust the weights (G ) of each
channel, because it makes sense that, for example, the hue of a color is more
important than its intensity. The function is defined as the sum of the weighted
differences between the moments of the two distributions:</p>
        <p>F9F , , ,
∑&amp;</p>
        <p>G |5 " 5 | H G |C " C | H G |
"
|</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4 Using Flickr user tags</title>
        <p>In some situations, the visual information is not enough to give a semantic
interpretation of an image and this is why we exploit user defined tags to improve the
judgment of the whole system. The problems that arise with these approaches are the
fact the number of user defined tags is relatively small (or 0), the tag can be in any
language, some of them are irrelevant or they are a concatenation of words (see Figure
3). These problems make the traditional methods used in the field of natural language
processing inapplicable in this situation.</p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] propose a linear SVM kernel that uses the most frequent user
tags from the training set, which proved to be a good method. The idea is to construct
a dictionary with user tags that appear at least k times (in our system we used k = 16)
in the associated images from the training set. This process eliminates irrelevant and
rare user tags and limits the dictionary to a number of n tags. Prior to the construction
of the dictionary, we used Bing Translator6 on every associated user tag, in order to
attempt translation in English and a stemming algorithm that will reduce inflected or
6 Bing Translator: http://www.bing.com/translator/
derived words to their root. After the dictionary is computed, an n-dimensional binary
vector will be assigned to each image, with the ith component 1 if the image is
annotated with the ith user tag from the dictionary and 0, otherwise. The linear SVM
kernel that classifies these vectors is:
        </p>
        <p>I JK , KBL = K MKB</p>
        <p>K MKB is the dot product between the transposed binary vector K and the KB vector.
The KG kernel counts the number of shared user tags between two associated images.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Classification</title>
      <sec id="sec-3-1">
        <title>3.1 Classification using SVMs</title>
        <p>
          Support vector machines [
          <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
          ] proved to be one of the best classification technique
used to address image classification problems as it can be very flexible and work with
large amounts of data. Because this task requires multi-label classification (an image
can be annotated with more than one concept), we choose to train an SVM classifier
for each of the 94 concepts proposed by the ImageCLEF organizers [15] (to train a
classifier for a concept c, we choose as positive examples the images that are
annotated with the c concept and as negative examples the rest of the training
images). Also, because of the highly unbalanced classification problem (the positive
examples are usually less than the negative examples), we implemented a sampling
method [16].
        </p>
        <p>We propose a combined SVM kernel that makes use of all the features described
above:</p>
        <p>IN9FO PQ@ R, S = TUVIUV R, S + T:QAI:QA R, S + TWUIWU R, S + TNFINF R, S</p>
        <p>
          Where TUV, T:QA, TWU, TNF ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ], (such
weights for the following kernel functions:
•
that TUV + T:QA + TWU + TNF = 1)
are
•
•
•
        </p>
        <p>IUV R, S = UV R , UV S is the cosine similarity defined in
section 2.1 for the TopSurf library;
I:QA R, S = exp −_||R − S|| ) is the RBF kernel and it is used with PEF
descriptors (section 2.2);
IWU(R, S) = U(`)aU(b) is the linear kernel defined in section 2.4 normalized by</p>
        <p>P
the number of tags in the dictionary;
INF(R, S) = exp(−_ F9F(R, S)) is the kernel that uses F9F function for
color moments (section 2.3) and _ is the regularization parameter.</p>
        <p>
          These functions and Kd=efghij kernel satisfy Mercer’s theorem [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] necessary to
ensure SVMs convergence.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Post processing</title>
        <p>In the post processing module of our system we ensure that the classifications made
by SVMs models are correct. For example, if an image is classified with
quality_noblur and quality_partialblur at the same time, we adjust the concept’s
probabilities so they sum up to 1. We learn about mutual exclusive concepts ( 5R )
from the training set. Let +k be the set of predictions made by SVMs, with T , T ∈
5R and T , T ∈ +k:
+k = +k \ mT ∶ (T , T ) ∈
5R, ( (T ) o ( (T )p</p>
        <p>We also compute the Voila – Jones face detection algorithm [17], in order to count
the number of persons in a given image (the concepts regarding the number of persons
in this year’s competition are: quantity_none, quantity_one, quantity_two,
quantity_three, quantity_smallgroup, quantity_largegroup) and to determine if
view_portrait concept is present.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Submitted runs and results</title>
      <p>Our system (Figure 4) has a modular and flexible structure and can easily be extended
with some other feature extractors’ algorithms:
We participated at this year ImageCLEF 2012, Photo Annotation Task by submitting
5 runs with different configurations:
• Submission1: Visual only configuration with the following parameters:
TUV = 0.6, T:QA = 0.2, TWU = 0.0, TNF = 0.2 and SVM’s regularization
parameter: C = 20 with post processing step;
• Submission2: Multimodal run with the parameters: TUV = 0.45, T:QA = 0.1,
TWU = 0.35, TNF = 0.1 and SVM’s regularization parameter C was chosen
separately for each of the 94 classifiers, with sampling for some of the
concepts;
• Submission3: The same configuration as for Submission2, with the sampling
strategy applied for each of the 94 classifiers;
• Submission4: Multimodal run with the parameters: TUV = 0.35, T:QA = 0.25,
TWU = 0.25, TNF = 0.15 and SVM’s regularization parameter C was chosen
separately for each of the 94 classifiers, with the sampling strategy applied
for each of the 94 classifiers;
• Submission5: Multimodal run with the parameters: TUV = 0.45, T:QA = 0.1,
TWU = 0.35, TNF = 0.1 with SVM’s regularization parameter C = 20 and
without the post processing step.</p>
      <p>
        Our best run was the one with Submission1 configuration and it was ranked 11th of
a total of 18 group participants [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The fact that our visual-only run achieved the best
of our scores shows that local invariant visual features are more appropriate for this
task than other type of features. Also, we noticed that using user tags for classifying
some of the concepts is, in fact, misleading. For example, for concept
weather_cloudysky the most frequent tags were: blue, Cannon, Nikon, clouds.
      </p>
      <p>All participants at ImageCLEF 2012 in Photo Annotation task have submitted several runs
using not only visual strategies based on features extracted from the images but as well textual
ones based on user defined tags that were given alongside the images. The best results however,
as it can also be observed from the table above, were achieved by the systems that managed to
combine both the visual and textual features together. What our system lacked was the fact that
we did not find the best balance between feature extraction algorithms (with their contribution
in the learning step) and also the fact that some of them should weight more or less depending
on the concept that is being learned.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Conclusions</title>
      <p>In this paper we combined several different state of the art algorithms for image
processing together with Support Vector Machines and kernel functions in order to
approach the task of automated image annotation. As images can be annotated with
more than one concept we tried to increase our system’s performance by using not
only local image feature descriptors (TopSurf), that for example, proved to be
unpractical at detecting feelings in an image, but also try analyzing the colors (Color
Moments) and the textures (Profile Entropy Features) in the image and even make use
of the user defined tag semantics and face detection algorithms.</p>
      <p>All experiments were made using the approach we presented in this paper and
careful attention was given to the selection of the threshold parameters of the SVM
kernel function that we used, IN9FO PQ@, TUV, T:QA, TWU and TNF.</p>
      <p>As future work, we will try and set different values for these parameters taking into
consideration the concept that the classifier is training for. For example, for concepts
that express feelings, Color Moments technique should have the deciding weight,
whereas for panoramic images a greater weight should be given to the texture
descriptor (PEF).</p>
      <p>Acknowledgement. The research presented in this paper was funded by the Sector
Operational Program for Human Resources Development through the project
“Development of the innovation capacity and increasing of the research impact
through post-doctoral programs” POSDRU/89/1.5/S/49944.
15. Gidudu, A., Hulley, G., Marwala, T.: Image Classification Using SVMs: One-against-One
Vs. One-against-All. In Proceeding of the 28th Asian Conference on Remote Sensing,
Malaysia, CD-Rom. (2007)
16. Witten, I., Frank, E., Hall, M.: Data Mining – Practical Machine Learning Tools and</p>
      <p>Techniques, Third Edition. Morgan Kaufmann, 629 pages. (2011)
17. Viola, P., Jones, M.: Rapid object detection using boosted cascade of simple features.</p>
      <p>Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), vol.
1, pp. 511–518, Kauai, Hawaii, USA. (2001)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Overview of the ImageCLEF 2012 Flickr Photo Annotation and Retrieval Task</article-title>
          .
          <source>CLEF 2012 working notes</source>
          , Rome, Italy. (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Iftene</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vamanu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croitoru</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <string-name>
            <surname>UAIC at ImageCLEF 2009 Photo Annotation</surname>
          </string-name>
          <article-title>Task</article-title>
          . In C. Peters et al. (Eds.):
          <source>CLEF</source>
          <year>2009</year>
          , LNCS 6242,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          (
          <article-title>Multilingual Information Access Evaluation Vol. II Multimedia Experiments)</article-title>
          .
          <source>Pp</source>
          .
          <volume>283</volume>
          -
          <fpage>286</fpage>
          . Springer, Heidelberg. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakker</surname>
            ,
            <given-names>E. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lew</surname>
            ,
            <given-names>M. S.:</given-names>
          </string-name>
          <article-title>TOP-SURF: a visual words toolkit</article-title>
          .
          <source>In Proceedings of the 18th ACM International Conference on Multimedia</source>
          , pp.
          <fpage>1473</fpage>
          -
          <lpage>1476</lpage>
          , Firenze,
          <string-name>
            <surname>Italy.</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Notes on the OpenSURF Library</article-title>
          . CSTR-
          <volume>09</volume>
          -001, University of Bristol.
          <source>January</source>
          <year>2009</year>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bay</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ess</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuytelaars</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Gool</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Speeded-up robust features (SURF)</article-title>
          .
          <source>Computer Vision</source>
          and Image Understanding.
          <source>Computer Vision and Image Understanding (CVIU)</source>
          , Vol.
          <volume>110</volume>
          , No.
          <issue>3</issue>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>359</lpage>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Csurka</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dance</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Willamowski</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bray</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Visual Categorization with Bags of Keypoints</article-title>
          . In Workshop on Statistical Learning in
          <source>Computer Vision</source>
          , ECCV. (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. van Gemert,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Snoek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Veenman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Smeulders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Geusebroek</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. M.</surname>
          </string-name>
          :
          <article-title>Comparing Compact Codebooks for Visual Categorization</article-title>
          .
          <source>Computer Vision and Image Understanding</source>
          , Vol.
          <volume>14</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>4</given-names>
          </string-name>
          , Pp.
          <fpage>450</fpage>
          -
          <lpage>462</lpage>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGill</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Introduction to modern information retrieval</article-title>
          .
          <source>McGraw-Hill</source>
          .
          <article-title>(</article-title>
          <year>1983</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ayache</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Efficient image concept indexing by harmonic and arithmetic profiles entropy</article-title>
          .
          <source>In: Proceedings of 2009 IEEE International Conference on Image Processing (ICIP</source>
          <year>2009</year>
          ), Cairo, Egypt, November 7-
          <issue>11</issue>
          ,
          <fpage>14</fpage>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Moddemeijer</surname>
          </string-name>
          , R.:
          <article-title>On estimation of entropy and mutual information of continuous distributions</article-title>
          .
          <source>Signal Processing</source>
          , vol.
          <volume>16</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>233</fpage>
          -
          <lpage>246</lpage>
          ,
          <year>March 1989</year>
          . (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Stricker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orengo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Similarity of color images. In Storage and Retrieval for Image and Video Databases</article-title>
          ,
          <source>Proc. SPIE 2420</source>
          , pp.
          <fpage>381</fpage>
          -
          <lpage>392</lpage>
          . (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Guillaumin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verbeek</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Multimodal semi-supervised learning for image classification</article-title>
          .
          <source>IEEE Conference on Computer Vision &amp; Pattern Recognition</source>
          . pp.
          <fpage>902</fpage>
          -
          <lpage>909</lpage>
          . Grenoble, France. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Cristianini</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          Shawe-Taylor.
          <article-title>: An Introduction to Support Vector Machines and other kernel-based learning methods</article-title>
          . Cambridge University Press, Cambridge, UK. (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tong</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Support vector machine active learning for image retrieval</article-title>
          .
          <source>In: Proc. ACM Multimedia</source>
          , Ottawa, Canada. (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>