<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A multimedia IR-based system for the Photo Annotation Task at ImageCLEF2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>X. Benavent</string-name>
          <email>xaro.benavent@uv.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Castellanos</string-name>
          <email>acastellanos@lsi.uned.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. de Ves</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Hernández-Aranda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>R. Granados</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Garcia-Serrano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Nacional de Educación a Distancia</institution>
          ,
          <addr-line>UNED</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat de Valéncia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The UNED-UV group at the ImageCLEF2013 Campaign have participated in the Scalable Concept Image Annotation subtask. We present a multimedia IR-based system for the annotation task. In this collection, the images do not have any textual description associated, so we have downloaded and preprocessed the web pages which contain the images. Regarding the concepts, we expanded their textual description with additional information from external resources as Wikipedia or WordNet and we generate a KLD concept model using recovered textual information. The multimedia IR-based system uses a logistic relevance algorithm to get a model for each of the concepts to be trained using visual image features. Finally, the fusion subsystem merges textual and visual scores for a certain image to belong a concept, and decides the presence of the concept in the images.</p>
      </abstract>
      <kwd-group>
        <kwd>Text-Based Information Retrieval</kwd>
        <kwd>Content-Based Information Retrieval</kwd>
        <kwd>Multimedia Fusion</kwd>
        <kwd>Logistic regression relevance algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The UNED-UV is a research group with researchers from two universities in Spain,
the Universidad Nacional de Educación a Distancia (UNED) and the Valencia
University (UV). The group is working together since ImageCLEF08 edition.</p>
      <p>At this 2013 ImageCLEF Campaign [2], we participate at the Photo Annotation
and Retrieval Task [6] in the Scalable Concept Image Annotation subtask. The
motivation for this edition is focused on the development of image annotation systems that
address the scalability problem in such a way that the annotation systems have to be
able to adapt their behavior to take into account new concepts that can appear in the
images to be annotated. There were two datasets to evaluate the systems, one
containing the same concepts used for training (95 concepts) and a second one containing
these 95 concepts and 21 additional ones.</p>
      <p>As the classification-based systems, traditionally used for image annotation, are not
suitable for this task, we use a multimedia IR-based annotation methodology that
produces a concept model that predicts the probability that a certain concept belongs
to an image. A merging algorithm fuses textual and visual probabilities, and this final
score is used to decide the presence of a certain concept in an image.</p>
      <p>Section 2 describes the system overview and the annotation methodology used for
the two approaches submitted using the textual and the multimodal information. After
that, section 3 shows the submitted runs and section 4 analyze the results obtained.
Finally, in section 5 we extract conclusions and outlines possible future research lines.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System Description</title>
      <p>The global system (shown at Fig. 1) is divided into three main subsystems: TBIR
(Text-Based Image Retrieval), CBIR (Content-Based Image Retrieval) and the
Merging module. The TBIR subsystem is in charge of annotating the images using only
textual information selected from the web pages where they were downloaded.</p>
      <p>As the list of concepts does not include example images to train each one of the
concepts, the TBIR subsystem is in charge of generating a training set of images for
each concept for the Multimedia approaches. These images are taken from the so
called 3k collection images (Devel + Test). The CBIR subsystem generates a model
for each of the concepts with the generated training set images. These concept models
are used to generate the lists of relevant images for each concept. Finally, these lists
are combined with the fusion subsystem following a late fusion approach based on the
OWA operator [7].
2.1</p>
      <sec id="sec-2-1">
        <title>Annotation using textual information</title>
        <p>TBIR subsystem is in charge of the textual annotation of images in the collection. In
this collection, the images do not have any textual description associated, so the first
step is to obtain the textual information for describing them. For this task, we have
downloaded and pre-processed the web pages which contain the images. Regarding
the concepts, we expanded their textual description with additional information from
external resources as Wikipedia or WordNet.</p>
        <p>As an image may be annotated by several concepts, the annotation strategy is based
on an information retrieval approach, which indexes the concepts and uses each image
as query. The result of the retrieval process is a ranked concepts list for each image,
ordered by the textual similarity or score ( ).</p>
        <p>The modules of the TBIR subsystem have the following functionalities:
Pages
web urls
Collection
Images
Images</p>
        <p>Image</p>
        <p>Expansion
TBIR
CBIR</p>
        <p>Feature
Extraction</p>
        <p>Preprocess</p>
        <p>Image Text /
Concept</p>
        <p>Queries
(Image Text /</p>
        <p>Concept)
Concept Modeling
Logistic Regression
Relevance feedback
Concept Estimator
Logistic Regression</p>
        <p>Relevance feedback</p>
        <p>Image Expansion. This module is in charge of downloading the web pages that
contain the images in the collection. Then, we extract the textual information directly
related to each image taking into account the text contained in the following HTML
attributes: "title" and "alt" of &lt;img&gt; tag; and the &lt;a&gt; tag if the image is within a link.
An image may be contained in several web pages, so we recover the textual
information of every web page. Moreover, we take into account the image name that
appears in its URL.</p>
        <p>Concept Expansion. In order to obtain additional information to describe the
concepts, we use Wikipedia and WordNet as external resources, in the following way:

</p>
        <p>Wikipedia. We extract the textual information from fields &lt;text&gt; and &lt;categories&gt;
contained in the corresponding Wikipedia pages of the concepts.</p>
        <p>WordNet. Lexical and semantic information about the concepts is extracted:
definition, synonyms, hypernyms, hyponyms and related concepts.</p>
        <p>
          Additionally, we have modelled the raw text obtained from the Wikipedia concepts
description to identify a list with the most representative terms (so-called the
Wikipedia-KLD list). For this, we applied a divergence-based approach (Kullback Leibler
Divergence or KLD [4]) to identify not only the representative terminology but also
the terminology that better differentiate each concept from the rest. KLD weights each
term according to their occurrence in a given content and their occurrence in the rest
of the contents following the formulation in (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ):
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
where is the probability of each term within a document (frequency of
divided by the whole of terms in the document ) and is the probability of the
same term within the collection (frequency of divided by the number of terms in
the collection ).
        </p>
        <p>Pre-processing. Textual information is pre-processed: 1) deletion of characters
with no statistical meaning, like punctuation marks or blanks; 2) deletion of semantic
empty words in English language (stopwords), 3) reduction of words to their base
form by stemming, and 4) conversion of all words into lower case.</p>
        <p>Indexing. The indexing process is carried out using Lucene. The images are
indexed using only one field with the text associated to each image. The concepts are
indexed using three fields, depending on the information used for expansion:
Wikipedia, WordNET and Wikipedia_KLD.</p>
        <p>Searching. This module is in charge of launching the queries against a concrete
index in order to obtain the corresponding textual results (Txt Results). When using
images as queries, a concepts list will be obtained; and when the queries are the
concepts, an images list will be generated. The latter is used to fuse these textual results
with the visual ones obtained from the CBIR subsystem. The applied ranking function
is BM25 and its extension for structured documents BM25F [5], using the default
parameters.</p>
        <p>For the concept description several approaches were tested. Finally, we have
represented (and indexed) each concept by:
 Concept: The name of the concept.
 WP_Description: Contains the raw text of the concept Wikipedia page, plus the</p>
        <p>Wikipedia categories of the concept.
 WP_KLD_Description: As we have modelled each concept using the raw
Wikipedia text and the Wikipedia categories (WP_Desccription), the 50 most
representative terms, according KLD weighting, are indexed for each concept.
 WN_Description: The textual information obtained for the concept at WordNet is
the one included at the &lt;definition&gt;, &lt;forms&gt;, &lt;hypernyms&gt;,&lt;hyponyms&gt; and
&lt;related&gt; WordNet components.</p>
        <p>The different textual information describing the images is (store in the five fields):
 Img: The textual description initially associated to the image: img_title, img_alt,
img_link e img_name.
 Webpage: Includes the general description about the webpage containing the
image: webpage_title, webpage_description y webpage_keywords.
 img+webpage: The two previous fields together.
 text: The whole webpage text (text element) containing the image.
 img+webpage+text: The three previous fields together.</p>
        <p>Several experiments were performed in order to compare the use of the previous
alternatives for the textual description of the images when using as queries to retrieve
concepts from the concepts index of the Devel collection. The best result was
obtained using only the img field as a query; so this is the field to be used in the runs.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Annotation using Multimedia information</title>
        <p>For the Multimedia approaches, the TBIR subsystem generates a training set for each
concept to be annotated. These images are taken from the 3k collection images (Devel
+ Test). The CBIR subsystem generates a model or predictor for each of the concepts
with the Logistic Regression Relevance Model algorithm [3]. Once, the concepts
models are trained, these models are used to predict the probability that a given image
belongs to a certain concept (Si). Both probabilities (St,Si) are combined by the
Fusion subsystem that finally decides if a certain concept is present or not to be
annotated.</p>
        <p>For the Logistic Regression Relevance algorithm, each of the concept models
needs two sets to be trained: a set of images that have the concept, being the
relevant or positive images, and a set of images that not belongs to a concept, being ,
the set of non-relevant images or negative images. Each image is represented by a
Kdimensional low-level features vector , . . , , . . , . The relevance probability for
a certain concept for a given image will be represented as . A logistic
regression model can estimate these probabilities. Let us consider for a binary Y, and k
explanatory variables , … , , the model for π(x) = P(Y=1| X ) (probability
Y=1) for the x values ⋯ , where logit
(π(x))=ln(π(x) / (1-π(x)). The model parameters are obtained by maximizing the
likelihood estimator (MLE) of the parameter vector β by using an iterative method.</p>
        <p>The positive o relevant images (I set), is given by the Text-Based Information
sub-system using the 3k collection images. This initial list is tailored up to the tenth
top images, and the final selection is human supervised. The non-relevant images, the
I set, is selected from the images that do not have the required concept and this list is
also tailored up to the twentieth top images, being the selection also human
supervised. A good selection of the images that represent a certain concept is very
important to make the estimator good and robust. For this reason, we have considered
important the human supervision for the training sets. Furthermore, these sets are
generated only once for training, and could be used to annotate any other collection
with these concepts.</p>
        <p>The explanatory variables x x , … , x to train the model are the visual
lowlevel features based on colour and texture information that are given by the
organization [6]: colour histograms and GIFT shape descriptor that describes the shape in an
image by calculating the Gabor transform. We have a low-level features vector of 544
components: 64 for the colour histograms and 480 for the GIFT descriptor.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Multimedia Fusion</title>
        <p>To merge the textual and the visual information, we have followed a late fusion
approach which combines the two monomodal results lists at decision level. The applied
late fusion algorithm tested in previous works [1] is based on the Mathematical
aggregation operator OWA [7]. The OWA transforms a finite number of inputs into a
single output without associating weights to any particular input; instead, the relative
magnitude of the inputs decides which weight corresponds to each input. In our
application, the inputs are the textual and image scores ( and ), and this property is
very interesting because we do not know, a priori, which subsystem will provide us
the best information. The aggregation weights used for our experiments correspond to
an 0.3 , which means that a weight of 0.3 is given to the higher probability
value and a weight of 0.7 to the lower one.</p>
        <p>Once the final fused list is obtained (containing, for each image, a ranked list of
associated concepts), we have to decide how many concepts will be used to annotate
each image. We consider two options: 1) select a fixed number of concepts; and 2)
calculate a relevance threshold that decides whether or not an image is annotated by a
concept. Both options have been evaluated with the Devel collection, since it has
ground truth.</p>
        <p>Relative to the first option, Fig. 2 shows the evolution of evaluation measures
MAP, mFsamp (by image) and mFcnpt (by concept) depending on the number of
annotated concepts (values between 1 and 20). We can see how the MAP value is
higher as the number of concepts is increased, while the value of the rest of measures
decreases, being the cut-point between 5 and 6 . On the other hand, the
mean number of concept annotation per image is 6.345 (calculated from Devel ground
truth), so, taking into account both factors, we decide to select 7 as fixed number
of concepts to annotate images.
For the second option, the threshold calculation is based on the percentage of the
maximum score by image. This option is more flexible, since an image is not
annotated with a fixed number of concepts. Fig. 3 shows the evolution of the evaluation
measures considering the percentages from 10% to 100%. We can see that the more
restrictive is the threshold (low values of percentage), the more MAP value increases
and the rest of measures decrease.</p>
        <p>The cut-off point is between 70 and 80%. If the number of concepts with which an
image is annotated is considered, for 70% the mean is 6.622 while for 80% is 3.356.
Therefore, we have selected the percentage of 70% as a threshold, since its mean
concept number is similar to the mean calculated from Devel ground truth (6.345).
In this section, we present the two approaches developed for the image annotation
subtask: the monomodal approach (TBIR-based annotation) and the multimodal
approach (TBIR and CBIR based annotation). The five runs submitted are:
 UNEDUV_1: Monomodal Approach. Query with the img field against the</p>
        <p>KLD_WP concept representation indexed.
 UNEDUV_2: Monomodal approach. Query with the img field against the</p>
        <p>KLD_WP+WN concept representation indexed.
 UNEDUV_3: Monomodal approach. Query with the img field against the WN
concept representation indexed.
 UNEDUV_4: Multimodal approach. UNEDUV_2 textual run is merged with the
visual run by OWA algorithm, using the seventh most representative concepts to
annotate every image.
 UNEDUV_5: Multimodal approach. UNEDUV_2 textual run is merged with the
visual run by OWA algorithm, using the concepts with 70% of representative
concept score.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>UNED-UV Results</title>
      <p>In this section, we describe the results obtained with the submitted runs using the
development and the test collections, measured according the mean F-measure for
both the samples (MF-samples) and the concepts (MF-concepts); and the mean
average precision for the samples (MAP-samples). The test set results is also evaluated
with a fourth measure (MF-unseen), the mean F-measure for the concepts that are not
in the development set. The values between the square brackets correspond to the
95% confidence intervals computed using Wilson's method. Table 1 shows the results
for the development set, divided into textual and multimodal runs, and Table 2 shows
the same results for the test set. The best obtained value for each of the measures it is
also included in the table, together with the average over all the presented
experiments by the rest of participants.</p>
      <p>All our submitted runs are beyond the baseline results for both the development set
and for the test set according to all measures (see the overall results at the
ImageCLEF webpage [6]). Looking into the overall participant’s results list, our best
runs are at positions 21, 16 and 27 ordered by the MF-Samples, MF-Concepts and
MAP-samples respectively for the development set, and at positions 24, 11 and 26 for
the test set. It means that our best runs are at the first third top results.</p>
      <p>Focusing on textual runs, UNEDUV_1 and 2 offer similar results according to all
the measures; however, UNEDUV_3 results, which are based only in WordNET
annotation, differs: values based on sample results are fewer than values of the other two
approaches, and the value based on concepts (MF-concepts) improves these results
(4-5 points higher). This behavior is observed in the development and in the test set.</p>
      <p>On the other hand, the multimodal-based approach, UNEDUV_4 and UNEDUV_5
runs, improve the textual based run, the UNEDUV_2, in almost all measures, being
this improvement higher in the UNEDUV_5. This means that the merging strategy of
using a relevance threshold to decide if a concept should be or not annotated performs
better than the one that uses a static number of concepts to be annotated. Similar
behavior is observed for the two sets, development and test.</p>
      <p>An important issue to highlight for the test set results at Table 2 is the MF-unseen
values. In general all our submitted approaches offer satisfactory values; especially
the WordNET-based approach (UNEDUV_3) value, that is the 3rd best overall value,
and the Multimedia approach (UNEDUV_5) at the 5th position. It highlights a good
generalization capacity for our systems to annotate unseen concepts.
Our best runs are at the first third top results for the different measurements. This
means that the multimedia IR-based system presented has obtained quite good results
regarding the current state of the art.</p>
      <p>For the textual approaches, the WordNET-based approach for concept expansion is
the one that has a better performance. As this textual baseline was not the one used for
the submitted multimodal approaches, it is going to be tested in the multimodal
IRbased system.</p>
      <p>The multimedia approaches slightly outperform its textual baseline, although not in
all measurements, being this behaviour needed to be further analysed. The fusion
subsystem has proved that a relevance threshold to decide the annotation of a concept
achieves better results than selecting a fixed number of concepts per image. It is
important to highlight the good generalization capacity of our system to annotate unseen
concepts.</p>
      <p>
        Acknowledgments. This work has been partially supported for Regional Government
of Madrid under Research Network MA2VIRMR (S2009/TIC-1542), HOLOPEDIA
(TIN 2010-21128-C02) and by project MCYT TEC2009-12980.
3. Leon T., Zuccarello P., Ayala G., de Ves E., Domingo J.: Applying logistic regression to
relevance feedback in image retrieval systems, Pattern Recognition (40), pp. 2621, 2007.
4. S. Kullback R. A. Leibler. On information and sufficiency. Annals of Mathematical
Statistics, 22(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ). 1951.
5. Robertson, S. E. S. Walker. Some simple effective approximations to the 2-Poisson model
for probabilistic weighted retrieval. In Proceedings of the SIGIR '94, W. Bruce Croft and
C. J. van Rijsbergen (Eds.). Springer-Verlag. NY, USA, 232-241. 1994.
6. Villegas, M. Paredes, R. Thomee, B. Overview of the ImageCLEF 2013 Scalable Concept
      </p>
      <p>Image Annotation Subtask. CLEF 2013 working notes, Valencia, Spain, 2013.
7. Yager, R. On ordered weighted averaging aggregation operators in multi criteria decision
making. IEEE Transactions Systems Man and Cybernetics (18), pp. 183-190. 1988.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Benavent</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>García-Serrano</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Granados</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benavent</surname>
            , J. de Ves,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Multimedia Information Retrieval based on Late Semantic Fusion Approaches: Experiments on a Wikipedia Image Collection</article-title>
          .
          <source>IEEE Transactions on Multimedia. DOI: 10.1109/TMM</source>
          .
          <year>2013</year>
          .
          <volume>2267726</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Thomee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Paredes</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Zellhofer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Goeau</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Martinez Gomez</surname>
            ,
            <given-names>J. Garcia</given-names>
          </string-name>
          <string-name>
            <surname>Varea</surname>
            , I. Cazorla,
            <given-names>M.</given-names>
          </string-name>
          <article-title>ImageCLEF 2013: the vision, the data and the open challenges</article-title>
          .
          <source>Proc. CLEF</source>
          <year>2013</year>
          ,
          <article-title>LNCS</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>