<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Telecom Bretagne at ImageCLEF 2010 WikipediaMM</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Institut Telecom/Telecom Bretagne, Departement Informatique</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <abstract>
        <p>In this paper, I describe the approach proposed by Telecom Bretagne for the WikipediaMM 2010 evaluation campaign [6]. One of the main challenges in large scale image retrieval is the mismatch between query terms and image textual descriptions from the database. This mismatch can be reduced using query expansion and here I present a Wikipedia based query expansion approach. In order to boost results' accuracy, the expansion is followed by a reranking step which uses query models extracted from Flickr.</p>
      </abstract>
      <kwd-group>
        <kwd>Adrian Popescu</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Retrieving images from heterogeneous and noisy databases was thoroughly
studied but remains an interesting research topic. Some open questions include:
{ how to deal with di cult queries?
{ how to perform query expansion in a manner that improves both precision
and recall?
{ what resources to use in order to model query content?
In this paper, I present techniques which represent possible answers to these
questions. A query expansion technique is adapted from previous work, and
is augmented a query modeling module based on Flickr tags. The categorical
structure of Wikipedia is exploited in order to nd and rank concepts from the
encyclopedia which are semantically similar to the initial query. The experiments
are focused on English queries but the same techniques are applicable to other
languages.</p>
      <p>
        With over 3,000,000 articles in its English version, Wikipedia is a rich
resource and is used in a variety of research tasks, such as: sense disambiguation,
ontology extraction or semantic relatedness. The last problem can be
formulated as follows: given an input (a concept or a longer text), nd the concepts
which are most closely related to the input. Wikipedia based techniques to nd
semantic relatedness include WikiRelate! [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Explicit Semantic Analysis (ESA)
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and Wikipedia Link-based Measure (WLM) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. WikiRelate! modi es
techniques previously applied to WordNet in order to suit Wikipedia's structure.
The authors of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] map queries to Wikipedia concepts representation in order to
nd related concepts. ESA is interesting because it nds related concepts for any
given query and not only for mono-conceptual queries and is thus suited for use
in Web information retrieval. WLM exploits only Wikipedia links to nd related
concepts.A comparison of the three techniques [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] shows that ESA achieves the
best performances, followed by WLM and WikiRelate!. Modeling Flickr content
is another very interesting area of research. Related to the present paper are
techniques that analyze Flickr tags to nd frequent topics and their correlations
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Correlations are then used in order to suggest new tags based on
supplied tags. The authors of [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] also map tags into WordNet to extract the
distribution of Flickr tags in di erent conceptual domains. They report that main
tag categories include artifacts and places. Wu et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] analyze image content
in order to improve tag choice using textual correlations.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Query modeling</title>
      <p>The proposed retrieval model has two main components: query modeling with
Flickr, respectively with Wikipedia. These two components are described in the
following subsections subsections.
2.1</p>
      <sec id="sec-2-1">
        <title>Query modeling with Flickr</title>
        <p>Flickr is a photo sharing site that contains over 4.8 billion photos as of August
2010. A part of these photos are tagged and these tags can be used to build
query models. The 2010 WikipediaMM textual topics, and { more generally {
image queries, contain visual cues (terms speci c to photography such as close
up, black and white), which are not useful during the query modeling stage. A
list of photographic terms is extracted from http://en.wikipedia.org/wiki/
Category:Film_techniques and http://en.wikipedia.org/wiki/Category:
Photographic_processes (respectively the corresponding categories for French
and German) and these terms are removed from the queries. Prepositions are
also stripped from the queries and the remaining words are of each topic are
used to query the Flickr API and download metadata for up to top 20,000
images associated to the query. The relatedness is de ned by counting the photos
that are annotated with the respective term. In table 1, we present the top 10
related terms for fractals, tennis player on court and cactus in desert. Most of
the terms presented in the table are closely related to the initial query and they
constitute an acceptable model of its content. They range from generic terms
such as abstract, digital, nature for fractals to speci c terms such as wimbledon,
federer or centre court for tennis player on court.</p>
        <p>One known problem in information retrieval is that the word form in the
queries is not the same as their form in the database and this mismatch hurts
recall. One common solution to this problem is to use stemming but this solution
has its drawbacks since the stemmed forms of the words can match completely
di erent words, especially when dealing with multilingual datasets. Stemming
was only used for the words in the query and was followed by a look-up in the
query models in order to nd word variants. Related words that have an edit</p>
        <p>Top related terms
TopicTopic text
ID
1 fractals
8
39
fractal, apophysis, abstract, mandelbrot, romanesco, art, green,
digital, cauli ower, nature
tennis player wimbledon, tennis court, racket, players, sport, federer, ball,
on court wta, atp, centre court, tournament
cactus in arizona, cacti, saguaro, tucson, cholla, sonoran, phoenix,
califordesert nia, barrel cactus, az</p>
        <p>Table 1. Flickr query models for English. Top 10 related terms are presented.
distance smaller than 3 with respect to words in the query or terms that begin
with the stem of the terms in the query were considered relevant. For instance,
the topic cactus in desert becomes cactus:cacti desert and each word variant
will be used in order to search relevant results. Although the discussion here
is focused on English queries, the same procedure was applied to French and
German versions of the topics.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Query modeling with Wikipedia</title>
        <p>
          Words in the topic do not cover the entire semantic eld of the underlying
concepts and many potentially relevant results are ignored. One particularly useful
relation for improving the conceptual coverage of the topic is the conceptual
inheritance (X isA Y). For instance, Elena Dementieva or Rafael Nadal are tennis
players and images annotated with their names are potentially relevant the topic
tennis player on court. To discover semantically similar concepts, I rely on own
previous work [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], which exploits the categorical structure of Wikipedia.
        </p>
        <p>The main exploited resource is Wikipedia, which provides its dumps for free
use. I downloaded the English version of March 2010, split them into individual
les and search for articles which are related to the words in the WikipediaMM
topics. Prior to that, topics were preprocessed in order to remove photographic
terms and prepositions and to nd term variants in the query models (see
previous subsection). Also, topic words were run trough WordNet and synonyms
were extracted for unambiguous ones (words that have a single WordNet entry).
This extraction was limited to unambiguous terms because noise can be
introduced by secondary senses of polysemous words. The enriched versions of topics
were compared to Wikipedia categories in order to nd articles categorized with
words in the topic. The number of common terms between the topic and an
article's categories represents a rough measure of their semantic similarity. It
is de ned here as a score between 0 (no terms in common) and 1 (all terms in
common) and used to propose a rst ranking of Wikipedia articles ( rst column
in table 2). Since many articles share the same coarse grained similarity score,
ties are broke by introducing a second score which is directly dependent of the
number and frequency of terms from the Flickr query model that appear in the
target article and inversely proportional to the log of the length of the article.
This second score was determined empirically and results would be probably
improved if a more principled similarity distance was used.</p>
        <p>The examples in 2 show that extracted Wikipedia concepts are generally
closely related to the initial topics. Also, the example for topic 39 illustrates
well the importance of the coarse similarity score because the top 9 concepts
(similarity = 1) are related to both terms in the query; whereas the last term,
Wharram Percy, a deserted village in England, is only related to desert. The
Wikipedia translation graph was used to also get similar Wikipedia concepts for
French and German.</p>
        <p>Images in the WikipediaMM collection come with brief textual descriptions.
In this setting, query expansion is an appealing way to improve recall and, if
performed in a judicious way, to also improve results precision.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Retrieval experiments</title>
      <p>Both textual and multimodal runs were submitted this year. Multimodal runs
involved a k-NN inspired visual reranking of textual results and actually degraded
the nal quality of results. Also, runs were submitted for English only queries
and for all three versions of each topic. The multilingual scores of images were
formed by adding up scores in individual languages. As with multimodal runs,
the overall performance was degraded with respect to English only runs.
Therefore, in this section we focus on textual English only runs. Unlike past years'
WikipediaMM campaigns, the texts of the articles that contained collection
images were provided to participants. This thorough textual context of the images
facilitates the retrieval task. However, since the main purpose of Telecom was to
test a retrieval approach in a database with short and noisy textual descriptions,
submitted runs were based on the metadata les only and not on article texts.</p>
      <p>Given a query, image results from the Wikipedia collection are retrieved by
searching for images which are described either by terms in the initial query or
by related concepts from Wikipedia. We assume that the relatedness of an image
to a query is proportional to the number of words shared by its description in
the collection and the query (in its extended form). A score between 0 and 1
is attributed to the image, in function of the number of topic words found. For
instance, if the initial query contained three words (tennis player court) and
two of them were retrieved (tennis and player or tennis and court ), the image
gets a score of 0.667. If a Wikipedia related concept is retrieved in the image's
metadata, the corresponding coarse similarity score is attributed to the image,
with a small penalization to account for the rank of the Wikipedia concept. For
instance, an image annotated with Roger Federer gets a score of 0.667 while
another image annotated with Venus Williams gets a score of 0.666. Relevant
images are found by launching queries with the initial terms and the expanded
queries in the following order:
{ all terms in the initial query and a related concept
{ the initial query
{ parts of the initial query (starting with largest subparts - and favoring rare
terms) and a related concept
{ related concept or parts of the initial query</p>
      <p>One e ect of this type of scoring is that many images have the same ranking
score and they need to be separated. To break ties, Flickr query models are used
to search related terms in the image description. For two images are annotated
with Venus Williams, which both have an initial score of 0.666, an image also
annotated with wimbledon will be ranked higher than an image annotated with
wta because the rst word is more closely related to Venus Williams than the
second.</p>
      <p>To study the in uence of the use of Wikipedia related concepts, two runs
were submitted:
{ telecom en ickr wiki { involves searches with up to top 1000 Wikipedia
related concepts, regardless of their coarse similarity score.
{ telecom en ickr wiki maj { is limited to search with only those concepts
among the top 1000 Wikipedia related concepts that have a similarity score
higher than 0.5. For instance, in the case of cactus in desert, only the top
nine terms from table 2 are used. This limitation was imposed in order to
study the e ect of semantic similarity on the retrieval results.</p>
      <p>The MAP and precision (@10 and @20) results for the two analyzed runs
are presented in table 3. The run that uses only Wikipedia concepts with a high
similarity to the initial query has better scores for both types of measures. The
relative improvement of the mean average precision when limiting the usage of
related concepts is of 8.5%. This di erence supports the hypothesis that query
expansion should be performed only with terms that are closely related to the
initial query.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and future work</title>
      <p>Here a query modeling technique that involves the use of Wikipedia related
concepts but also of Flickr tags is introduced. The role of the Wikipedia concepts
is to improve recall when working with scarce annotations. Flickr models are used
to enrich initial queries with words that derivations of the words in the initial
query, to propose a ne grained ranking of Wikipedia concepts and to break ties
between initial scores provided after searching the database with terms in the
enriched version of the initial query and Wikipedia concepts.</p>
      <p>
        One interesting future work direction is to replace the ad-hoc concept ranking
schemes used in the experiments with more principled methods. Also necessary
for the validation of the approach is the comparison to standard retrieval models
which are already implemented in freely available indexing frameworks. Third,
it would be interesting to evaluate the usage of richer textual descriptions (i.e
Wikipedia articles that contain the images or textual windows that surround
the images in these articles). As noted, only image metadata were used and, as
underlined in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], these metadata provide only a scarce textual representation
of the images in the collection, which probably penalizes the quality of the nal
results.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Aknowledgement</title>
      <p>This research is part of the French National Agency for Research (ANR) project
Georama ( ANR-08-CORD- 009 ).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>E.</given-names>
            <surname>Gabrilovich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Markovich</surname>
          </string-name>
          .
          <article-title>Computing Semantic Relatedness using Wikipediabased Explicit Semantic Analysis</article-title>
          .
          <source>In Proc. of IJCAI</source>
          ,
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Milne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>An e ective, low-cost measure of semantic relatedness obtained from Wikipedia links</article-title>
          .
          <source>In Proc. of WIKIAI</source>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Le Borgne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.-A.</given-names>
            <surname>Mo</surname>
          </string-name>
          <article-title>ellic, Conceptual Image retrieval over a Large Scale Database</article-title>
          .
          <source>In Evaluating Systems for Multilingual and Multimodal Information Access</source>
          ,
          <source>In Proc. of the 9th Workshop of the Cross-Language Evaluation Forum, Lecture Notes in Computer Science</source>
          , 2009
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Sigurbjornsson</surname>
          </string-name>
          , B.,
          <string-name>
            <surname>van</surname>
            <given-names>Zwol</given-names>
          </string-name>
          , R..
          <source>Flickr Tag Recommendation based on Collective Knowledge. In Proc. of WWW</source>
          <year>2008</year>
          (Beijing, China).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Strube</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto. WikiRelate! Computing Semantic Relatedness Using</surname>
          </string-name>
          <article-title>Wikipedia</article-title>
          .
          <source>In Proc. of AAAI</source>
          ,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tsikrika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kludas</surname>
          </string-name>
          .
          <article-title>Overview of the WikipediaMM task at ImageCLEF 2010</article-title>
          . CLEF working notes,
          <year>2010</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>R. H. Van Leuken</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Olivares</surname>
          </string-name>
          , R. van Zwol.
          <article-title>Visual Diversi cation of Image Search Results</article-title>
          .
          <source>In Proc. of WWW</source>
          ,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hua</surname>
            ,
            <given-names>X.-S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
          </string-name>
          , W.-Y. and
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Flickr Distance</article-title>
          .
          <source>Proc. of ACM Multimedia</source>
          <year>2008</year>
          , Vancouver, Canada
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>