<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DUTH at ImageCLEF 2011 Wikipedia Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Avi Arampatzis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Konstantinos Zagoris</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Savvas A. Chatzichristofis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical and Computer Engineering, Democritus University of Thrace</institution>
          ,
          <addr-line>Xanthi 67100</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1837</year>
      </pub-date>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>As digital information is increasingly becoming multimodal, the days of single-language
text-only retrieval are numbered. Take as an example Wikipedia where a single topic
may be covered in several languages and include non-textual media such as image,
audio, and video. Moreover, non-textual media may be annotated with text in several
languages in a variety of metadata fields such as object caption, description, comment,
and filename. Current search engines usually focus on limited numbers of modalities at
a time, e.g. English text queries on English text or maybe on textual annotations of other
media as well, not making use of all information available. Final rankings are usually
results of fusion of individual modalities, a task which is tricky at best especially when
noisy modalities are involved.</p>
      <p>In this paper we present the experiments performed by Democritus University of
Thrace (DUTH), Greece, in the context of our participation to the ImageCLEF 2011
Wikipedia Retrieval task.1 The ImageCLEF 2011 Wikipedia collection is the same as
in 2010. It has image as its primary medium, consisting of 237; 434 items, associated
with noisy and incomplete user-supplied textual annotations and the Wikipedia articles
containing the images. Associated annotations are written in any combination of
English, German, French, or any other unidentified language. This year there are 50 new
test topics, each one consisting of a textual and a visual part: three title fields (one per
language—English, German, French), and 4 or 5 example images. The exact details of
the setting of the task, e.g., research objectives, collection etc., are provided at the task’s
webpage.</p>
      <p>
        We kept building upon and improving the experimental multimodal search engine
we introduced last year, www.mmretrieval.net (Fig.1). The engine allows
multiple image and multilingual queries in a single search and makes use of the total
available information in a multimodal collection. All modalities are indexed separately and
searched in parallel, and results can be fused with different methods. The engine
demonstrates the feasibility of the proposed architecture and methods, and furthermore enables
a visual inspection of the results beyond the standard TREC-style evaluation. Using the
engine, we experimented with different score normalization and combination methods
for fusing results. We eliminated the least effective methods based on our last year’s
participation to ImageCLEF [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and improved upon whatever worked best.
      </p>
      <sec id="sec-1-1">
        <title>1 http://www.imageclef.org/2011/Wikipedia</title>
        <p>The rest of the paper is organized as follows. In Section 2 we describe the
MMretrieval engine, give the details on how the Wikipedia collection is indexed and a brief
overview of the search methods that the engine provides. In Section 3 we describe in
more detail the fusion methods we experimented with and justify their use. A
comparative evaluation of the methods is provided in Section 4; we used the 2010 topics for
tuning. Experiments with the 2011 topics are summarized in Section 5. Conclusions are
drawn in Section 6.</p>
        <p>
          www.MMRetrieval.net: A Multimodal Search Engine
During last year’s ImageCLEF Wikipedia Retrieval, we introduced an experimental
search engine for multilingual and multimedia information, employing a holistic web
interface and enabling the use of highly distributed indices [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Modalities are searched
in parallel, and results can be fused via several selectable methods. This year, we built
upon the same engine eliminating the least effective methods and trying to improve
whatever worked best last year.
        </p>
        <sec id="sec-1-1-1">
          <title>2.1 Indexing</title>
          <p>
            To index images, we employ the family of descriptors known as Compact Composite
Descriptors (CCDs). CCDs consist of more than one visual features in a compact vector,
and each descriptor is intended for a specific type of image. We index with two
descriptors from the family, which we consider them as capturing orthogonal information
content, i.e., the Joint Composite Descriptor (JCD) [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] and the recently proposed Spatial
Color Distribution (SpCD) [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]. JCD is developed for color natural images, while SpCD
is considered suitable for colored graphics and artifficially generated images. Thus, we
have 2 image indices.
          </p>
          <p>The collection of images at hand, i.e. the ImageCLEF 2010/2011 Wikipedia
collection, comes with XML metadata consisting of a description, a comment, and multiple
captions, per language (English, German, and French). Each caption is linked to the
wikipedia article where the image appears in. Additionally, a raw comment is supplied
which may contain some of the per-language comments and any other comment in an
unidentified language. Any of the above fields may be empty or noisy. Furthermore, a
name field is supplied per image containing its filename. We do not use the supplied
&lt;license&gt; field.</p>
          <p>For text indexing and retrieval, we employ the Lemur Toolkit V4.11 and Indri V2.11
with the tf.idf retrieval model.2 In order to have clean global (DF) and local statistics
(TF, document length), we split the metadata and articles per language and index them
separately. Thus, we have 4 indices: one per language which includes metadata and
articles together but allows limiting searches in either of them, plus one for the
unidentified language metadata including the name field (which can be in any language). For
English text, we enable Krovetz stemming; no stemming is done for other languages
in the current version of the system. We also Krovetz-stem the unidentified language
metadata, assuming that most of it is probably English.
2.2</p>
        </sec>
        <sec id="sec-1-1-2">
          <title>Searching</title>
          <p>The web application is developed in the C#/.NET Framework 4.0 and requires a fairly
modern browser as the underlying technologies which are employed for the interface
are HTML, CSS and JavaScript (AJAX). Fig.2 illustrates an overview of the
architecture. The user provides image and text queries through the web interface which are
dispatched in parallel to the associated databases. Retrieval results are obtained from
each of the databases, fused into a single listing, and presented to the user.</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>2 http://www.lemurproject.org</title>
        <p>Users can supply no, single, or multiple query images in a single search, resulting
in 2 i active image modalities, where i is the number of query images. Similarly, users
can supply no text query or queries in any combination of the 3 languages, resulting in
3 l active text modalities, where l is the number query languages. Each supplied query
results to 3 modalities: it is run against the corresponding language metadata, articles,
as well as, the unidentified language metadata. The current alpha version assumes that
the user provides multilingual queries for a single search, while operationally query
translation may be done automatically.</p>
        <p>The results from each modality are fused by one of the supported methods. Fusion
consists of two components: score normalization and combination. In CombSUM, the
user may select a weigh factor W 2 [0; 100] which determines the percentage
contribution of the image modalities against the textual ones.</p>
        <p>For efficiency reasons, only the top-2500 results are retrieved from each modality. If
a modality returns less than 2500 items, all non-returned items are assigned zero scores
for the modality. When a modality returns 2500 items, all non-occurring items in the
top-2500 are assigned half the score of the 2500th item.
3</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Fusion</title>
      <p>Let i = 1; 2; : : : be the index running over example images, and j running over the
visual descriptors (only two in our setup), i.e. j 2 f1; 2g. Let DESCji be the score of a
collection image against the ith example image for the jth descriptor.
3.1</p>
      <sec id="sec-2-1">
        <title>Score Combination</title>
      </sec>
      <sec id="sec-2-2">
        <title>CombSUM</title>
        <p>The parameter w controls the relative contribution of the two media; for w = 1 retrieval
is based only on text while for w = 0 is based only on image.</p>
      </sec>
      <sec id="sec-2-3">
        <title>CombDUTH</title>
        <p>Image Modalities Assuming that the descriptors capture orthogonal information, we
add their scores per example image. Then, to take into account all example images,
the natural combination is to assign to each collection image the maximum similarity
seen from its comparisons to all example images; this can be interpreted as looking for
images similar to any of the example images. Summarizing, the score s for a collection
image against the topic is defined as:</p>
        <p>0
s = max @X DESCjiA</p>
        <p>i j</p>
        <p>Let l 2 f1; 2; 3g be the index running over provided natural languages (or example
text queries, i.e. three in our setup), and m 2 f1; 2; 3g running over the textual data
streams per language (we consider three: metadata, articles, and undefined language
metadata). Let TEXTml be the score of a collection item against the text query in the
lth language for the mth text stream.</p>
        <p>Fusion consists of two successive steps: score normalization and score combination.
1
!
(1)
(2)
(3)
(4)
Text Modalities Assuming that the text streams capture orthogonal information, we
add their scores per language. Then, to take into account all the languages, the natural
combination is to assign to each collection item the maximum similarity seen from its
comparisons to all text queries; this can be interpreted as looking for items in any of the
languages. Summarizing, the score s for a collection image against the topic is defined
as:
s = max
l</p>
        <p>X TEXTml
m
Combining Media Incorporating text, again as an orthogonal modality, we add its
contribution. Summarizing, the score s for a collection image against the topic is defined
as:
s = (1
0</p>
        <p>1
w) miax @ 1j Xj DESCjiA + w mlax</p>
        <p>
          Query Difficulty Inverse document frequency (IDF) is a widely used and robust term
weighting function capturing term specificity [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Analogously, query specificity (QS)
or query IDF can be seen as a measure of the discriminative power of a query over a
collection of documents. A query’s IDF is a log estimate of the inverse probability that
a random document from a collection of N documents would contain all query terms,
assuming that terms occur independently. QS is a good pre-retrieval predictor for query
performance [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For a query with k terms 1; : : : k, QS is defined as
        </p>
        <p>QSk = log</p>
        <p>Yk N !
i=1 dfi
= Xk log N
i=1
dfi
(5)
where dfi is the document frequency (DF), i.e. the number of collection documents in
which the term i occurs.</p>
        <p>In the Query Difficulty (QD) normalization, we divide all scores per modality by
QS, using the df statistics corresponding to the modality. This will promote the scores
of ‘easy’ modalities and demote the scores of ‘difficult’ modalities for the query.</p>
        <p>For image modalities, we do a similar normalization as defined in the above
equation, except that the k terms are replaced by each descriptor’s bins.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments with the 2010 Topics</title>
      <sec id="sec-3-1">
        <title>MinMax+CompSUM</title>
        <p>Tables 5, 6, 7, and 8 summarize the MinMax+CompDUTH results.
Per-modalitytype is the weakest MinMax normalization, followed by per-query-language. Best early
precision is achieved by per-modality (best P10) at w = 0:5 and per-index-language
(best P20) at w = 0:6. Per-modality at w = 0:7 achieves the best MAP, while
perindex-language achieves the best bpref at w = 0:7. Although per-index-language has
lower MAP than per-modality, its MAP comparable to per-modality; moreover,
perindex-language achieves a higher bpref which signals that we may be retrieving
unjudged relevant items. All in all, we conclude that per-index-language is the strongest
MinMax normalization.
4.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Overall Comparison of CompSUM, CompDUTH, and MinMax Types</title>
        <p>Overall, best early precision is achieved by per-index-language MinMax with
CompSUM at w = 0:7, and all other measures are optimized by per-modality MinMax with
CompSUM at w = 0:8. However, since the 2011 topic set consists of 4 or 5 example
images per topic, CompDUTH may show larger effectiveness differences than these on
the 2010 topic set; consequently, we will retain CompDUTH runs with 2011 topic set,
using per-index-language MinMax and w = 0:6; 0:7; 0:8. All these will result to 5 runs
in total.
4.4</p>
        <p>QD Normalization</p>
        <p>Tables 9 and 10 summarize the QD normalization results with both combination
methods. In early precision, the QD normalization works much better with CompSUM
than with CompDUTH. The best CompSUM results are achieved for w = 0:4; this run
has also the best P10 we have reported so far. In all other measures, although CompSUM
is slightly better than CompDUTH, their effectiveness is comparable.</p>
        <p>In comparison to the MinMax normalizations, the QD normalization achieves the
best initial precision results (when CompSUM is used for combination), and
comparable effectiveness to the best MinMax normalization in all other measures.</p>
        <p>In summary, we will retain QD+CompSUM at w = 0:4 and QD+CompDUTH at
w = 0:3 and 0:5; thus, we will have 3 QD runs in total.
4.5</p>
      </sec>
      <sec id="sec-3-3">
        <title>Summary</title>
        <p>While we have experimented with radically different normalization and combination
methods, our results have not shown a large variance. This suggests that we are
‘pushing’ at the effectiveness ceiling of the 2010 dataset. It is worth noting that most of
the runs reported so far have a better MAP and bpref than last year’s best automatic
run submitted to ImageCLEF, and a slightly lower but comparable initial precision.3
Nevertheless, a visual inspection of our results reveals that with CompDUTH we are
retrieving un-judged items which are sometimes relevant, a fact that most of the times
does not seem to get picked up by bpref.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments with the 2011 Topics</title>
      <p>
        3 Last year’s best MAP, P10, P20, and bpref were 0.2765, 0.6114, 0.5407, and 0.3137,
respectively; they were all achieved by the XRCE group [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We reported our experiences and research conducted in the context of our
participation to the controlled experiment of the ImageCLEF 2010 Wikipedia Retrieval task. As
second-time participants, we improved upon and extended our experimental search
engine, http://www.mmretrieval.net, which combines multilingual and
multiimage search via a holistic web interface and employs highly distributed indices.
Modalities are search in parallel, and results can be fused via several methods.</p>
      <p>All in all, we are modestly satisfied with our results. Although our best MAP run
ranked our system as the second-best among the other participants’ systems (excluding
all relevance feedback and query expansions runs), we believe that the content-based
image retrieval part of the problem has a large room for improvement. A promising
direction may be using new image modalities such as those based on the
bag-of-visualwords paradigm and other similar approaches. Furthermore, we consider score
normalization and combination important problems; while effective methods exist in
traditional text retrieval, those problems are not trivial in multimedia setups.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arampatzis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zagoris</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Multimedia search with noisy modalities: Fusion and multistage retrieval</article-title>
          .
          <source>In: Braschler et al. [2]</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Braschler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pianta</surname>
          </string-name>
          , E. (eds.):
          <article-title>CLEF 2010 LABs and Workshops</article-title>
          , Notebook Papers,
          <fpage>22</fpage>
          -
          <lpage>23</lpage>
          September 2010, Padua, Italy (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boutalis</surname>
            ,
            <given-names>Y.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Selection of the proper compact composite descriptor for improving content-based image retrieval</article-title>
          .
          <source>In: SPPRA</source>
          . pp.
          <fpage>134</fpage>
          -
          <lpage>140</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boutalis</surname>
            ,
            <given-names>Y.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <string-name>
            <surname>SpCD - Spatial Color Distribution Descriptor -</surname>
          </string-name>
          <article-title>A fuzzy rule-based compact composite descriptor appropriate for hand drawn color sketches retrieval</article-title>
          .
          <source>In: Proceedings ICAART</source>
          . pp.
          <fpage>58</fpage>
          -
          <lpage>63</lpage>
          . INSTICC Press (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Clinchant</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Csurka</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ah-Pine</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacquet</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perronnin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Sa</given-names>
            ´nchez, J.,
            <surname>Minoukadeh</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Xrce's participation in wikipedia retrieval, medical image modality classification and ad-hoc retrieval tasks of imageclef 2010</article-title>
          . In: Braschler et al. [
          <volume>2</volume>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cronen-Townsend</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Croft</surname>
          </string-name>
          , W.B.:
          <article-title>Predicting query performance</article-title>
          .
          <source>In: SIGIR</source>
          . pp.
          <fpage>299</fpage>
          -
          <lpage>306</lpage>
          . ACM (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Spa¨rck Jones,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>A statistical interpretation of term specificity and its application in retrieval</article-title>
          .
          <source>Journal of Documentation</source>
          <volume>28</volume>
          ,
          <fpage>11</fpage>
          -
          <lpage>21</lpage>
          (
          <year>1972</year>
          ), http://www.soi.city.ac.uk/˜ser/ idf.html
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zagoris</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arampatzis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          <article-title>: www.mmretrieval.net: a multimodal search engine</article-title>
          . In: SISAP. pp.
          <fpage>117</fpage>
          -
          <lpage>118</lpage>
          . ACM (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>