<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DEMIR at ImageCLEFwiki 2011: Evaluating Different Weighting Schemes in Information Retrieval</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tolga Berber</string-name>
          <email>tberber@cs.deu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ali Hosseinzadeh Vahid</string-name>
          <email>ali_h_vahid@yahoo.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Okan Ozturkmenoglu</string-name>
          <email>okan.ozturkmenoglu@deu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roghaiyeh Gachpaz Hamed</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adil Alpkocak</string-name>
          <email>alpkocak@cs.deu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dokuz Eylul University Dept. of Computer Engineering, DEMIR Research Group Tinaztepe</institution>
          ,
          <addr-line>35160 Izmir</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>This paper present the participation details of DEMIR (Dokuz Eylul University Multimedia Information Retrieval) research team at ImageCLEFwiki2011. This year we investigate on evaluating of different weighting models on text retrieval performance. In the case of low-level feature selection, we extracted different features and examined their performance to choose the proper feature for our experiments. Thereupon to apply late fusion for best gained result of image and textual features. In these experiments we found that choice of proper weighting model may crucially affect the performance of any information retrieval system and also we found that although linear weighted fusion is simplest and frequently used method. The results clearly show that combining text-based and content-based image retrieval with a proper fusion technique improves the performance.</p>
      </abstract>
      <kwd-group>
        <kwd>Information Retrieval</kwd>
        <kwd>Low-level Features</kwd>
        <kwd>Linear Weighted Fusion</kwd>
        <kwd>Combination Algorithms &amp; Methods</kwd>
        <kwd>Feature Extraction and Selection</kwd>
        <kwd>Late fusion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In this paper we present the experiments performed by Dokuz Eylul University
Multimedia Information Retrieval (DEMIR) Group, Turkey, in the context of our
participation to the ImageCLEF 2011Wikipedia Retrieval task. The main focus of this
work is to improve results by examine of different weighting models for retrieved
text data and then choose the best low level feature of figures for fusion with text data
result. During the combination of text and low level features we check the variation of
methods to gain the best result.</p>
      <p>The rest of the paper is organized as follows: In Section 2 we describe our text
retrieval and weighting models examination. In Section 3 we describe the image
features extraction &amp; selection phase. In section 4 we present and compare late fusion
combination methods and we conclude and propose the future work in section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Different Weighting Models in Text Retrieval</title>
      <p>Since the choice of the weighting model may crucially affect the performance of any
information retrieval system specially the text based one, first of all we decided to
work on evaluating the relative merits and drawbacks of different weighting models
using Terrier IR Platform, open source search engine written in Java and is developed
at the School of Computing Science, University of Glasgow.</p>
      <p>
        Terrier provides implementation of the following weighting models: [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>
        We performed our experiments using three ImageCLEF2010 Wikipedia track
monolingual test collection and using all them together. We start from a traditional
bag-of-words representation of pre- processed texts that preprocessing includes
stemming (Porter stemmer [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for English, Snowball for German and French) and
stop word removal. As illustrated in Figure2, we acquired that best result in almost all
experiment using IFB2 model, so we use it to submit our base-line run on Image
CLEF2011 Wikipedia track textual metadata. (RunID 5 in Table1)
Feature extraction is one of the major aspects of a typical content-based information
retrieval (CBIR) system. We call these low-level features because most of them are
extracted directly from digital representations of objects in the database and have little
or nothing to do with human perception. We utilized the Img(Rummager)
application[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is developed in the Automatic Control Systems &amp; Robotics Laboratory
at the Democritus University of Thrace-Greece, and extract features explained below
for all images in ImageCLEF2011 test collection and query examples:
 CEDD: This feature is called “Color and Edge Directivity Descriptor” and
incorporates in histogram color and texture information. CEDD size is limited to
54 bytes per image, rendering this descriptor suitable for use in large image
databases. Important attribute of the CEDD is the low computational power needed
for its extraction, in comparison to the needs of the most MPEG-7 descriptors[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
 FCTH: This feature is named “Fuzzy Color and Texture Histogram” and results
from the combination of 3 fuzzy systems include histogram, color and texture
information. FCTH size is limited to 72 bytes per image, also suitable for use in
large image databases.[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
 BTDH: The Scalable Fuzzy Brightness and Texture Directionality Histogram, was
specially conceived for representing radiology images. It combines brightness and
texture characteristics and their spatial distribution in one compact vector by using
a two-unit fuzzy system. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
 EHD: This Edge Histogram Descriptor proposed for MPEG-7 expresses only the
local edge distribution in the image and is designed to contain only 80 bins for this
purpose. The EHD basically represents the distribution of 5 types of edges in each
local area called a sub-image that is defined by dividing the image space into 4x4
non-overlapping blocks. Thus, the image partition always yields 16 equal-sized
sub-images regardless of the size of the original image. Edges in the sub-images
are categorized into 5 types: vertical, horizontal, 45-degree diagonal, 135-degree
diagonal and non-directional edges. Thus, the histogram for each sub-image
represents the relative frequency of occurrence of the 5 types of edges in the
corresponding sub-image and contains 5 bins.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
 SCD: “Scalable Color Descriptor” is one of the most basic descriptions of color a
feature is a color histogram encoded by a Haar transform. It uses the HSV colors
space uniformly quantized to 255 bins.
 CLD: “Color Layout Descriptor” represents the spatial layout of color images in a
very compact form. It is based on generating a tiny (8x8) thumbnail of an image,
which is encoded via Discrete Cosine Transform (DCT) and quantized. As well as
efficient visual matching, this also offers a quick way to visualize the appearance
of an image, by reconstructing an approximation of the thumbnail, by inverting the
DCT.
      </p>
      <p>After extracting features, we gain an n-dimensional feature space per feature. For
query processing, we had to mapping all the objects in the database and the query
onto this space and then evaluating the similarity/ difference between the vector
corresponding to the query and the vectors representing the data. We selected the
Euclidean distance, one of commonly used similarity and distance functions for
measuring distances between points in the 3D space, as distance/similarity function
and based on obtained similarity scores, we found that CEDD and FCTH are the best
descriptors for image retrieval based on low level features only. Therefore we
submitted our visual only base point run for CEDD feature. (RunID 6 in table 1)
Moreover we use these features for multimodal fusion in next experiments, as explain
below:</p>
      <p>DEMIR at ImageCLEFwiki 2011: Evaluating</p>
      <p>
        Different Weighting Schemes in Information Retrieval 5
Multimedia fusion is referred to as integration of multiple media, their associated
features, or the intermediate decisions in order to perform an analysis task, has gained
much attention of many researchers in recent times. The fusion of multiple modalities
can provide complementary information and increase the accuracy of the overall
decision-making process [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The fusion of different modalities is generally performed at two levels: feature
level or early fusion and decision level or late fusion. Some researchers have also
followed a hybrid approach by performing fusion at the feature as well as the decision
level. In the feature level or early fusion approach, the features, some distinguishable
properties of a media stream, extracted from input data are first combined and then
sent as input to a single analysis unit that performs the analysis task. In the decision
level or late fusion approach, the analysis units first provide the local decisions D1 to
Dn that are obtained based on individual features F1 to Fn. Then a decision fusion unit
combines local decisions to make a fused decision vector that is analyzed further to
obtain a final decision D about the task or the hypothesis. To achievement the
advantages of both the feature level and the decision level fusion strategies, several
researchers have opted to use a hybrid fusion strategy, which is a combination of both
feature and decision level strategies.</p>
      <p>
        The decision level fusion strategy has many advantages over feature fusion. For
instance, the decisions (at the semantic level) usually have the same representation.
Therefore, the fusion of decisions becomes easier. Moreover, the decision level fusion
strategy offers scalability (i.e. graceful upgrading or degradation) in terms of the
modalities used in the fusion process, which is difficult to achieve in the feature level
fusion. Another advantage of late fusion strategy is that it allows us to use the most
suitable methods for analyzing each single modality and this provides much more
flexibility than the early fusion. Because of these profits, we exerted Linear Weighted
Fusion, one of the simplest and most widely used methods on our extracted CEDD
and FCTH similarity scores and similarity scores that gained from text retrieval as
explained in previous chapters. Since different retrieval results can generate quite
different ranges of similarity values, a normalization method should be applied to get
accurate and correct results. Hence we apply Max-Min normalization on similarity
values to ensure that the range of these features is between 0 and 1. Then we applied
Fagin’s Combination Algorithms [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] for Ranked Input Sets putting on different
combination method and based on our previous study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], we found that the method
called Comb-SUM for summing the similarity values provide the best results. So we
combined the CEDD and FCTH features normalized similarity scores with textual
similarity score in this manner and submitted two other runs (RunID 3, 4 in Table 1).
On the other hand, the Weighted Sum function applied in the same manner but
differing on multiplying each individual similarity with a weight value and submitted
two the best runs of our experiments. As shown in Table 1, we applied Weighted
Comb-SUM combination method with multiplying CEDD feature by 2 and result of
retrieved textual feature by 3 and submitted result as RunID 2 in Table1. Also in
RunID 1 in Table 1 as our best ranked run, first we use Comb-SUM combination for
fused CEDD and FCTH features similarity scores then combined them using
Weighted Comb-SUM with two folds of retrieved textual features.
      </p>
      <p>Bpref
0.2564
0.2573
0.2554
0.2529
0.2476
0.0115</p>
      <p>So it is apparent when we combine the results of different modalities, all of the
performance evaluation factors in retrieval system improved.
5</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this year, we examined effects of different weighting models on text retrieval and
found that the role of proper weighting model selection is to improve the performance
of text retrieval systems. Also, we compare MAP of different extracted low-level
features normalized similarity scores and due to this comparison we select CEDD and
FCTH descriptors as suitable features to utilize for fusion to textual results. Also due
to analogy of combination methods in our previous studies, we acquire choosing a
suitable combination method for fusion improved the results. The results clearly show
that combining text-based and content-based image retrieval with a proper fusion
technique improves the performance.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. The Terrier IR Platform, http://terrier.org/docs/v2.2.1/</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Amati</surname>
          </string-name>
          , Giambattista:
          <article-title>Probability models for information retrieval based on divergence from randomness</article-title>
          ,
          <source>PhD Thesis</source>
          , University of Glasgow Faculty of Information and
          <source>Mathematical Sciences Department of Computing Science</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          :
          <article-title>An algorithm for suffix stripping</article-title>
          ,
          <source>Program: electronic library and information systems</source>
          , vol.
          <volume>14</volume>
          ,
          <issue>iss</issue>
          . 3, pp.
          <fpage>130</fpage>
          --
          <lpage>137</lpage>
          (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boutalis</surname>
            ,
            <given-names>Y.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lux</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Img(Rummager): An Interactive Content Based Image Retrieval System</article-title>
          .
          <source>In: 2nd International Workshop on Similarity Search and Applications</source>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>153</lpage>
          . IEEE Computer Society, Washington (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boutalis</surname>
            ,
            <given-names>Y.S.:</given-names>
          </string-name>
          <article-title>FCTH: Fuzzy Color and Texture Histogram - A Low Level Feature for Accurate Image Retrieval</article-title>
          .
          <source>In: 9th International Workshop on Image Analysis for Multimedia Interactive Services</source>
          , vol., no., pp.
          <fpage>191</fpage>
          -
          <lpage>196</lpage>
          . Klagenfurt,
          <string-name>
            <surname>Austria</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chatzichristofis</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boutalis</surname>
            <given-names>Y.S.:</given-names>
          </string-name>
          <article-title>Content based radiology image retrieval using a fuzzy rule based scalable composite descriptor</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          , vol.
          <volume>46</volume>
          ,
          <issue>iss</issue>
          . 2, pp.
          <fpage>493</fpage>
          --
          <lpage>519</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Won</surname>
            <given-names>C. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            <given-names>D. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park S</surname>
          </string-name>
          .J.:
          <article-title>Efficient Use of MPEG-7 Edge Histogram Descriptor</article-title>
          ,
          <source>ETRI Journal</source>
          , vol.
          <volume>24</volume>
          , no.
          <issue>1</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pradeep</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Atrey</surname>
          </string-name>
          , Anwar Hossain M.
          <article-title>: Multimodal fusion for multimedia analysis</article-title>
          ,
          <source>Multimedia Systems</source>
          , vol
          <volume>16</volume>
          , pp.
          <fpage>345</fpage>
          --
          <lpage>379</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Fagin</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lotem</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naor</surname>
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Optimal aggregation algorithms for middleware</article-title>
          ,
          <source>In: Journal of Computer and System Sciences</source>
          , vol.
          <volume>66</volume>
          , pp.
          <fpage>614</fpage>
          --
          <lpage>656</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ulker</surname>
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Analysis and comparison of combination algorithms for joining ranked inputs</article-title>
          ,
          <source>MSc Thesis</source>
          , Dokuz Eylül University Department of Computer Engineering, Izmir, Turkey (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>