<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RGU at ImageCLEF2010 Wikipedia Retrieval Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jun Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dawei Song</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leszek Kaliciak</string-name>
          <email>l.kaliciakg@rgu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computing, The Robert Gordon University</institution>
          ,
          <addr-line>Aberdeen</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This working notes paper describes our first participation in the ImageCLEF2010 Wikipedia Retrieval Task[1]. In this task, we mainly test our Quantum Theory inspired retrieval function on cross media retrieval. Instead of heuristically combining the ranking scores independently from different media types, we develop a tensor product based model to represent textual and visual content features of an image as a non-separable composite system. Such system incorporates the statistical/semantic dependencies between certain features. Then the ranking scores of the images are computed in a way as quantum measurement does. Meanwhile, we also test a new local feature that we have developed for content based image retrieval.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>In our participation to ImageCLEF we have submitted runs in the Wikipedia Retrieval
Task, which are based on our Quantum Theory inspired retrieval model. This model
applies tensor product to represent the textual and visual features of an image as a
norder tensor in order to capture the non-separability of textual and visual features. The
order of the tensor depends on the visual features that are going to be incorporated in
the image retrieval.</p>
      <p>We have tested this retrieval model on the ImageCLEF2007 data collection that has
20,000 images in total, and achieved some promising results. Therefore we would like
to test it on a larger scale dataset.
Mathematically, the tensor product is used to construct a new vector space or a new
tensor, where the relationship of the vector spaces can be expressed. In quantum
mechanics, the tensor product can be used to expand Hilbert spaces or construct a composite
system with separate systems.</p>
      <p>Suppose an image represented in a single feature space to be a single system Si, then
that image can be constructed as a composite system by tensor product the separated
systems from different feature spaces:</p>
      <p>S = S1</p>
      <p>S2</p>
      <p>
        Sn
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
Where Si is a system in a single feature space.
where Pi wt2i = 1, and the amplitude wti is proportional to the probability that the
document is about the term ti. wti can be set up with any term weighting scheme, and
we adopted TF-IDF in our experiment.
      </p>
      <p>Similarly, the histogram of content feature can also be represented as a superposed
state in a content feature space:
jT i =</p>
      <sec id="sec-1-1">
        <title>X wti jtii</title>
        <p>i
jF i =</p>
      </sec>
      <sec id="sec-1-2">
        <title>X wfi jfii</title>
        <p>i
where Pi wf2i = 1, fi is the a particular feature bin, and wfi is proportional to the
number of pixels falling into the corresponding bin of the feature space.</p>
        <p>When more than one features are used to represent an image, the representation will
be:</p>
        <p>jDi = jT i jF1i jF2i jFni
In our pilot study, we only combine one content feature with the textual feature.
jDi = jT i
= X
ij
jF i
ij jti
fj i</p>
        <p>Next lets look at how we present a single system in the Hilbert space with Dirac
notation. Here we will not introduce the detail of Dirac notation, readers who are
interested can refer to Van Rijsbergen’s book “The Geometry of Information Retrieval”[2].</p>
        <p>
          With quantum mechanics, traditionally vector based representation of documents
are represented as superposed states. The textual feature of an image is:
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(3)
(4)
(5)
(6)
(7)
(8)
Here, i and j are the dimensionalities of textual and content features. When textual
and content feature are completely independent, then ij = wti wfj . However, this
does not hold generally. Extra operation is necessary to reflect the non-separability of
the two features. It can be operationalised as the co-occurrence or correlation of the
features. The tensor product enables the expansion of feature spaces in a seamless way
and incorporates the correlations between the feature spaces.
        </p>
        <p>With superposed representation, the similarity of a document and a query can be
viewed as the probability that the document projects onto the sub-space that is expanded
by the query features.</p>
        <p>The probability that a document collapses to a state is:</p>
        <p>P (tijd) = jhtijT ij2 = wt2i
It can be described as a projection onto a space spanned by jtii:</p>
        <p>P (tijd) = htij djtii = wt2i
where d = jdihdj is the density matrix of document d, and wt2i is the probability that
the term ti appears in the document or the probability that the document is about the
term.</p>
        <p>With the textual and visual composite system, the density matrix of a document is:
Then the similarity between a document and a query is:</p>
        <p>D = jDihDj = X i2j jtifj ihtifj j
(9)
(10)
(11)
(12)
(13)
(14)
(15)
(16)
i
= X</p>
        <p>jk
d =</p>
        <p>X wi2jtiihtij</p>
        <p>i
= X wi2 X Uij jej i X</p>
        <p>j k
jkjej ihekj</p>
        <p>hekjUik
sim(d; q) = tr(X di2jtiihtijQ)</p>
        <p>i
= tr(X di2jtiihtij X qj2jtj ihtj j)</p>
        <p>i j
= X di2qi2</p>
        <p>i</p>
        <p>In the current experiment, we assume that all the terms are orthogonal to simplify
the calculation, then:</p>
        <p>For a visual feature, each of its dimensions can be treated as orthogonal to other
dimensions. Because the blue color pixel can never be counted as a red pixel. While
for textual feature, two terms can be semantically related, e.g. cup and mug may refer
to the same thing in one document. To represent the document with orthogonal textual
basis, we can use a transformation matrix to fulfill the requirement:
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Experiment Settings and Results</title>
      <sec id="sec-2-1">
        <title>Text Processing</title>
        <p>When we associated texts with images, we not only used annotation documents, but
also used Wikipedia pages, hoping the original full text document can provide more
semantic information.</p>
        <p>Although in the Wikipedia task, the images may be annotated by more than one
language, due to our language expertise we only parse the English language. If the
images are not annotated by English, or do not appeared in English Wikipedia page,
then this image will not be indexed. It also means that this image will never be retrieved
in the runs using both text and content features as queries.</p>
        <p>For the annotation files, we parsed all the terms from name, description, comment
and caption entries if they contain any term. For the Wikipedia dump files, We parsed all
the terms in the Wikipedia page, including the the link names within the page. However
we do not go further into the linked page. The tags particular to the Wikipedia webpage
are removed.</p>
        <p>Same for the queries, we only use their English titles during the retrieval stage.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Content Feature</title>
        <p>Apart from the content features provided by the organizer, we also used our local
feature, which is based on the “bag of features”/“bag of visual words” approach. The
feature extraction consists of following stages: image sampling, description of local
patches, feature vector construction, visual dictionary generation and histogram
computation. The number of sample points is 900 per image and the sampling is purely
random. We open a square window - local patch (10 by 10 pixels wide) centred at
the sample point. Each local patch is represented in a form of three colour moments
computed for individual colour channels in HSV colour space. Thus obtained vector
representation of a local patch has 9 dimensions. We apply the K-means clustering
to the training set in order to obtain the codebook of visual words. Finally, we create
a histogram of visual word counts by calculating manhattan distance between image
patches and cluster centroids and generate a vector representation for each image from
the collection. Thus obtained vector representation of an image has 40 dimensions.
3.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Experiment runs and Results</title>
        <p>With the textual and visual feature available, we submitted the following runs, some
of which are based on content feature only and some are content and textual feature
mixed.</p>
        <p>– Text and content mixed retrieval</p>
        <p>T+F L : retrieve on annotation first, then re-rank with our local feature
T+F C : retrieve on annotation first, then re-rank with cime
TXF L : quantum-like measurement on tensor product space of annotation
vector and our local feature vector
TXF C : quantum-like measurement on tensor product space of annotation
vector and feature cime vector
combine: T+F C based retrieval. When the length of result from T+F C is less
than 1000, then the images from content retrieval are appended into the result
list.</p>
        <p>W+F C : retrieve on Wikipedia file first, then re-rank with cime
– Content only retrieval
c leszek: city block distance with our new local feature
c add: city block distance with all content feature provided by ImageCLEF
organiser</p>
        <p>For the mixed retrieval, we did not run the mixed retrieval process on the whole
collection due to the huge size. We ran the text retrieval first, then applied mixed retrieval
to the re-ranking.</p>
        <p>From table 1, we can see that the retrieval result based on content feature has
extremely low MAP. The retrieval results from text and content mixed retrieval in
ImageCLEF2010 is also considerably lower than our results from ImageCLEF2007 whose
MAPs that are around 14%.</p>
        <p>After further looking into the tensor product based experiments, we did find a bug in
the code. Because a “+” operator is missing, which resulted in that only the last textual
and content feature dimension had been used to re-rank the image. This accounts the
poor performance of the submission runs.</p>
        <p>We have corrected the code and re-run the experiments. The text based result is
MAP=0.0939 and P@10=0.3485; The tensor product based result with our local feature
is MAP=0.0665 and P@10=0.2000. There is a slightly improvement after removing the
bug from the experiment, but it is still a lot of worse than text based retrieval, which
needs a further investigation.</p>
        <p>Run Modality field(s) MAP P@10
combine Mixed TITLEIMG 0.0617 0.2271</p>
        <p>T+F L Mixed TITLEIMG 0.0617 0.2257
TXF C Mixed TITLEIMG 0.0486 0.1443
TXF L Mixed TITLEIMG 0.0484 0.1500
T+F C Mixed TITLEIMG 0.0325 0.1143
W+F C Mixed TITLEIMG 0.0031 0.0086
c leszek Visual IMG 0.0069 0.0614
c add Visual IMG 0.0003 0.0100</p>
        <p>Table 1. Our runs in ImageCLEF2010
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and Future Works</title>
      <p>In this notes paper, we reported our quantum theory inspired multimedia retrieval
framework, which provides a formal and flexible way to expand the feature spaces, and
seamlessly integrate different features. The similarity measurement between query and
document follows quantum measurement.</p>
      <p>We did not include the correlations between dimensions across different feature
spaces in this year submission. We would like to investigate such issue in the continuing
study, and further the study with entanglement concept in quantum mechanics.</p>
      <p>In the current experiments, we assumed each word is orthogonal, which does not
hold and can be relaxed in the future. We can solve this problem with either
dimensionality reduction or utilize the thesaurus to remove the synonyms. This will also facilitate
the ranking computation.
This work is funded in part by the UK’s Engineering and Physical Sciences Research
Council (grant number: EP/F014708/2).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Popescu</surname>
          </string-name>
          , Theodora Tsikrika, and
          <string-name>
            <given-names>Jana</given-names>
            <surname>Kludas</surname>
          </string-name>
          .
          <article-title>Overview of the wikipedia retrieval task at imageclef 2010</article-title>
          .
          <source>In Working Notes of CLEF</source>
          <year>2010</year>
          , Padova, Italy,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>C. J. van Rijsbergen</surname>
          </string-name>
          , editor.
          <source>The Geometry of Information Retrieval</source>
          . Cambridge University Press,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>