<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UEC, Tokyo at MediaEval 2013 Retrieving Diverse Social Images Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Keiji Yanai</string-name>
          <email>yanai@cs.uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Do Hang Nga</string-name>
          <email>dohang@mm.cs.uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The University of Electro-Communications</institution>
          ,
          <addr-line>Tokyo 1-5-1 Chofugaoka, Chofu-shi, Tokyo 182-8585</addr-line>
          ,
          <country country="JP">JAPAN</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>In this paper, we describe our method and results for the MediaEval 2013 Retrieving Diverse Social Images Task. To accomplish the task objective, we adopt VisualRank [5] and Ranking with Sink Points [2], which are common methods to select representative and diverse photos. To obtain an affinity matrix for both ranking methods, we used only the officially-provided features including visual features and tag features. We submitted three required runs including only visual feature run, only textual feature run and textualvisual fused feature run.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        In this paper, we describe our method and results for the
MediaEval 2013 Retrieving Diverse Social Images Task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
The objective of this task is to select relevant and diverse
photos from the given photos regarding the specific
locations. To do that, we adopt VisualRank [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Ranking
with Sink Points [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The reason why we adopted these
method is that we had used these methods for ranking
geotagged photos [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. First we calculate a similarity matrix
using the given features, and we apply VisualRank to select
the most representative photo. Then we re-rank the
remaining photos by Ranking with Sink Points after removing the
first-ranked photo. We repeat re-ranking by Ranking with
Sink Points and removing the first-ranked photos until 50
photos are selected.
      </p>
      <p>To obtain a similarity matrix for both ranking methods,
we used only the officially-provided features including visual
features and tag features. We submitted three required runs
including only visual feature run, only textual feature run
and textual-visual fused feature run, which are the minimum
requirements to participate this task.</p>
    </sec>
    <sec id="sec-2">
      <title>RANKING METHOD</title>
      <p>
        To obtain representative and diverse photos in the
upper rank, we adopt VisualRank [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Ranking with Sink
Points [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this section, we explain both methods and
features briefly.
      </p>
    </sec>
    <sec id="sec-3">
      <title>VisualRank</title>
      <p>
        VisualRank is an image ranking method based on
PageRank [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. PageRank calculates ranking of Web pages using
hyper-link structure of the Web. The rank values are
estimated as the steady state distribution of the random-walk
Markov-chain probabilistic model.
      </p>
      <p>VisualRank uses a similarity matrix of images instead of
hyper-link structure. Eq.(1) represents an equation to
compute VisualRank.</p>
      <p>
        ri+1 = αSri + (1 − α)p, (0 ≤ α ≤ 1)
(1)
S is the column-normalized similarity matrix of images, p
is a damping vector, r is the ranking vector each element of
which represents a ranking score of each image, and α plays
a role to control the extent of effect of p. The final value of r
is estimated by updating r iteratively with Eq.(1). Because
S is column-normalized and the sum of elements of p is 1,
the sum of ranking vector r does not change. Although p
is set as a uniform vector in VisualRank as well as normal
PageRank, it is known that p can plays a bias vector which
affects the final value of r [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Ranking with Sink Points</title>
      <p>
        Because VisualRank is a ranking method considering only
representativeness of items, higher ranks are sometimes
occupied with items which are similar to each other. This is,
VisualRank cannot accomplish ranking considering diversity
of items. Therefore, we adopt Ranking with Sink Points [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
which can be regarded as an extension of PageRank [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to
make obtained ranking relevant and diverse.
      </p>
      <p>To address the diversity in ranking, the concept of sink
points is useful. The sink points are data objects whose
ranking scores are fixed at zero during the ranking process.
Hence, the sink points will never spread any ranking score to
their neighbors. Intuitively, we can imagine the sink points
as the “black holes” on the ranking manifold, where ranking
scores spreading to them will be absorbed and no ranking
scores would escape from them.</p>
      <p>First we apply VisualRank to select the most
representative photo with the obtained affinity matrix. Then we
re-rank the remaining photos by Ranking with Sink Points
as shown in Eq.(2), after removing the first-ranked photo as
“a sink point”. We repeat re-ranking by Ranking with Sink
Points and removing the first-ranked photos until 50 photos
are selected as following:
ri+1 = αSIiri + (1 − α)p
(2)
Ii is an indicator matrix which is a diagonal matrix with its
(i, i) − element equal to 0 if xi ∈ Xs and 1 otherwise. Xs is
a set of “sink points”.</p>
      <p>Note that α is set as 0.85 in the experiments.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Visual Features</title>
      <p>We used the ten kinds of visual features officially provided
by the task organizers such as Global Histogram of Oriented</p>
      <p>
        Gradient and Color Moments on HSV Color Space. The
detail on official visual features is explained in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>With histogram intersection, we calculate similarities for
each of visual features. Finally we construct an affinity
matrix by averaging similarity on ten kinds of visual features.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Textual Features</title>
      <p>We use social TF-IDF weights provided by the task
organizers. We extract bag-of-words vectors from Flickr
metadata with social TF-IDF weights for all the given images.
We calculate an affinity matrix with cosine similarity
between bag-of-words vectors within each place.</p>
      <p>To obtain an affinity matrix for the visual-textural-fused
runs, we simply averaged both visual-feature-based affinity
matrix and textual-feature-based affinity matrix.</p>
    </sec>
    <sec id="sec-7">
      <title>EXPERIMENTAL RESULTS</title>
      <p>Tables 1 and 2 show the evaluated results of our three
submission runs by experts and crowds, respectively. Note
that the results by experts is based on evaluation for the
entire dataset of 346 locations, while the results by the crowds
is based on evaluation for only 50 locations in the dataset
and are obtained by averaging evaluations by three crowd
persons.</p>
      <p>Basically, the results by only visual were better than the
results by only textual and the results by visual-textual,
although the difference were not so large.</p>
      <p>We show the top six photos of an successful example by
the proposed method with three kinds of features: textual,
visual and visual-textual features in Figure 1. These photos
represents “The Gate of Forbidden City in Beijing, China.”
In this example, the photos selected by the visual-textual
feature is more representative and diverse than the photos
selected by the only textual or only visual features. This
indicates that our proposed methods works successfully.</p>
      <p>In the case of the above example, most of the photos
included in the given photo set are relevant and only a few
noise photos are included. However, given photo sets of
some landmark include many noise photos. In such case, the
proposed methods sometimes failed to select relevant photos
and selected noise photos in the upper ranking. Therefore,
removal of noise photos is one of our important future works.</p>
    </sec>
    <sec id="sec-8">
      <title>CONCLUSIONS</title>
      <p>ing dataset, GPS data coordinates and Wikipedia photos.
In fact, if you had enough time, we should have used the
training data for estimating optimal parameters such as α
in the VisualRank formulation and a mixing weight of visual
similarity and textual similarity.
5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Brin</surname>
          </string-name>
          and
          <string-name>
            <surname>L. Page.</surname>
          </string-name>
          <article-title>The anatomy of a large-scale hypertextual web search engine</article-title>
          .
          <source>In Proc. of the Seventh International World Wide Web Conference</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.-Q.</given-names>
            <surname>Cheng</surname>
          </string-name>
          , P. Du,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. Chen.</surname>
          </string-name>
          <article-title>Ranking on data manifold with sink points</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>25</volume>
          (
          <issue>1</issue>
          ):
          <fpage>177</fpage>
          -
          <lpage>191</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Haveliwala</surname>
          </string-name>
          .
          <article-title>Topic-sensitive PageRank: A context-sensitive ranking algorithm for web search</article-title>
          .
          <source>IEEE trans. on Knowledge and Date Engneering</source>
          ,
          <volume>15</volume>
          (
          <issue>4</issue>
          ):
          <fpage>784</fpage>
          -
          <lpage>796</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Menendez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescu</surname>
          </string-name>
          .
          <article-title>Retrieving diverse social images at mediaeval 2013: Objectives, dataset and evaluation</article-title>
          . In MediaEval 2013 Workshop, CEUR-WS.org, ISSN:
          <fpage>1613</fpage>
          -
          <lpage>0073</lpage>
          , Barcelona, Spain, October
          <volume>18</volume>
          -19
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jing</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Baluja</surname>
          </string-name>
          . Visualrank:
          <article-title>Applying pagerank to large-scale image search</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>30</volume>
          (
          <issue>11</issue>
          ):
          <fpage>1870</fpage>
          -
          <lpage>1890</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kawakubo</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Yanai. Geovisualrank</surname>
          </string-name>
          :
          <article-title>A ranking method of geotagged images considering visual similarity and geo-location proximity</article-title>
          .
          <source>In Proc. of the ACM International World Wide Web Conference</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>