<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MIL at ImageCLEF 2014: Scalable System for Image Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Atsushi Kanehira</string-name>
          <email>kanehira@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Masatoshi Hidaka</string-name>
          <email>hidaka@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yusuke Mukuta</string-name>
          <email>mukuta@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuichiro Tsuchiya</string-name>
          <email>tsuchiya@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetsuaki Mano</string-name>
          <email>mano@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tatsuya Harada</string-name>
          <email>harada@mi.t.u-tokyo.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Machine Intelligence Lab., The University of Tokyo</institution>
        </aff>
      </contrib-group>
      <fpage>372</fpage>
      <lpage>379</lpage>
      <abstract>
        <p>In this working note, we describe details of our method in ImageCLEF2014 Scalable Concept Image Annotation task. We are given images from the Web and some additional information, including web pages, in which images exist. Using this information, we must construct an annotation system that has high performance and scalability. To assign labels to each image, we use the page title and attributes of an image tag extracted from the web page. As visual features, we propose the use of the combination of two complementary features, which are Fisher Vector and deep convolutional neural network based feature. They are generative and discriminative feature respectively. We then train linear classi ers using Passive{Aggressive with Averaged Pairwise Loss. After training, we calculate the score of each concept for test data and label some concepts having the best scores. Results show that the combination of two features contributes to the improvement of recognition performance.</p>
      </abstract>
      <kwd-group>
        <kwd>ImageCLEF</kwd>
        <kwd>deep convolutional neural network</kwd>
        <kwd>Fisher vector</kwd>
        <kwd>Image annotation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        For the ImageCLEF2014 Scalable Concept Image Annotation task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], our
task is to construct an image annotation system that yields high-performance
with scalability.
      </p>
      <p>
        As visual features, we use a convolutional neural network (CNN) based
feature as well as the Fisher Vector (FV) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For scalability, our method of assigning
labels to training images is simple. We use only the page title and attributes of
the image tag extracted from the web page. To train linear classi ers, we use
Passive{Aggressive with Averaged Pairwise Loss (PAAPL) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] because of its
scalability and robustness to noise of label assignment.
      </p>
      <p>In our experiment, we combine two types of features, which are FV and
CNNbased features. Actually, FV is often used in image recognition tasks because of
its recognition performance. However, many results of recent studies show that
deep CNN achieves high performance on many tasks. Therefore, we expect that
the feature, which is the neuron activation pattern in the hidden layers of the
network, has high-representational ability.</p>
      <p>These two features are extracted in completely di erent ways. We obtain FV
by coding local descriptors considering their probabilistic distribution. Therefore
FV can be regarded as the feature expressing generative information of an image.
In contrast to FV because deep CNN based feature is extracted from the network
trained for recognition task, we can regard it as a discriminative feature.</p>
      <p>Assuming that these two types of features, which represent di erent kinds
of information, mutually compensate for representational ability, we propose
their combined use. Our contribution is the usage of a combination of features
that have complementary properties to improve the performance of annotation
systems.</p>
      <p>The remainder of this working note is structured as follows. Section 2 presents
a description of two types of visual features: FV and deep CNN based features.
In Section 3, we explain details of how we obtain labels from training data.
Then, in section 4, we introduce a multi-label linear classi er training method:
PAAPL. In section 5, we present the results of experiments, using either or both
of these visual features. Finally, in section 6, we discuss the analysis of the results
obtained in our experiment.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Visual Feature</title>
      <sec id="sec-2-1">
        <title>Fisher Vector</title>
        <p>As a visual feature, we use Fisher Vector (FV) because FV can achieve better
recognition performance than Bag of Visual Words with a linear classi er. In
general, the linear classi er is less costly than a nonlinear one such as
kernelSVM when the amount of the training sample increases. Therefore, FV is suitable
for this task, which requires scalability.</p>
        <p>In our experiments, we extract four local descriptors: SIFT, GIST, LBP,
and C-SIFT. The dimensions of all these local descriptors are reduced to 64
dimensions using Principal Component Analysis (PCA). These local descriptors
are densely extracted from ve scales of patches (squares 16, 25, 36, 49, 64 pixels
on a side) sliding with a step of six pixels. Using some of the obtained local
descriptors, we train a Gaussian Mixture Model (GMM) with 256 components,
which have diagonal matrices as covariance matrices. After training GMM, we
extract FV from each image by calculating the gradient of log-likelihood of local
descriptors with respect to parameters of GMM. Then we normalize it using
the Fisher information matrix. Power normalization and L2 normalization are
applied to the extracted FVs. To include spatial information, we divide images
into 1 1, 2 2, and 3 1 cells, extract features from each region and concatenate
them into one vector. The nal dimension of FV is 262,144 (64 256 2 8).</p>
        <p>As described above, FV is obtained by coding local descriptors considering
their probabilistic distribution. Therefore FV can be regarded as a feature
expressing generative information of the image.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Deep CNN based feature</title>
        <p>In addition to FV, we use a deep convolutional neural network (CNN) based
feature extracted from a deep CNN model that had been pre-trained with ImageNet
dataset.</p>
        <p>
          In recent years, many studies that speci cally examine deep CNN have shown
that such models can perform better than conventional feature representation in
object recognition and other tasks. On ImageNet Large Scale Visual Recognition
Challenge, one system [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] dramatically outperformed all other methods including
a state-of-the-art method using FV.
        </p>
        <p>However, training deep architecture of CNN requires large-scale data to
prevent over tting and to ensure generalization ability. Trained models will have
low recognition performance if the training data are few.</p>
        <p>For this task, because we must label training data automatically, the number
of reliably labeled data we can obtain using our method is roughly 100,000 on
the development set and about 200,000 on the test set. This fact implies that the
approximate number we can use in training models is only 1,000 per concept,
on average. Therefore, because of limitations in the amount of data in the given
dataset, it is di cult to train deep CNN models to produce high recognition
performance.</p>
        <p>
          According to [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], features extracted from the activation of a CNN pre-trained
in supervised fashion can be re-purposed to generic tasks. In his experiment, he
shows that such feature representation has such high generality that it
outperforms conventional methods in several tasks despite their simple training
algorithm.
        </p>
        <p>
          In consideration of the discussion presented above, we use feature
representations extracted from a deep CNN model pre-trained with the ImageNet dataset.
Following the method of [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], we extract features from sixth and seventh layers
of the network having architecture that is the same as that proposed by [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
which won ILSVRC2012 and which includes ve convolutional and three fully
connected layers. In some layers, the Hinge function is used as the activation
function. It is designated as ReLU.
        </p>
        <p>Contrary to FV, because deep CNN-based features are extracted from the
network, which is trained for recognition task, we can regard it as a feature that
expresses discriminative information of an image.</p>
        <p>
          In our experiment, we use DeCAF, an open source library produced by [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], to
extract deep feature representation. We use features of four types. The feature
vector dimensions are 4,096.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Label Assignment</title>
      <p>Because no explicit labels are assigned to training images, we must label them
using additional information or external sources. In this section, we describe how
to assign labels to images. The pipeline of label assignment is shown in Fig.1.
We label images in the following two steps.</p>
      <sec id="sec-3-1">
        <title>Concept C!</title>
        <p>webpage!</p>
        <p>WordNet
Database</p>
        <p>Parse
XML!</p>
        <p>WC=
{C,synonym ( C ) , hyponym ( C )}
T={words related to images}</p>
      </sec>
      <sec id="sec-3-2">
        <title>Label C!</title>
        <p>Assign</p>
        <p>Label! if T ∩ WC ≠ 0</p>
        <p>In the rst step, we parse the xml les of the web page, in which an image
exists. Then we extract page titles and attributes of the image tag, which include
src, title, and alt. The hope is that these attributes include important information
about what the image represents. Then we split them into a set of single words
T . For example, if there is an image tag in an xml le shown in Fig.2, then we
obtain</p>
        <p>T = fQueen; Prince; corgi; family; abroadg:
!"
"
!"!!#
!"#$%&amp;'""
#
#
#
#####"$%&amp;#'()*)+(&amp;$,--.!/"01234567##
#######89:;#*#&lt;=::#&gt;3+&gt;(?&lt;#</p>
        <p>"
#######&gt;:9*@=::#&gt;3+&gt;(?A#9B;#CD;;E#&gt;E?#6($E);#####
#######F?G&gt;(?#G$9B#+E;!#+"H!#9!B#;#H&gt;%$:IJ'#)+(&amp;$#KL#
#
#
#
#
#
"
!"#$%#&amp;"</p>
        <p>D(:#+H#$%&gt;&amp;;"
89:;#+H#$%&gt;&amp;;"
=:9;(E&gt;8M;#G+(?#GB;E#
$%&gt;&amp;;#)&gt;EN9#3;#(;&gt;?"</p>
        <p>
          In the second step, we collect a set of synonyms and hyponyms for each
concept C using WordNet [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We denote the collected sets by WC , which is
expressed as
        </p>
        <p>WC = fC; synonym(C); hyponym(C)g
where synonyms (C) and hyponyms (C) respectively represent sets of synonyms
and hyponyms of the concept C. For example, given a target concept \dog", we
obtain</p>
        <p>Wdog = fdog; puppy; corgi; :::g:
We assign the concept C as the label to the image if at least one word in WC
appears in T .
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Training Classi er</title>
      <p>
        In this section, we introduce the multi-label linear classi er training method
Passive{Aggressive with Averaged Pairwise Loss (PAAPL) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Because PAAPL
is based on Passive{Aggressive (PA) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] method, which is known to be robust
to outliers, PAAPL also has robustness to outliers. In addition, PAAPL has
scalability because the trained classi er is linear. Such properties of PAAPL are
suitable for our task, in which it is necessary to construct a scalable system that
handles data including some outliers.
      </p>
      <p>First, we describe the model update rule of PA. Given the t-th training
sample, we designate the visual feature by xt. We de ne Yt as the set of labels
assigned to the t-th sample, and Yt as the set of labels not assigned. The model
(weight) of the linear classi er corresponding to concept label C before updating
by the t-th sample is denoted by wtC .
1. Fetch the t-th training sample. Then compute scores for each label using
current models. Because classi ers we are training are linear, scores are given
by simple calculation of the inner product of weight and feature.
2. Based on scores, nd a combination of labels rt 2 Yt and st 2 Yt in the
following way.</p>
      <p>rt = arg min wtr xt</p>
      <p>r2Yt
st = arg max wts xt</p>
      <p>s2Yt
3. For a combination of rt and st, compute the hinge-loss l
l(wtrt ; wtst ; (xt; Yt)) =
(0
1
if wtrt xt</p>
      <p>wtst xt &gt; 1
(wtrt xt
wtst xt) otherwise
4. Update models using hinge-loss according to the following rule.
wtr+t1 = wtrt +
wts+t1 = wtst
l
l
2jxtj2 + D1 xt
2jxtj2 + D1 xt
Therein, D is a Passive{Aggressive parameter that reduces the negative
inuence of noisy labels.</p>
      <p>Then we describe the method of training classi ers with PAAPL.
1. Pick the t-th training sample, compute scores for target labels using current
models.
2. For a randomly selected combination of labels rt 2 Yt and st 2 Yt,
hingeloss is calculated as PA and remove rt and st from Yt and Yt. Continue this
process until jYtj = 0 or jYtj = 0.
3. For combinations satisfying the condition that the hinge-loss is not 0, update
the models according to the update rule of PA.</p>
      <p>In PAAPL, convergence of models is faster than in PA because PAAPL
updates multiple pairs of models for one sample, whereas PA updates only one pair
of models.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>At the training phase, we rst extract visual features and assign labels. Then,
linear classi ers are trained. At the test phase, we calculate scores for test
images using the linear classi ers we trained. The concepts are labeled on those
images. When training classi ers, the number of iterations is set to 5 and
Passive Aggressive parameter D is set to 1:0 105. After the training process, we
average scores from di erent models trained using di erent visual features. Also,
we decide concepts by selecting those with scores in the top 4% of all given
concepts to each test sample. In our experiments, we compare two types of visual
features. The settings, except for the combinations of visual features, are xed
throughout all of our experiments.</p>
      <p>First, for each type of feature, we respectively search for the best
combination. For FV, we use four local descriptors, SIFT, C-SIFT, GIST, and LBP.
Because the respective properties of these four features di er, we try all possible
combinations of them. As for deep CNN-based features, we extract them from
sixth and seventh layers. From each layer, we obtain features of two types, such
as activations and outputs of each unit. Therefore we also obtain four types of
visual features. In contrast to the case of FV, we need not try all possible
combinations because combinations of features from the same layer do not make sense
theoretically: they are expected to have similar properties. As shown in Tables
Table 1 and Table 2, in both cases, the results of combining all features achieves
higher performance than the others. These results respectively correspond to our
Run 1 and Run 2.</p>
      <p>Finally, we combine these two types of visual features. Using the information
presented above, we combine all four FVs and four deep CNN based features.
The nal results are presented in a Table in Table 3, which correspond to all
Runs we submitted. We achieved better performance using both features than
by using either one.</p>
      <p>As a result, we achieved the second score among all participants with our
best run.</p>
      <p>C-SIFT GIST LBP SIFT MF-samples</p>
      <p>sixth (ReLU) sixth seventh (ReLU) seventh MF-samples
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
X
In this working note, we described our annotation method for the ImageCLEF
2014 Scalable Concept Image Annotation task. As visual features, we used FV
and deep CNN based feature. Assuming that these two types of features
mutually express di erent kinds of information complementarily, we tried combining
them. In our experiment, we showed how the combination of features contributes
to the improvement of recognition performance. Results show that the
combination of generative features and discriminative features proved e ective in image
recognition tasks.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Caputo</surname>
          </string-name>
          , Henning Muller, Jesus Martinez-Gomez, Mauricio Villegas, Burak Acar, Novi Patricia, Neda Marvasti, Suzan Uskudarl , Roberto Paredes, Miguel Cazorla,
          <article-title>Ismael Garcia-Varea, and Vicente Morell</article-title>
          .
          <source>ImageCLEF</source>
          <year>2014</year>
          :
          <article-title>Overview and analysis of the results</article-title>
          .
          <source>In CLEF proceedings, Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Mauricio</given-names>
            <surname>Villegas</surname>
          </string-name>
          and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Paredes</surname>
          </string-name>
          .
          <article-title>Overview of the ImageCLEF 2014 Scalable Concept Image Annotation Task</article-title>
          .
          <source>In CLEF 2014 Evaluation Labs and Workshop</source>
          , Online Working Notes,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F.</given-names>
            <surname>Perronnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sanchez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Mensink</surname>
          </string-name>
          .
          <article-title>Improving the sher kernel for large-scale image classi cation</article-title>
          .
          <source>European Conference on Computer Vision</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ushiku</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Harada</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kuniyoshi</surname>
          </string-name>
          .
          <article-title>E cient image annotation for automatic sentence generation</article-title>
          .
          <source>The 20th ACM International Conference on Multimedia</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <surname>Geo</surname>
            rey
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>In NIPS</source>
          , Vol.
          <volume>1</volume>
          , p.
          <fpage>4</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Je</given-names>
            <surname>Donahue</surname>
          </string-name>
          , Yangqing Jia, Oriol Vinyals, Judy Ho man, Ning Zhang, Eric Tzeng, and
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Darrell</surname>
          </string-name>
          .
          <article-title>Decaf: A deep convolutional activation feature for generic visual recognition</article-title>
          .
          <source>arXiv preprint arXiv:1310.1531</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>C.</given-names>
            <surname>Fellbaum</surname>
          </string-name>
          .
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . MIT Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>K.</given-names>
            <surname>Crammer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dekel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Keshet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shalev-Shwartz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          . Online Passive{
          <article-title>Aggressive Algorithms</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          , Vol.
          <volume>7</volume>
          , pp.
          <volume>551</volume>
          {
          <issue>585</issue>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>