<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NII-UIT at MediaEval 2014 Violent Scenes Detection Affect Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vu Lam</string-name>
          <email>lqvu@fit.hcmus.edu.vn</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duy-Dinh Le</string-name>
          <email>ledduy@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sang Phan</string-name>
          <email>plsang@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shin'ichi Satoh</string-name>
          <email>satoh@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Duc Anh Duong</string-name>
          <email>ducda@uit.edu.vn</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of</institution>
          ,
          <addr-line>Informatics, 2-1-2 Hitotsubashi, Chiyoda-ku, Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
          <addr-line>101-8430</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Information, Technology</institution>
          ,
          <addr-line>KM20 Ha Noi highway, Linh, Trung Ward,Thu Duc District, Ho Chi Minh</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Science</institution>
          ,
          <addr-line>227 Nguyen Van Cu, Dist.5, Ho Chi Minh</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>Violent scene detection (VSD) is a challenging problem because of the heterogeneous content, large variations in video quality, and semantic meaning of the concepts. The Violent Scenes Detection Task of MediaEval [1] provides a common dataset and evaluation protocol thus enables a fair comparison of methods. In this paper, we describe our VSD system used in MediaEval 2014 and brie y discuss the performance results obtained in main subjective tasks. In this year, we focus on improving the trajectory-based motion features that have been proven e ective in previous year's evaluation. Besides that, we also adopt SIFT-based and audio features as in last year's system. We combined these features using late fusion. Our results show that the trajectory-based motion features still have very competitive performance and the combination with still image features and audio features can improve overall performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        We consider the Violent Scenes Detection (VSD) task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
as a concept detection task. For evaluation, we use our
NIIKAORI-SECODE framework, which has been achieved good
performances on other benchmarks such as ImageCLEF and
PASCALVOC. Firstly, videos are divided into equal
segments with 5-second length. In each segment, keyframes
are extracted by sampling 5 keyframes per second. For still
image features, local descriptors are extracted and encoded
for all keyframes in each segment and then segment-based
features are formed from their keyframe-based features by
applying average or max pooling. Motion feature and audio
feature are extracted directly from the whole segment. For
all features, we use the popular SVM algorithm for learning.
Finally, the probability output scores of the learned classi er
are used for ranking retrieved segments.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>FEATURE EXTRACTION</title>
      <p>We use features from di erent modalities to test if they are
complementary for violent scenes detection. Currently, we
have developed our VSD system to incorporate still image
feature, motion feature, and audio feature.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Still Image Features</title>
      <p>
        In this year, we use only SIFT-based features for VSD
because they could capture di erent characteristics of
images. We use popular SIFT-based features with both
Hessian Laplace interest points and dense sampling at multiple
scales. Besides the standard SIFT descriptor, we also use
Opponent-SIFT and Color-SIFT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We employ the
bagof-words model with a codebook size of 1000 and the
softassignment technique to generate a xed-dimension feature
representation for each keyframe. Beside encoding the whole
image, we also divide it into grids of 3x1 and 2x2 to encode
spatial information. Finally, in order to generate a single
representation for each segment, we use two pooling
strategies: average pooling and max pooling.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Motion Feature</title>
      <p>
        We use the Improved Trajectories [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to extract dense
trajectories. A combination of Histogram of Oriented
Gradients (HOG), Histogram of Optical Flow (HOF) and
Motion Boundary Histogram (MBH) is used to describe each
trajectory. We encode HOGHOF and MBH features
separately using the Fisher Vector encoding. The codebook size
is 256, trained using a Gaussian Mixture Model (GMM).
The feature representation of each descriptor after applying
PCA has 65,536 dimensions. Finally, these two features are
concatenated to form the nal feature vector with 131,072
dimensions.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Audio Feature</title>
      <p>We use the popular Mel-frequency Cepstral Coe cients
(MFCC) for extracting audio features. We choose a length
of 25ms for audio segments and a step size of 10ms. The
13dimensional MFCC vectors along with each rst and second
derivatives are used for representing each audio segment.
Raw MFCC features are also encoded using Fisher vector
encoding. We use a GMM to train the codebook with 256
clusters. For audio features, we do not use PCA. The nal
feature descriptor has 19,968 dimensions. Our motion and
audio framework are shown in Fig 2.</p>
    </sec>
    <sec id="sec-6">
      <title>CLASSIFICATION</title>
      <p>
        LibSVM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used for training and testing at segment
level. To generate training data, segments of which at least
80% are marked as violent according to the ground truth.
Extracted features are scaled to [0; 1] using the SVM-scale
tool of LibSVM. The remaining segments are considered as
negative. For still image features, we use a chi-square kernel
to calculate the distance matrix. For audio and motion
features, which are encoded using Fisher vector, a linear kernel
is used. The optimal gamma and cost parameters for
learning SVM classi ers are found by conducting a grid search
with 5-fold cross validation on the training dataset.
      </p>
    </sec>
    <sec id="sec-7">
      <title>SUBMITTED RUNS</title>
      <p>We select two training sets: (A) uses 14 videos, (B) uses
24 videos. We use the VSD 2013 test dataset (7 videos) as
validation set. We employ a simple late fusion strategy on
the above features, using equal weights and learnt weights.
We submitted ve runs in total (Fig 1): (R1) using training
set A, we rst select the best still image feature and fuse it
with motion and audio features; (R2) using training set B,
we fuse all still image features with motion and audio using
6.
equal weight; (R3) same as R1 but using training set B;
(R4) using training set B, we fuse all still image features with
motion and audio using learnt fusion weights from validation
set; (R5) using training set B, we fuse motion and audio
features with equal weights.
5.</p>
    </sec>
    <sec id="sec-8">
      <title>RESULTS AND DISCUSSIONS</title>
      <p>The detailed performance for each submitted run is shown
in Figure 3. Our best run is the fusion run of best single
still image features (RGBSIFT), motion and audio features
(R1). There is not a big gap among submitted runs. We see
that, the performance of motion features with Fisher vector
encoding is alway good and signi cantly better than others.
In all submitted runs, we used motion features as a base
to fuse with others. Audio and still image features did not
achieve good performance, but they can be complementary
to motion features. Another interesting observation is that
runs trained on fewer videos (training set A - 14 videos) have
better performance than the runs in which set (24 videos)
was used. This indicates that the second training set might
contain ambiguous violent scene's annotations, which harms
the detection performance.</p>
    </sec>
    <sec id="sec-9">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This research is partially funded by Vietnam National
University Ho Chi Minh City (VNU-HCM) under grant
number B2013-26-01.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sjo</surname>
          </string-name>
          berg,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Quang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Demarty</surname>
          </string-name>
          .
          <article-title>The MediaEval 2014 A ect Task: Violent Scenes Detection</article-title>
          . In MediaEval 2014 Workshop, Barcelona, Spain, October
          <volume>16</volume>
          -17
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>K. Van de Sande</surname>
            , T. Gevers,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Snoek</surname>
          </string-name>
          ,
          <article-title>"Evaluating Color Descriptors for Object and Scene Recognition,"</article-title>
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on , vol.
          <volume>32</volume>
          , no.
          <issue>9</issue>
          , pp.
          <volume>1582</volume>
          ,
          <issue>1596</issue>
          ,
          <string-name>
            <surname>Sept</surname>
          </string-name>
          . 2010
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <article-title>Action Recognition with Improved Trajectories</article-title>
          .
          <source>In Proceedings of the 2013 IEEE International Conference on Computer Vision</source>
          (ICCV '
          <article-title>13)</article-title>
          . IEEE Computer Society, Washington, DC, USA,
          <fpage>3551</fpage>
          -
          <lpage>3558</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.-J.</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>LIBSVM : a library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>2</volume>
          :
          <issue>27</issue>
          :1{
          <fpage>27</fpage>
          :
          <fpage>27</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>