<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TUB-IRML at MediaEval 2014 Violent Scenes Detection Task: Violence Modeling through Feature Space Partitioning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Esra Acar</string-name>
          <email>esra.acar@tu-berlin.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sahin Albayrak</string-name>
          <email>sahin.albayrak@dai-labor.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DAI Laboratory, Technische Universität Berlin</institution>
          <addr-line>Ernst-Reuter-Platz 7, TEL 14, 10587 Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>This paper describes the participation of the TUB-IRML group to the MediaEval 2014 Violent Scenes Detection (VSD) a ect task. We employ low- and mid-level audio-visual features fused at the decision level. We perform feature space partitioning of training samples through k -means clustering and train a di erent model for each cluster. These models are then used to predict the violence level of videos by employing two-class support vector machines (SVMs) and a classi er selection approach. The experimental results obtained on Hollywood movies and short Web videos show the superiority of mid-level audio features over visual features in terms of discriminative power, and a further enhanced performance resulting from the fusion of audio-visual cues at the decision-level. Finally, the results also demonstrate a performance gain obtained by partitioning the feature space and training multiple models, compared to a unique violence detection model.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The MediaEval 2014 VSD task aims at detecting violent
segments in movies and short Web videos. Detailed
description of the task, the dataset, the ground truth and evaluation
criteria are given in the paper by Sjoberg et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>THE PROPOSED METHOD</title>
    </sec>
    <sec id="sec-3">
      <title>Video Representation</title>
      <p>We represent the audio content using mid-level
representations, whereas the visual content is represented at two
di erent levels: low-level and mid-level.</p>
      <p>
        Mid-level audio representations are based on
MelFrequency Cepstral Coe cient (MFCC) features extracted
from the audio signals of video segments of 0.6 second length.
We experimentally veri ed that the 0.6 second time window
was short enough to be computationally e cient and long
enough to retain su cient relevant information. In order
to generate the mid-level representations, we apply an
abstraction process which uses an MFCC-based Bag-of-Audio
Words approach with a sparse coding scheme. We employ
the dictionary learning technique presented in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In order
to learn the dictionary of size k (k = 1024 in this work)
for sparse coding, 400 k MFCC feature vectors are
sampled from the training data (this 400 k gure was
determined experimentally). In the coding phase, we construct
the sparse representation of audio signals by using the LARS
algorithm [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In order to generate the nal sparse
representation of video segments, which is a set of MFCC feature
vectors, we apply the max-pooling technique.
      </p>
      <p>
        We use motion-related descriptors for the visual
representation of video segments. One of the motion
descriptors is Violent Flow (ViF) which is proposed for real-time
detection of violent crowd behaviors. We compute a ViF
descriptor for each video segment to represent statistics of
ow-vector magnitude changes over time. For a detailed
explanation of the computation of this descriptor, the reader
is referred to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        We also use static content representations. More
speci cally, we employ a ect-related static visual
descriptors. We compute mean and standard deviation of
saturation, brightness and hue in the HSL color space. We also
compute the colorfulness of the keyframe of video segments
using the method in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], where the keyframe is deemed to
be the frame in the middle of a video segment.
      </p>
      <p>
        Mid-level visual representations are based on
histogram of oriented gradient (HoG) and histogram of oriented
optical ow (HoF) features extracted from the visual content
of video segments of 0.6 second length. HoG and HoF
descriptors are densely sampled and computed for subvolumes
of video segments (HoG descriptors are subsampled every 6
frames and HoF descriptors are subsampled every 2 frames
as recommended in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]). The Horn-Schunk method [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is
applied to compute optical ow vectors which are used for
the extraction of HoF descriptors. The resulting HoG and
HoF descriptors are subsequently used to generate mid-level
HoG and HoF representations separately.
2.2
      </p>
      <p>Violence Detection Model
\Violence" is a concept which can audio-visually be
expressed in diverse manners. Therefore, learning multiple
models for the \violence" concept instead of a unique model
constitutes a more judicious choice. This justi es that we
rst perform feature space partitioning by clustering video
segments in our training dataset and learn a di erent model
for each violence sub-concept (i.e., cluster). We use
twoclass SVMs in order to learn violence models (i.e., one SVM
for each sub-concept).</p>
      <p>In the learning step, the main issue is the problem of
imbalanced data. This is caused by the fact that, in the
training dataset, the number of non-violent video shots is
much higher than the number of violent ones. We choose, in
the current framework, to perform random undersampling
to balance the number of violent and non-violent samples
(with a balance ratio of 1:2).</p>
      <p>In the test phase, the main challenge is to combine the
classi cation results of the violence models. We perform
a classi er selection to solve this. More precisely, we rst
determine the nearest cluster to a video segment of the test
set using Euclidean distance measures. Once the classi er
for the video sample is determined, the output of the chosen
model is used as the nal prediction for that video sample.</p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND DISCUSSION</title>
      <p>Table 1 reports the mean Average Precision (MAP2014)
and MAP@100 metrics on the Hollywood movie dataset and
the Web video dataset. We observe that the mid-level audio
representation based on MFCC and sparse coding (Run1 )
provides promising performance and outperforms all other
representations (Run2 and Run3 ) that we use in this work.
We also note that the performance is further improved by
fusing these mid-level audio cues with low- and mid-level
visual cues at the decision level by linear fusion (Run4 ).
However, the fusion of a ect-related color features with the
other audio-visual features does not help improving the
performance (Run5 ).</p>
      <p>The results on the Web video dataset in terms of MAP
metrics even demonstrate superior results compared to the
ones obtained on the Hollywood movie dataset. Therefore,
we can conclude that our violence detection method
generalizes particularly well to other contents, including other
types of video content not used for training the models.
Another interesting observation is that a ect-related color
features seem to provide better results in terms of MAP metrics
on the Web video dataset in comparison to the Hollywood
movie dataset (Run3 ).</p>
      <p>Table 1 also provides a comparison of our methods in
terms of MAP2014 and MAP@100 metrics with an
SVMbased unique violence detection model (i.e., a model where
no feature space partitioning is performed). We can
conclude that our method outperforms the SVM-based
detection method where the feature space is not partitioned and
all violent and non-violent samples are used to build a unique
model.</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we presented an approach based on
feature space partitioning for the detection of violent content
in movies and short Web videos at the video segment level.
We employed low- and mid-level audio-visual features to
represent videos. We showed that the mid-level audio
representation based on MFCC and sparse coding provides
promising performance in terms of MAP2014 and MAP@100
metrics and also outperforms our visual representations. We
also fused these mid-level audio cues with low- and
midlevel visual cues at the decision level using linear fusion for
further improvement and achieved better results than
unimodal video representations in terms of the MAP metrics.
We observed from the overall evaluation results that our
method performs better when violent content is better
expressed in terms of audio features (a typical example would
be a gun shot scene). Hence, as a future work, we need to
extend/improve our visual representation set with more
discriminative representations. Another possibility for future
work is to further investigate the feature space partitioning
concept and optimize the distribution or number of
subconcepts in order to enhance the classi cation performance
of our method.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The research leading to these results has received funding
from the European Community FP7 under grant agreement
number 261743 (NoE VideoSense).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Efron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hastie</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Johnstone</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          .
          <article-title>Least angle regression</article-title>
          .
          <source>The Annals of statistics</source>
          ,
          <volume>32</volume>
          (
          <issue>2</issue>
          ):
          <volume>407</volume>
          {
          <fpage>499</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hasler</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Suesstrunk</surname>
          </string-name>
          .
          <article-title>Measuring colorfulness in natural images</article-title>
          .
          <source>In Electronic Imaging</source>
          <year>2003</year>
          , pages
          <fpage>87</fpage>
          {
          <fpage>95</fpage>
          .
          <string-name>
            <surname>Int</surname>
          </string-name>
          .
          <source>Society for Optics and Photonics</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hassner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Itcher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Kliper-Gross</surname>
          </string-name>
          .
          <article-title>Violent ows: Real-time detection of violent crowd behavior</article-title>
          .
          <source>In Computer Vision and Pattern Recognition Workshops (CVPRW)</source>
          ,
          <year>2012</year>
          IEEE Computer Society Conference on, pages
          <fpage>1</fpage>
          <article-title>{6</article-title>
          . IEEE,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B. K.</given-names>
            <surname>Horn</surname>
          </string-name>
          and
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Schunck</surname>
          </string-name>
          .
          <article-title>Determining optical ow</article-title>
          .
          <source>In 1981 Technical Symposium East</source>
          , pages
          <volume>319</volume>
          {
          <fpage>331</fpage>
          .
          <string-name>
            <surname>Int</surname>
          </string-name>
          .
          <source>Society for Optics and Photonics</source>
          ,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mairal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Sapiro</surname>
          </string-name>
          .
          <article-title>Online learning for matrix factorization and sparse coding</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          ,
          <volume>11</volume>
          :
          <fpage>19</fpage>
          {
          <fpage>60</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sjo</surname>
          </string-name>
          berg,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-G.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. L.</given-names>
            <surname>Quang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.-H.</given-names>
            <surname>Demarty</surname>
          </string-name>
          .
          <article-title>The MediaEval 2014 A ect Task: Violent Scenes Detection</article-title>
          . In MediaEval 2014 Workshop, Barcelona, Spain, October
          <volume>16</volume>
          -17
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Uijlings</surname>
          </string-name>
          , I. Duta,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rostamzadeh</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Sebe</surname>
          </string-name>
          .
          <article-title>Realtime video classi cation using dense hof/hog</article-title>
          .
          <source>In ICMR, page 145. ACM</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>