<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Media Interestingness via Biased Discriminant Embedding and Supervised Manifold Regression</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yang Liu</string-name>
          <email>csygliu@comp.hkbu.edu.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhonglei Gu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobey H. Ko</string-name>
          <email>tobeyko@hku.hk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Hong Kong Baptist University</institution>
          ,
          <addr-line>HKSAR</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Industrial and Manufacturing Systems Engineering, University of Hong Kong</institution>
          ,
          <addr-line>HKSAR</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Research and Continuing Education, Hong Kong Baptist University</institution>
          ,
          <addr-line>Shenzhen</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>In this paper, we describe our model designed for automatic prediction of media interestingness. Specifically, a two-stage learning framework is proposed. In the first stage, supervised dimensionality reduction is employed to discover the key discriminant information embedded in the original feature space. We present a new algorithm dubbed biased discriminant embedding (BDE) to extract discriminant features with discrete labels and use supervised manifold regression (SMR) to extract discriminant features with continuous labels. In the second stage, SVM is utilized for prediction. Experimental results validate the efectiveness of our approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>Manifold</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Predicting the interestingness of multimedia content has long
been studied in the psychology community [
        <xref ref-type="bibr" rid="ref1 ref6 ref7">1, 6, 7</xref>
        ]. More
recently, we witness an explosion of multimedia content due
to the accessibility of low cost multimedia creation tools, the
automatic prediction of media interestingness thus started
to attract attention in the computer science community
because of its many useful applications to content providers,
marketing, and managerial decision-makers.
      </p>
      <p>
        In this paper, we propose to use dimensionality reduction
to extract low-dimensional features for MediaEval 2017
Predicting Media Interestingness Task. Specifically, we propose a
new algorithm called biased discriminant embedding (BDE)
for discrete labels and utilize supervised manifold regression
(SMR) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for continuous labels.
2.1
      </p>
      <p>DIMENSIONALITY REDUCTION</p>
      <p>Biased Discriminant Embedding
Given the data matrix X = [x1, x2, ..., x], where x ∈ R
denotes the feature vector of the -th image or video, and
label vector l = [1, 2, ..., ], where  ∈ {0, 1} denotes the
corresponding label of x, with 1 for interesting and 0 for
non-interesting, biased discriminant embedding (BDE) aims
to learn a  ×  transformation matrix W, which maximizes
the biased discriminant information in the reduced subspace.
The motivation for proposing the biased discrimination is that
in media interestingness prediction, we are probably more
interested in the interesting class than the non-interesting
Copyright held by the owner/author(s).
W = arg min ∑︁</p>
      <p>W</p>
      <p>,=1
interestingness level of x and that of x.
where  = | − | measures the similarity between the</p>
      <p>For each high-dimensional data point x, we can obtain
its low-dimensional representation by y =
apply SVM to y for interestingness prediction.</p>
      <p>W x. Then we
3</p>
    </sec>
    <sec id="sec-2">
      <title>EXPERIMENTS</title>
      <p>
        For each image data sample, we construct a 1299-D feature
vector by selecting features from the feature set provided
by the task organizers, including 128-D color histogram
features, 300-D denseSIFT features, 512-D gist features, 300-D
hog2× 2, and 59-D LBP features. For the video data, we treat
each frame as a separate image, and calculate the average
and standard deviation over all frames in this shot, and thus
we have a 2598-D feature set for each video. We
normalize each dimension of the training data to the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]
before dimensionality reduction, where
− 
by ˆ = − 
 and  denote the minimum and maximum values
in the corresponding dimension, respectively. Details about
the dataset description can be found in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>For Run 1 of image data, we use the normalized 1299-D
feature vector as the input of SVM. For Runs 2-5 of image
data, we reduce the original data to the 23-D, 25-D, 26-D,
27-D subspaces via BDE (for discrete labels) and SMR (for
continuous labels), respectively. For Run 1 of video data, we
one. The objective function of BDE is given as follows:
W = arg max</p>
      <p>W
︃(</p>
      <p>W SW )︃
notes the biased within-class scatter, S = ∑︀
x)(x − x)
de,=1( ×
| − |)(x − x)(x −
scatter, and  = (−|| x −
x)
 denotes the biased between-class
x||2/2 ) measures the
closeness between two data samples x and x. The optimization
problem could be solved by generalized eigen-decomposition.
2.2</p>
      <p>
        Supervised Manifold Regression
Supervised manifold regression (SMR) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] aims to find the
latent subspace, where two data points should be close to
each other if they possess similar interestingness levels. The
objective function of SMR is given as follows:
‖W (x − x)‖2 · (︀   + (1 −  )︀) ,
(1)
(2)
(a) BDE on image data
(b) SMR on image data
(c) BDE on video data
(d) SMR on video data
      </p>
      <p>Run 1
Run 2
Run 3
Run 4
Run 5</p>
      <p>Images
MAP@10
0.1184
0.132
0.1332
0.1315
0.1369</p>
      <p>MAP
0.2812
0.2916
0.2898
0.2884
0.291</p>
      <p>Videos
MAP@10
0.0556
0.0468
0.0468
0.0463
0.0445</p>
      <p>
        MAP
0.1813
0.1761
0.1761
0.1742
0.1746
use the normalized 2598-D feature vector as the input of
SVM. For Runs 2-5 of video data, we reduce the original
data to the 23-D, 25-D, 26-D, 27-D subspaces via BDE (for
discrete labels) and SMR (for continuous labels), respectively.
To predict the binary interestingness labels, we use  -SVC [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
with an RBF kernel. We set  = 0.1 and  = 100 (for
image data)/64 (for video data). To predict the continuous
interestingness level, we use  -SVR [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] with an RBF kernel.
We set  = 1,  = 0.01, and  = 1/. Table 1 reports
the evaluation results of the proposed model provided by
the task organizers. For image data, the reduced features
perform better than the original ones, which indicates that
the subspaces learned by BDE and SMR capture important
information in terms of media interestingness. For video
data, the performance of reduced features is slightly worse
than that of the original ones. The reason might be that
video data are more complex than image data so that such a
low-dimensional representation cannot fully capture the key
discriminant information embedded in the original space.
      </p>
      <p>We further analyze the contribution of each dimension
in the original feature space. The contribution of the -th
dimension is defined as  = ∑︀  | |, where
  denotes the -th eigenvalue,  denotes the (, )-th
element of W, and | · | denotes the absolute value operator.
From Figures 1(a) and 1(c), we can observe that color
histogram and LBP features contribute more than the others
while the GIST features contribute the least in the discrete
prediction task. In continuous prediction (Figures 1(b) and
1(d)), the color histogram and GIST features contribute the
most among the five feature sets.
4</p>
      <p>DISCUSSION AND OUTLOOK
This paper introduces our model designed for media
interestingness prediction. For the future work, we aim to improve
the performance of video interestingness prediction by
incorporating the video temporal information. Moreover, as
the ground truth (labels) of interestingness are provided by
human beings, they generally vary with each individual and
are somewhat subjective. We are therefore particularly
interested in refining the human labeled ground truth (especially
for continuous case) via machine learning technologies.</p>
    </sec>
    <sec id="sec-3">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported in part by the National Natural
Science Foundation of China under Grant 61503317, and in
part by the Faculty Research Grant of Hong Kong Baptist
University (HKBU) under Project FRG2/16-17/032.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Daniel</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Berlyne</surname>
          </string-name>
          .
          <year>1960</year>
          .
          <article-title>Conflict, arousal and curiosity</article-title>
          .
          <source>McGraw-Hill.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chih-Chung Chang</surname>
          </string-name>
          and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          <volume>2</volume>
          (
          <year>2011</year>
          ),
          <volume>27</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>27</fpage>
          . Issue 3.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sjoberg</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            , T.-T. Do,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Gygli</surname>
            , and
            <given-names>N. Q. K</given-names>
          </string-name>
          <string-name>
            <surname>Duong. MediaEval 2017 Predicting Media</surname>
          </string-name>
          <article-title>Interestingness Task</article-title>
          .
          <source>In Proc. of the MediaEval 2017 Workshop</source>
          . Dublin, Ireland, Sept.
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.-M.</given-names>
            <surname>Cheung</surname>
          </string-name>
          .
          <article-title>Supervised Manifold Learning for Media Interestingness Prediction</article-title>
          .
          <source>In Proc. of the MediaEval 2016 Workshop</source>
          . Hilversum, Netherlands, Oct.
          <volume>20</volume>
          -
          <fpage>21</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Bernhard</given-names>
            <surname>Sch</surname>
          </string-name>
          ¨olkopf, Alex J.
          <string-name>
            <surname>Smola</surname>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Williamson</surname>
          </string-name>
          , and
          <string-name>
            <surname>Peter L. Bartlett</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>New Support Vector Algorithms</article-title>
          .
          <source>Neural Comput</source>
          .
          <volume>12</volume>
          ,
          <issue>5</issue>
          (
          <year>2000</year>
          ),
          <fpage>1207</fpage>
          -
          <lpage>1245</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Paul J.</given-names>
            <surname>Silvia</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Exploring the psychology of interest</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Craig</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>Phoebe</given-names>
            <surname>Ellsworth</surname>
          </string-name>
          .
          <year>1985</year>
          .
          <article-title>Patterns of cognitive appraisal in emotion</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>48</volume>
          ,
          <issue>4</issue>
          (
          <year>1985</year>
          ),
          <fpage>813</fpage>
          -
          <lpage>838</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>