<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GIBIS at MediaEval 2017: Predicting Media Interestingness Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jurandy Almeida</string-name>
          <email>jurandy.almeida@unifesp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ricardo M. Savii</string-name>
          <email>ricardo.manhaes@unifesp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GIBIS Lab, Institute of Science and Technology, Federal University of São Paulo - UNIFESP 12247-014, São José dos Campos</institution>
          ,
          <addr-line>SP -</addr-line>
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper describes the GIBIS team experience in the Predicting Media Interestingness Task at MediaEval 2017. In this task, the teams were required to develop an approach to predict whether images or videos are interesting or not. Our proposal relies on late fusion with rank aggregation methods for combining ranking models learned with diferent features and by diferent learning-to-rank algorithms.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        In this paper, we explore the use of rank aggregation methods
for predicting the interestingness of images and videos. For that,
content-based representations for images and videos are obtained
by diferent features, which are used to train diferent
learningto-rank algorithms, creating rankers capable of predicting the
interestingness degree of images and videos. Then, the information
provided by diferent pairs of feature-ranker are combined by rank
aggregation methods, yielding more efective predictions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        This work is developed in the context of the MediaEval 2017
Predicting Media Interestingness Task, whose goal is to automatically
select the most interesting frames or portions of videos according
to a common viewer by using features derived from audio-visual
content or associated textual information. Details about data, task,
and evaluation are described in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>PROPOSED APPROACH</title>
      <p>
        The start point for our proposal is the work of Almeida [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], where
motion features were extracted from videos and then used to train
four diferent ranking models, which were combined with a
majority voting strategy [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The key idea exploited in the Almeida’s
work was the use of multiple learning-to-rank algorithms, and their
combination was pointed out as promising.
      </p>
      <p>
        Here, we extend the work of Almeida [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] by exploring rank
aggregation methods for combining ranking models learned with
diferent features and by diferent learning-to-rank algorithms.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Features</title>
      <p>
        Images. For the image subtask, we used only the pre-computed
features provided by the task organizers [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Five low-level features
were considered: Dense SIFT, Histogram of Gradients (HoG), Local
Binary Patterns (LBP), GIST, and Color Histogram. Also, two deep
learning features were used and they refer to Convolutional Neural
Network (CNN) features extracted from the last layers (i.e., fc7 and
prob) of the pre-trained AlexNet model [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Videos. For the video subtask, we used nine pre-computed features
provided by the task organizers [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. One of them represents audio
information: Mel-Frequency Cepstral Coeficients (MFCC). Seven
features are the same used for images and encode visual content:
ifve low-level features (Dense SIFT, HoG, LBP, GIST, and Color
Histogram) and two deep learning features (CNN-fc7 and CNN-prob).
These eight features are frame-based representations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. To obtain
a single video representation, we built a Bag-of-Features (BoF) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
model for each feature. In the BoF framework, visual words [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] are
obtained by quantizing a feature space according to a pre-learned
dictionary. Thus, a video is represented as a normalized frequency
histogram of visual words associated with each feature. In this work,
we construct a codebook of 4000 visual words using a random
selection. In addition, we considered three video-based representations.
One of them is also a pre-computed feature provided by task
organizers, denoted C3D [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The two others refer to additional
visual features we extracted from videos: Histogram of Motion
Patterns (HMP) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Bag-of-Attributes (BoA) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Learning-to-Rank Algorithms</title>
      <p>
        Each of the above features was used as input to train four diferent
learning-to-rank algorithms, which are the same used in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
ifrst three are based on pairwise comparisons: Ranking SVM [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
RankNet [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and RankBoost [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The latter approach considers
lists of objects by using ListNet [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        The SVMr ank package1 [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] was used for running Ranking SVM.
The RankLib package2 was used for running RankNet, RankBoost,
and ListNet. Ranking SVM was configured with a linear kernel. The
others were configured with their default parameter settings.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Rank Aggregation Models</title>
      <p>
        Let C = {o1, o2, . . . , on } be a collection of n objects (i.e., images
or videos). Let R = {r1, r2, . . . , rm } be a set of m feature-ranker
pairs. Let ρj (i) be the interestingness degree assigned by the
featureranker pair rj ∈ R to the object oi ∈ C. Based on the score ρj , a
ranked list τj can be computed. The ranked list τj can be defined as a
permutation of the collection C, which contains the most interesting
objects according to the feature-ranker pair rj . A permutation τj
is a bijection from the set C onto the set [n] = {1, 2, . . . , n}. For a
permutation τj , we interpret τj (i) as the position (or rank) of the
object oi in the ranked list τj . We can say that, if oi is ranked before
ok in the ranked list τj , that is, τj (i) &lt; τj (k), then ρj (i) ≤ ρj (k) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Given the diferent scores ρj and their respective ranked lists τj
computed by distinct pairs rj ∈ R, a rank aggregation method aims
to compute a fused score F (i) to each object oi [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this work, we
used three diferent methods based on score and rank information:
m
(1) Borda Method [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]: F (i) = Í τj (i),
j=0
1https://www.cs.cornell.edu/people/tj/svm_light/svm_rank.html (As of August 2017)
2https://sourceforge.net/p/lemur/wiki/RankLib/ (As of August 2017)
m
(2) Multiplicative Approach [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]: F (i) = Î (1 + ρj (i)),
j=1
m
(3) Weighted Sum Model [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]: F (i) = Í (τj (i) × ρj (i)).
j=1
      </p>
    </sec>
    <sec id="sec-6">
      <title>3 EXPERIMENTS &amp; RESULTS</title>
      <p>Five diferent runs were submitted for each subtask configured
as shown in Table 13. For both subtasks, the first run is the best
feature-ranker pair in isolation and the others refer to the fusion
of the top performing feature-ranker pairs with rank aggregation
methods. All the evaluated approaches were calibrated through a
3-fold cross validation on the development data.</p>
      <p>The development data is composed of 7,396 videos from 78 movie
trailers. For the image subtask, the middle keyframe of each video
was extracted, forming a dataset with 7,396 images. Each of the
features (Section 2.1) was used as input to train each of the
learningto-rank algorithms (Section 2.2). In this way, we obtained 28
featureranker pairs (i.e., 7 features × 4 rankers) for the image subtask
and 44 feature-ranker pairs (i.e., 11 features × 4 rankers) for the
video subtask. Next, each of the feature-ranker pairs was used
to predict the interestingness degree of test images and videos.
Finally, the prediction scores of the top performing feature-ranker
pairs in isolation were combined using rank aggregation methods
(Section 2.3), producing fused prediction scores.</p>
      <p>
        To assess the efectiveness of each approach, we computed the
Mean Average Precision (MAP). For that, we transformed prediction
scores into binary decisions using the strategy proposed in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
First, the prediction scores associated with images and videos of a
same movie trailer were normalized using a z-score normalization.
Then, an empirical threshold of 0.7 was applied to the normalized
prediction scores, producing binary decisions.
      </p>
      <p>Table 2 presents MAP scores obtained for each run on the
development data. For both subtasks, the fusion of the top performing
feature-ranker pairs (runs 2 to 5) performed better than the best
feature-ranker pair in isolation (run 1). The only exception was the
run 2 of the video subtask, which was a required run for the task
where the use of audio features (i.e., MFCC) was mandatory. All
3The run 1 of the image subtask and the run 2 of the video subtask were the required
runs for the task, while the other runs were optional.</p>
      <p>J. Almeida and R. Savii
the machine-learned rankers using MFCC achieved poor results.
By analyzing the confidence intervals, it can be noticed that the
results achieved by the rank aggregation methods seem promising.</p>
    </sec>
    <sec id="sec-7">
      <title>4 CONCLUSIONS</title>
      <p>Our approach has explored rank aggregation methods for
combining feature-ranker pairs. Obtained results demonstrate that the
proposed approach is promising. Future work includes the
investigation of a smarter strategy for selecting the pairs to be combined.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS</title>
      <p>We thank the São Paulo Research Foundation - FAPESP (grant 2016/
06441-7) and the Brazilian National Council for Scientific and
Technological Development - CNPq (grant 423228/2016-1) for funding.
This work has also benefited from the support of the Association
for the Advancement of Afective Computing (AAAC) and the ACM
Special Interest Group on Information Retrieval (SIGIR).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Almeida</surname>
          </string-name>
          .
          <year>2016</year>
          . UNIFESP at MediaEval 2016:
          <article-title>Predicting Media Interestingness Task</article-title>
          .
          <source>In Proc. of the MediaEval 2016 Workshop</source>
          . http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1739</volume>
          /MediaEval_2016_paper_28.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Leite</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Torres</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Comparison of Video Sequences with Histograms of Motion Patterns</article-title>
          .
          <source>In IEEE Intl. Conf. Image Processing (ICIP'11)</source>
          .
          <fpage>3673</fpage>
          -
          <lpage>3676</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Valem</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. C. G.</given-names>
            <surname>Pedronette</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Rank Aggregation Framework for Video Interestingness Prediction</article-title>
          .
          <source>In Intl. Conf. Image Analysis and Processing (ICIAP'17)</source>
          .
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Boureau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bach</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <article-title>LeCun, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Learning MidLevel Features for Recognition</article-title>
          .
          <source>In IEEE Intl. Conf. Computer Vision and Pattern Recognition (CVPR'10)</source>
          .
          <fpage>2559</fpage>
          -
          <lpage>2566</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. J. C.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shaked</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Renshaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lazier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Deeds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. N.</given-names>
            <surname>Hullender</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Learning to rank using gradient descent</article-title>
          .
          <source>In Intl. Conf. Machine Learning (ICML'05)</source>
          .
          <fpage>89</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          , T-Y. Liu,
          <string-name>
            <surname>M-F. Tsai</surname>
            , and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Learning to rank: from pairwise approach to listwise approach</article-title>
          .
          <source>In Intl. Conf. Machine Learning (ICML'07)</source>
          .
          <fpage>129</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>C-H. Demarty</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sjöberg</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
          </string-name>
          ,
          <string-name>
            <surname>T-T. Do</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Gygli</surname>
            , and
            <given-names>N. Q. K.</given-names>
          </string-name>
          <string-name>
            <surname>Duong</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>MediaEval 2017 Predicting Media Interestingness Task</article-title>
          .
          <source>In Proc. of the MediaEval 2017 Workshop</source>
          . Dublin, Ireland.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Duarte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. A. B.</given-names>
            <surname>Penatti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Almeida</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Bag of Attributes for Video Event Retrieval</article-title>
          .
          <source>CoRR abs/1607</source>
          .05208 (
          <year>2016</year>
          ). http://arxiv. org/abs/1607.05208
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Fishburn</surname>
          </string-name>
          .
          <year>1967</year>
          .
          <article-title>Additive Utilities with Incomplete Product Set: Applications to Priorities and Assignments</article-title>
          .
          <source>Operations Research Society of America (ORSA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Freund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Iyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>An Eficient Boosting Algorithm for Combining Preferences</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>4</volume>
          (
          <year>2003</year>
          ),
          <fpage>933</fpage>
          -
          <lpage>969</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Y-G. Jiang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Mei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Rui</surname>
            , and
            <given-names>S-F.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Super Fast Event Recognition in Internet Videos</article-title>
          .
          <source>IEEE Transactions on Multimedia 17</source>
          ,
          <issue>8</issue>
          (
          <year>2015</year>
          ),
          <fpage>1174</fpage>
          -
          <lpage>1186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Training linear SVMs in linear time</article-title>
          .
          <source>In ACM SIGKDD Intl. Conf. Knowledge Discovery and Data Mining (ACM SIGKDD'06)</source>
          .
          <fpage>217</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lam</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Suen</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Application of majority voting to pattern recognition: an analysis of its behavior and performance</article-title>
          .
          <source>IEEE Trans. Systems, Man, and Cybernetics</source>
          , Part A
          <volume>27</volume>
          ,
          <issue>5</issue>
          (
          <year>1997</year>
          ),
          <fpage>553</fpage>
          -
          <lpage>568</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D. C. G.</given-names>
            <surname>Pedronette</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Torres</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Image Re-Ranking and Rank Aggregation based on Similarity of Ranked Lists</article-title>
          .
          <source>Pattern Recognition</source>
          <volume>46</volume>
          ,
          <issue>8</issue>
          (
          <year>2013</year>
          ),
          <fpage>2350</fpage>
          -
          <lpage>2360</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Sivic</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Video Google: A Text Retrieval Approach to Object Matching in Videos</article-title>
          .
          <source>In IEEE Intl. Conf. Computer Vision (ICCV'03)</source>
          .
          <fpage>1470</fpage>
          -
          <lpage>1477</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bourdev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Torresani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning Spatiotemporal Features with 3D Convolutional Networks</article-title>
          .
          <source>In IEEE Intl. Conf. Computer Vision (ICCV'15)</source>
          .
          <fpage>4489</fpage>
          -
          <lpage>4497</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H. P.</given-names>
            <surname>Young</surname>
          </string-name>
          .
          <year>1974</year>
          .
          <article-title>An axiomatization of Borda's rule</article-title>
          .
          <source>Journal of Economic Theory</source>
          <volume>9</volume>
          ,
          <issue>1</issue>
          (
          <year>1974</year>
          ),
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>