<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MediaEval 2016 Predicting Media Interestingness Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claire-Hélène Demarty</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mats Sjöberg</string-name>
          <email>mats.sjoberg@helsinki</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Ionescu</string-name>
          <email>bionescu@imag.pub.ro</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thanh-Toan Do</string-name>
          <email>do@sutd.edu.sg</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanli Wang</string-name>
          <email>hanliwang@tongji.edu.cn</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ngoc Q. K. Duong</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Frédéric Lefebvre</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Technicolor</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rennes</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HIIT, University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LAPI, University Politehnica of Bucharest</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Singapore University of Technology and Design, Singapore &amp; University of Science</institution>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Tongji University</institution>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>This paper provides an overview of the Predicting Media Interestingness task that is organized as part of the MediaEval 2016 Benchmarking Initiative for Multimedia Evaluation. The task, which is running for the rst year, expects participants to create systems that automatically select images and video segments that are considered to be the most interesting for a common viewer. In this paper, we present the task use case and challenges, the proposed data set and ground truth, the required participant runs and the evaluation metrics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The ability of multimedia data to attract and keep
people's interest for long periods of time is gaining more and
more importance in the eld of multimedia, where concepts
such as memorability [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], aesthetics [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], interestingness [
        <xref ref-type="bibr" rid="ref11 ref13">13,
11</xref>
        ], attractiveness [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a ective value [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], are intensely
studied, especially in the context of the ever growing
market value of social media and advertising. In particular,
although interestingness has been studied for a long time in
the psychology community [
        <xref ref-type="bibr" rid="ref2 ref21 ref22">21, 2, 22</xref>
        ] and more recently but
actively in the image processing community [
        <xref ref-type="bibr" rid="ref1 ref10 ref18 ref23 ref5 ref6">23, 5, 1, 6, 10,
18</xref>
        ], no common de nition exists in the literature. Moreover
datasets publicly available are only a few and no benchmark
exists for the evaluation of what makes a media interesting.
      </p>
      <p>In this paper we introduce the 2016 MediaEval1
Predicting Media Interestingness Task, which is a pioneer
benchmarking initiative for automatic prediction of image and
video interestingness. The task which is in its rst year
derives from a practical use case at Technicolor2. It
involves helping professionals to illustrate a Video on Demand
(VOD) web site by selecting some interesting frames and/or
video excerpts for the posted movies. The frames and
excerpts should be suitable in terms of helping a user to make
his/her decision about whether he/she is interested in
watching the underlying movie. The data in this task is therefore
adapted to this particular context which provides a more
focused de nition for interestingness.</p>
      <sec id="sec-1-1">
        <title>1http://multimediaeval.org/ 2http://www.technicolor.com/</title>
        <p>2.</p>
        <p>The task requires participants to deploy algorithms that
automatically select images and video segments of
Hollywoodlike movies which are considered to be the most interesting
for a common viewer. Interestingness of the media is judged
based on the visual appearance, audio information and text
accompanying the data. Therefore, the multimodal facet of
interestingness prediction can be investigated.</p>
        <p>Two di erent subtasks are provided, which correspond to
the two types of available media content, namely:
predicting image interestingness | given a set of
keyframes extracted from a movie, the task requires to
automatically identify those images for the given movie
that viewers report to be the most interesting. To solve
the task, participants can make use of visual content
as well as external metadata, e.g., Internet data about
the movie, social media information, etc;
predicting video interestingness | given the video shots
of a movie, the task requires to automatically identify
those shots that viewers report to be the most
interesting in the given movie. To solve the task, participants
can make use of visual and audio data as well as
external data, e.g., subtitles, Internet data, etc.</p>
        <p>In both cases, the task is a binary classi cation task, thus
participants are expected to label the provided data as being
interesting or not (Note that prediction will be carried out on
a per movie basis). However, a con dence value is required
for the provided prediction.
3.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>DATA DESCRIPTION</title>
      <p>The 2016 data is extracted from Creative Commons
licensed trailers of Hollywood-like movies. It consists of a
development data intended for designing and training the
methods (information is extracted from 52 trailers) and a
testing data which is used for the nal benchmarking (with
information from 26 trailers). The choice of using trailers
instead of full movies is driven by the need to nd data,
both freely distributable and still representative in content
and quality of Hollywood movies. Trailers are the results of
some manual ltering of movies to keep interesting scenes,
but also less attractive shots to balance their content. We
therefore believe they are still representative for the task.</p>
      <p>For the predicting video interestingness subtask, the data
consists of the video shots obtained after the manual
segmentation of the videos (video shots are the continuous
frame sequences recorded between a camera turn on and
o ), i.e., 5,054 shots for the development data, and 2,342
shots for the test data.</p>
      <p>For the predicting image interestingness subtask, the data
consists of collections of key-frames extracted from the video
shots used for the video subtask. One single key-frame is
extracted per shot, therefore leading to 5,054 key-frames for
the development set and 2,342 for the test set. This single
key-frame is chosen as the middle frame, as it is highly likely
to capture the most important information of the shot.</p>
      <p>
        To facilitate participation from various communities, we
also provide some pre-computed content descriptors, namely:
low level features | dense SIFT (Scale Invariant Feature
Transform) which are computed following the original work
in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], except that the local frame patches are densely
sampled instead of using interest point detectors. A codebook
of 300 codewords is used in the quantization process with a
spatial pyramid of three layers [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]; HoG descriptors
(Histograms of Oriented Gradients) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] are computed over densely
sampled patches. Following [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], HoG descriptors in a 2
2 neighborhood are concatenated to form a descriptor of
higher dimension; LBP (Local Binary Patterns) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]; GIST
are computed based on the output energy of several
Gaborlike lters (8 orientations and 4 scales) over a dense frame
grid like in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]; color histogram computed in the HSV space
(Hue-Saturation-Value); MFCC (Mel-Frequency Cepstral
Coe cients) computed over 32ms time-windows with 50%
overlap. The cepstral vectors are concatenated with their rst
and second derivatives; fc7 layer (4,096 dimensions) and
prob layer (1,000 dimensions) of AlexNet [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; mid level face
detection and tracking related features 3 | obtained by face
tracking-by-detection in each video shot with a HoG
detector [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and the correlation tracker proposed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
4.
      </p>
    </sec>
    <sec id="sec-3">
      <title>GROUND TRUTH</title>
      <p>All data was manually annotated in terms of
interestingness by human assessors. A dedicated web-based tool was
developed to assist the annotation process. Overall, more
than 312 annotators participated to the annotation for the
video data and 100 for the images. The cultural distribution
is over 29 di erent countries in the world.</p>
      <p>
        We use a pair-wise comparison protocol [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] where
annotators are provided with a pair of images/shots at a time
and asked to tag which of the content is more interesting
for them. The process is repeated by scanning the whole
dataset. As an exhaustive comparison of all possible pairs is
basically impossible due to the required human resources, a
boosting selection was used instead, i.e., a modi ed version
of the adaptive square design method [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], for which several
annotators participate to each iteration.
      </p>
      <p>
        To achieve the nal ground truth, pair-based annotations
are aggregated with the Bradley-Terry-Luce (BTL) model
computation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] resulting in an interestingness degree for
each image/shot. The nal binary decisions are obtained
after the following processing steps: (i) the interestingness
values are ranked in increasing order and normalized
between 0 and 1; (ii) the resulting curve is smoothed with a
short averaging window, and the second derivative is
computed; (iii) for both shots and images, and for all videos,
a threshold empirically set to 0.01 is applied on the second
derivative to nd the rst point whose value is above the
3http://multimediaeval.org/mediaeval2016/
persondiscovery/
threshold. This position corresponds to the limit between
non interesting and interesting shots/images. The
underlying motivation for this empirical rule is the following: the
non interesting population has rather similar interestingness
values which increase slowly, while a gap happens when one
switches from this non interesting population to the
population of more interesting samples. The second derivative
was chosen preferably to the rst derivative, as it allowed to
select those gaps more precisely.
      </p>
      <p>Ground truth is provided in binary format, i.e., 1 for
interesting and 0 for non interesting, for each image and video
in the two subtasks.
5.</p>
    </sec>
    <sec id="sec-4">
      <title>RUN DESCRIPTION</title>
      <p>Each participating team is expected to submit up to 5
runs for both subtasks altogether. Among these 5 runs, two
runs are required, one per subtask: for the predicting image
interestingness subtask, the required run is built on visual
information only and no external data is allowed; for the
predicting video interestingness subtask only audio and visual
information is allowed (no external data) for the required
run.</p>
      <p>Note that in this context, external data can be understood
as: (i) additional datasets and annotations dedicated to
interestingness classi cation; (ii) pre-trained models, features,
detectors obtained from such dedicated additional datasets;
and (iii) additional metadata that could be found on the
Internet on the provided content (e.g., from IMDB4).</p>
      <p>On the contrary, CNN features trained on generic datasets
such as ImageNet (typically the provided CNN features) are
allowed for use in the required runs. By generic datasets, we
mean datasets that were not designed to support research in
the task area, i.e., for the classi cation/study of image and
video interestingness.
6.</p>
    </sec>
    <sec id="sec-5">
      <title>EVALUATION</title>
      <p>The o cial evaluation metric is the mean average
precision (MAP) over the interesting class, i.e., the mean over
the average precision scores computed for each trailer. This
metric, adapted to retrieval tasks, ts perfectly the chosen
use case in which we want to help a user choose between
different samples by providing him a list of suggestions, ranked
according to interestingness. For assessing the performance,
we use the trec_eval tool provided by NIST5. In addition
to MAP, other commonly used metrics such as precision and
recall will be provided to participants.
7.</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS</title>
      <p>The 2016 Predicting Media Interestingness task provides
participants with a comparative and collaborative
evaluation framework for predicting content interestingness with
explicit focus on multimedia approaches. Details on the
methods and results of each individual participant team can
be found in the working note papers of the MediaEval 2016
workshop proceedings.
8.</p>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank Yu-Gang Jiang and Baohan Xu
from the Fudan University, China, and Herve Bredin, from
LIMSI, France for providing the features that accompany the
released data, and Alexey Ozerov and Vincent Demoulin for
their valuable inputs to the task de nition.</p>
      <sec id="sec-7-1">
        <title>4http://www.imdb.com/ 5http://trec.nist.gov/trec\_eval/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>X.</given-names>
            <surname>Amengual</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosch</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. L. de la Rosa</surname>
          </string-name>
          .
          <article-title>Review of Methods to Predict Social Image Interestingness and Memorability</article-title>
          , pages
          <volume>64</volume>
          {
          <fpage>76</fpage>
          . Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Berlyne</surname>
          </string-name>
          .
          <article-title>Con ict, arousal and curiosity</article-title>
          .
          <source>Mc-Graw-Hill</source>
          ,
          <year>1960</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bradley</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Terry</surname>
          </string-name>
          .
          <article-title>Rank analysis of incomplete block designs: the method of paired comparisons</article-title>
          .
          <source>Biometrika</source>
          , (
          <volume>39</volume>
          (
          <issue>3-4</issue>
          )):
          <volume>324</volume>
          {
          <fpage>345</fpage>
          ,
          <year>1952</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bylinskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Isola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bainbridge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <article-title>Intrinsic and extrinsic e ects on image memorability</article-title>
          .
          <source>Vision research</source>
          , (
          <volume>116</volume>
          ):
          <volume>165</volume>
          {
          <fpage>178</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Chamaret</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Demoulin</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Marquant.</surname>
          </string-name>
          <article-title>Experiencing the interestingness concept within and between pictures</article-title>
          .
          <source>In Proceeding of SPIE, Human Vision and Electronic Imaging</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Fedorovskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Quek</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Snyder.</surname>
          </string-name>
          <article-title>The e ect of familiarity on perceived interestingness of images</article-title>
          . volume
          <volume>8651</volume>
          , pages
          <fpage>86511C</fpage>
          {
          <year>86511C</year>
          {
          <fpage>12</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Dalal</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Triggs</surname>
          </string-name>
          .
          <article-title>Histograms of oriented gradients for human detection</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Danelljan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hager</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Khan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Felsberg</surname>
          </string-name>
          .
          <article-title>Accurate scale estimation for robust visual tracking</article-title>
          .
          <source>In British Machine Vision Conference</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ordonez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Berg</surname>
          </string-name>
          .
          <article-title>High level describable attributes for predicting aesthetics and interestingness</article-title>
          .
          <source>In IEEE International Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Grabner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nater</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Druey</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. Van</given-names>
            <surname>Gool</surname>
          </string-name>
          .
          <article-title>Visual interestingness in image sequences</article-title>
          .
          <source>In Proceedings of the 21st ACM International Conference on Multimedia</source>
          , pages
          <volume>1017</volume>
          {
          <fpage>1026</fpage>
          , New York, NY, USA,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gygli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Grabner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Riemenschneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nater</surname>
          </string-name>
          , and
          <string-name>
            <surname>L. van Gool.</surname>
          </string-name>
          <article-title>The interestingness of images</article-title>
          .
          <source>In ICCV International Conference on Computer Vision</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.-G.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rui</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.-F.</given-names>
            <surname>Chang</surname>
          </string-name>
          .
          <article-title>Super fast event recognition in internet videos</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          ,
          <volume>177</volume>
          (
          <issue>8</issue>
          ):1{
          <fpage>13</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.-G.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Yan</surname>
          </string-name>
          .
          <article-title>Understanding and predicting interestingness of videos</article-title>
          .
          <source>In AAAI Conference on Arti cial Intelligence</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kalayci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. K.</given-names>
            <surname>Ekenel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Gunes</surname>
          </string-name>
          .
          <article-title>Automatic analysis of facial attractiveness from video</article-title>
          .
          <source>In IEEE ICIP International Conference on Image Processing</source>
          , pages
          <volume>4191</volume>
          {
          <fpage>4195</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <article-title>Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <volume>2169</volume>
          {
          <fpage>2178</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Barkowsky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Callet</surname>
          </string-name>
          .
          <article-title>Boosting paired comparison methodology in measuring visual discomfort of 3dtv: performances of three di erent designs</article-title>
          .
          <source>In SPIE Electronic Imaging, Stereoscopic Displays and Applications</source>
          , volume
          <volume>8648</volume>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International Journal on Computer Vision</source>
          , (
          <volume>60</volume>
          ):
          <volume>91</volume>
          {
          <fpage>110</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Machajdik</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          .
          <article-title>A ective image classi cation using features inspired by psychology and art theory</article-title>
          .
          <source>In Proceedings of the 18th ACM International Conference on Multimedia</source>
          , pages
          <volume>83</volume>
          {
          <fpage>92</fpage>
          , New York, NY, USA,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pietikainen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Maenpaa</surname>
          </string-name>
          .
          <article-title>Multiresolution gray-scale and rotation invariant texture classi cation with local binary patterns</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>,</volume>
          (
          <volume>24</volume>
          (
          <issue>7</issue>
          )):
          <volume>971</volume>
          {
          <fpage>987</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          .
          <article-title>Modeling the shape of the scene: a holistic representation of the spatial envelope</article-title>
          .
          <source>International Journal of Computer Vision</source>
          , (
          <volume>42</volume>
          ):
          <volume>145</volume>
          {
          <fpage>175</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Silvia</surname>
          </string-name>
          .
          <article-title>Exploring the psychology of interest</article-title>
          . Oxford University Press,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C.</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Ellsworth</surname>
          </string-name>
          .
          <article-title>Patterns of cognitive appraisal in emotion</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          ,
          <volume>48</volume>
          (
          <issue>4</issue>
          ):
          <volume>813</volume>
          {
          <fpage>838</fpage>
          ,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          .
          <article-title>The quest for visual interest</article-title>
          .
          <source>In Proceedings of the 23rd ACM International Conference on Multimedia</source>
          , pages
          <volume>919</volume>
          {
          <fpage>922</fpage>
          , New York, NY, USA,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ehinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          .
          <article-title>Sun database: Large-scale scene recognition from abbey to zoo</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <volume>3485</volume>
          {
          <fpage>3492</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Yazdani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Skodras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Fakotakis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Ebrahimi</surname>
          </string-name>
          .
          <article-title>Multimedia content analysis for emotional characterization of music video clips</article-title>
          .
          <source>EURASIP Journal on Image and Video Processing</source>
          , (
          <volume>26</volume>
          ),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>