<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MediaEval 2017 Predicting Media Interestingness Task</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Claire-Hélène Demarty</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Science and Helsinki Institute for Information Technology HIIT, University of Helsinki</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ETH Zurich, Switzerland &amp; Gifs.com</institution>
          ,
          <country country="US">US</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>LAPI, University Politehnica of Bucharest</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Technicolor</institution>
          ,
          <addr-line>Rennes</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Adelaide</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>In this paper, the Predicting Media Interestingness task which is running for the second year as part of the MediaEval 2017 Benchmarking Initiative for Multimedia Evaluation, is presented. For the task, participants are expected to create systems that automatically select images and video segments that are considered to be the most interesting for a common viewer. All task characteristics are described, namely the task use case and challenges, the released data set and ground truth, the required participant runs and the evaluation metrics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Predicting the interestingness of media content has been an
active area of research in the computer vision community for several
years now [
        <xref ref-type="bibr" rid="ref1 ref10 ref7 ref8">1, 7, 8, 10</xref>
        ] and it has even been studied earlier in the
psychological community [
        <xref ref-type="bibr" rid="ref16 ref17 ref2">2, 16, 17</xref>
        ]. However, there were multiple
competing definitions of interestingness, only a few publicly
available datasets, and until last year, no public benchmark existed to
assess the interestingness of content. In 2016, a task for the
Prediction of Media Interestingness was proposed in the MediaEval 2016
Benchmarking Initiative for Multimedia Evaluation. This task was
also an opportunity to propose a clear definition of interestingness,
compatible with a real-world industry use case at Technicolor1.
The 2017 edition of the MediaEval benchmark includes a follow-up
of the Predicting Media Interestingness Task. This paper gives an
overview of the task description in its second year, together with
a description of the data and ground truth. The required runs and
chosen evaluation metrics are also detailed. In all cases, changes in
this year’s edition are highlighted compared to last year’s edition.
      </p>
    </sec>
    <sec id="sec-2">
      <title>TASK DESCRIPTION</title>
      <p>The Predicting Media Interestingness Task was proposed for the
ifrst time last year. This year’s edition is a follow-up which builds
incrementally upon the previous experience. The task requires
participants to automatically select images and/or video segments
that are considered to be the most interesting for a common viewer.
Interestingness of media is to be judged based on visual appearance,
audio information and text accompanying the data, including movie
metadata. To solve the task, participants are strongly encouraged
to deploy multimodal approaches.</p>
      <p>As in 2016, interestingness should be assessed according to a
practical use case at Technicolor, which involves helping
professionals to illustrate a Video on Demand (VOD) web site by selecting
some interesting frames and/or video excerpts for the movies. The
frames and excerpts should be suitable in terms of helping a user
to make his/her decision about whether he/she is interested in
watching the whole movie. Once again, two subtasks are be ofered
to participants, which correspond to two types of available media
content, namely images and videos. Participants are encouraged to
submit to both subtasks. In both cases, the task will be considered
as a binary classification and a ranking task. Prediction will be
carried out on a per movie basis. The two taskes are:</p>
      <p>Predicting Image Interestingness Given a set of key-frames
extracted from a certain movie, the task involves automatically (1)
identifying those images that viewers report to be interesting and
(2) ranking all images according to their level of interestingness.
To solve the task, participants can make use of visual content as
well as accompanying metadata, e.g., Internet data about the movie,
social media information, etc.</p>
      <p>Predicting Video Interestingness Given a set of video
segments extracted from a certain movie, the task involves
automatically (1) identifying the segments that viewers report to be
interesting and (2) ranking all segments according to their level of
interestingness. To solve the task, participants can make use of
visual and audio data as well as accompanying metadata, e.g.,
subtitles, Internet data about the movie, etc.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DATA DESCRIPTION</title>
      <p>The data is extracted from Creative Commons licensed
Hollywoodlike videos: 103 movie trailers and 4 continuous extracts of ca. 15min
from full-length movies. For the video interestingness subtask, the
data consists of video segments obtained after a manual
segmentation. These segments correspond to shots (video shots are the
continuous frame sequences recorded between the camera being
turned on and being turned of) for all videos but four. Their average
duration is of one second. The four last videos, which correspond
to the full-length movie extracts cited above, were manually
segmented into longer segments (243) with an average duration of
11.4s, to better take into account a certain unity of meaning and
the audio information of the resulting segments. For the image
subtask, the data consists of collections of key-frames extracted
from the video segments used for the video subtask (one key-frame
per segment). This will allow the comparison of results from both
subtasks. The extracted key-frame corresponds to the frame in the
middle of each video segment. In total, 7,396 video segments and
7,396 key-frames are released in the development set, whereas the
test set consists of 2435 video segments and the same number of
key-frames.</p>
      <p>
        To facilitate participation from various communities, we also
provide some pre-computed content descriptors, namely: low level
features — dense SIFT (Scale Invariant Feature Transform) which are
computed following the original work in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], except that the local
frame patches are densely sampled instead of using interest point
detectors. A codebook of 300 codewords is used in the quantization
process with a spatial pyramid of three layers [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]; HoG descriptors
(Histograms of Oriented Gradients) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] are computed over densely
sampled patches. Following [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], HoG descriptors in a 2 × 2
neighborhood are concatenated to form a descriptor of higher dimension;
LBP (Local Binary Patterns) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]; GIST are computed based on the
output energy of several Gabor-like filters (8 orientations and 4
scales) over a dense frame grid like in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]; color histogram computed
in the HSV space (Hue-Saturation-Value); MFCC (Mel-Frequency
Cepstral Coeficients) computed over 32ms time-windows with
50% overlap. The cepstral vectors are concatenated with their first
and second derivatives; fc7 layer (4,096 dimensions) and prob layer
(1,000 dimensions) of AlexNet [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]; mid level face detection and
tracking related features2 — obtained by face tracking-by-detection in
each video shot with a HoG detector [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and the correlation tracker
proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In addition to these frame-based features, we
provide C3D [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] features, which were extracted from fc6 layer (4,096
dimensions) and averaged on a segment level.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>GROUND TRUTH</title>
      <p>
        Both video and image data was manually and independently
annotated in terms of interestingness by human assessors, to make it
possible to study the correlation between the two subtasks. A
dedicated web-based annotation tool was developed by the organising
team for the previous edition of the task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This year some
incremental improvements were added, and the tool was released as free
and open source software3. Overall, more than 252 annotators
participated in the annotation for the video data and 189 for the images.
The cultural distribution is over 22 diferent countries in the world.
As in last year’s setup we use a pair-wise comparison protocol [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
where annotators are provided with a pair of images/shots at a time
and asked to tag which one in the pair is the more interesting for
them. As a change from last year, we now ask the question in a way
more directly connected to the commercial application: “Which
image/video makes you more interested in watching the whole
movie?”, with the intent to make the decision criteria clearer to
the annotators. As an exhaustive annotation of all possible pairs is
practically impossible due to the required human resources, a
boosting selection was used instead. In particular, we used a modified
version of the adaptive square design method [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], in which several
annotators participated in each iteration. In this method the number
of comparisons for each iteration is reduced from all possible pairs
3
n(n − 1)/2 ∼ O(n2) to a subset of pairs n(√n − 1) ∼ O(n 2 ), where n
is the number of segments or images. For the development set, we
started from iteration 6, as we could reuse the annotations done last
year. To achieve the ranking used as the basis for the next round, the
pair-based annotations are aggregated with the Bradley-Terry-Luce
(BTL) model computation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] resulting in an interestingness degree
      </p>
      <sec id="sec-4-1">
        <title>2http://multimediaeval.org/mediaeval2016/persondiscovery/ 3https://github.com/mvsjober/pair-annotate</title>
        <p>
          C.H. Demarty et al.
for each image/shot. Previously the same procedure was also used
to get the final interestingness values. This year we used an
alternative method, which took into account all pair comparisons from
all rounds done this year into a single large BTL calculation. This
was done mainly because we discovered afterwards that some
annotations from earlier rounds had to be discarded, because of some
unserious annotators. These annotators occasionally switched to
cheating, where they simply always selected the first, or the
second item as the most interesting one without actually assessing
the media contents. In the development set as many as 10% of the
annotations were marked as invalid and not included in the final
BTL calculation. We added some heuristic anti-cheating measures
to the system, although it is not possible to perfectly detect all
cheating. Unfortunately, in the iterative approach, we could only
discard annotations from the most recent round, as it would be
based on the previous round’s BTL output, which is why we
developed another solution to compute the final BTL ranking. The final
binary decisions are obtained using a thresholding scheme that
tries to detect the boundary where interestingness values make the
“jump” between the underlying distributions of the non interesting
and interesting populations. See last year’s overview paper for a
more detailed description [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>RUN DESCRIPTION</title>
      <p>Every team can submit up to 10 runs, 5 per subtask. For each subtask,
a required run is defined: Image subtask - required run: classification
is to be carried out with the use of the visual information. External
data is allowed. Video subtask - required run: classification is to be
achieved with the use of both audio and visual information. External
data is allowed. Apart from these required runs, any additional run
for each subtask will be considered as a general run, i.e., anything
is allowed, both from the method point of view and the information
sources.
6</p>
    </sec>
    <sec id="sec-6">
      <title>EVALUATION</title>
      <p>For both subtasks, the oficial evaluation metric will be the mean
average precision at 10 (MAP@10) computed over all videos, and
over the top 10 best ranked images/video shots. MAP@10 is selected
because it reflects the VOD use case, where the goal is to select a
small set of the most interesting images or video segments for each
movie. To provide a broad overview of the systems’ performances,
other common metrics will also be provided. All metrics will be
computed by using the trec_eval tool from NIST4.
7</p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSIONS</title>
      <p>In the 2017 Predicting Media Interestingness task a complete and
comparative framework for the evaluation of content
interestingness is proposed. Details on the methods and results of each
individual participant team can be found in the working note papers of
the MediaEval 2017 workshop proceedings.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank Yu-Gang Jiang and Baohan Xu from the Fudan
University, China, Hervé Bredin, from LIMSI, France, and Michael Gygli for
providing the features that accompany the released data. Part of the task
was funded under research grant PN-III-P2-2.1-PED-2016-1065, agreement
30PED/2017, project SPOTTER.</p>
      <sec id="sec-8-1">
        <title>4http://trec.nist.gov/trec_eval/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Xesca</given-names>
            <surname>Amengual</surname>
          </string-name>
          , Anna Bosch, and Josep Lluís de la Rosa.
          <year>2015</year>
          .
          <article-title>Review of Methods to Predict Social Image Interestingness</article-title>
          and Memorability. Springer,
          <fpage>64</fpage>
          -
          <lpage>76</lpage>
          . https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -23192-
          <issue>1</issue>
          _
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Daniel</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Berlyne</surname>
          </string-name>
          .
          <year>1960</year>
          .
          <article-title>Conflict, arousal and curiosity</article-title>
          .
          <source>Mc-Graw-Hill.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Bradley</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Terry</surname>
          </string-name>
          .
          <year>1952</year>
          .
          <article-title>Rank Analysis of Incomplete Block Designs: the method of paired comparisons</article-title>
          .
          <source>Biometrika</source>
          <volume>39</volume>
          (
          <issue>3-4</issue>
          ) (
          <year>1952</year>
          ),
          <fpage>324</fpage>
          -
          <lpage>345</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Dalal</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Triggs</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Histograms of oriented gradients for human detection</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision</source>
          and Pattern Recognition.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Danelljan</surname>
          </string-name>
          , Gustav Hager, Fahad Shahbaz Khan, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Felsberg</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Accurate scale estimation for robust visual tracking</article-title>
          .
          <source>In British Machine Vision Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Mats Sjöberg, Bogdan Ionescu,
          <string-name>
            <surname>Thanh-Toan</surname>
            <given-names>Do</given-names>
          </string-name>
          , Hanli Wang,
          <string-name>
            <surname>Ngoc Q.K. Duong</surname>
            , and
            <given-names>Frédéric</given-names>
          </string-name>
          <string-name>
            <surname>Lefebvre</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>MediaEval 2016 Predicting Media Interestingness Task</article-title>
          .
          <source>In Proceedings of the MediaEval 2016 Workshop</source>
          . Hilversum, Netherlands.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Sagnik</given-names>
            <surname>Dhar</surname>
          </string-name>
          , Vicente Ordonez, and Tamara L Berg.
          <year>2011</year>
          .
          <article-title>High level describable attributes for predicting aesthetics and interestingness</article-title>
          .
          <source>In IEEE International Conference on Computer Vision</source>
          and Pattern Recognition.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gygli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Grabner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Riemenschneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nater</surname>
          </string-name>
          , and
          <string-name>
            <surname>L. van Gool.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>The Interestingness of Images</article-title>
          .
          <source>In ICCV International Conference on Computer Vision.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Yu-Gang</surname>
            <given-names>Jiang</given-names>
          </string-name>
          , Qi Dai, Tao Mei, Yong Rui, and
          <string-name>
            <surname>Shih-Fu Chang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Super Fast Event Recognition in Internet Videos</article-title>
          .
          <source>IEEE Transactions on Multimedia 177</source>
          ,
          <issue>8</issue>
          (
          <year>2015</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Y-G. Jiang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Zheng</surname>
            , and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Yan</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Understanding and Predicting Interestingness of Videos</article-title>
          .
          <source>In AAAI Conference on Artificial Intelligence .</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lazebnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ponce</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>2169</fpage>
          -
          <lpage>2178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Jing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Marcus</given-names>
            <surname>Barkowsky</surname>
          </string-name>
          , and Patrick Le Callet.
          <year>2013</year>
          .
          <article-title>Boosting paired comparison methodology in measuring visual discomfort of 3DTV: performances of three diferent designs</article-title>
          .
          <source>In SPIE Electronic Imaging, Stereoscopic Displays and Applications</source>
          , Vol.
          <volume>8648</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International Journal on Computer Vision</source>
          <volume>60</volume>
          (
          <year>2004</year>
          ),
          <fpage>91</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pietikainen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Maenpaa</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Multiresolution grayscale and rotation invariant texture classification with local binary patterns</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>24</volume>
          (
          <issue>7</issue>
          ) (
          <year>2002</year>
          ),
          <fpage>971</fpage>
          -
          <lpage>987</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Modeling the shape of the scene: a holistic representation of the spatial envelope</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>42</volume>
          (
          <year>2001</year>
          ),
          <fpage>145</fpage>
          -
          <lpage>175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Paul J.</given-names>
            <surname>Silvia</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Exploring the psychology of interest</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Craig</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>Phoebe</given-names>
            <surname>Ellsworth</surname>
          </string-name>
          .
          <year>1985</year>
          .
          <article-title>Patterns of cognitive appraisal in emotion</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>48</volume>
          ,
          <issue>4</issue>
          (
          <year>1985</year>
          ),
          <fpage>813</fpage>
          -
          <lpage>838</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Du</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and
          <string-name>
            <given-names>Manohar</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning spatiotemporal features with 3d convolutional networks</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ehinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>SUN database: Large-scale scene recognition from abbey to zoo</article-title>
          .
          <source>In IEEE CVPR Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>3485</fpage>
          -
          <lpage>3492</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>