<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of MediaEval 2020 Predicting Media Memorability Task: What Makes a Video Memorable?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alba G. Seco de Herrera</string-name>
          <email>alba.garcia@essex.ac.uk</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rukiye Savran Kiziltepe</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jon Chamberlain</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Gabriel Constantin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claire-Hélène Demarty</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Faiyaz Doctor</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Ionescu</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan F. Smeaton</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>InterDigital</institution>
          ,
          <addr-line>R&amp;I</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University Politehnica of Bucharest</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Essex</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>This paper describes the MediaEval 2020 Predicting Media Memorability task. After first being proposed at MediaEval 2018, the Predicting Media Memorability task is in its 3rd edition this year, as the prediction of short-term and long-term video memorability (VM) remains a challenging task. In 2020, the format remained the same as in previous editions. This year the videos are a subset of the TRECVid 2019 Video-to-Text dataset, containing more action rich video content as compared with the 2019 task. In this paper a description of some aspects of this task is provided, including its main characteristics, a description of the collection, the ground truth dataset, evaluation metrics and the requirements for participants' run submissions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Media platforms such as social networks, media advertisements,
information retrieval and recommendation systems deal with
exponential growth. Enhancing the relevance of multimedia occurrences
in our everyday lives requires new ways to organise – in particular,
to retrieve – digital content. Like other video metrics of importance,
such as aesthetics or interestingness, memorability can be regarded
as a useful aspect to help make a choice between competing videos.
This is even truer when one considers specific use cases of creating
commercials or educational content. Because the impact of
diferent multimedia content, images or videos, on human memory is
unequal, the capability of predicting the memorability of a given
piece of video content is of high importance for professionals in
the field of advertising and other fields. Beyond advertising, other
applications, such as film-making, education, content retrieval, etc.,
may also be influenced by this task.</p>
      <p>
        The Predicting Media Memorability task addresses this problem.
The task is part of the MediaEval benchmark and, following the
success of previous editions [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ], creates a common benchmarking
protocol and provides a ground truth dataset for short-term and
long-term memorability using common definitions.
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        The computational understanding of video memorability follows
on from the study of image memorability prediction, which has
attracted increasing attention since the seminal work of Isola et
al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Models have achieved very good results at predicting image
memorability [
        <xref ref-type="bibr" rid="ref15 ref8">8, 15</xref>
        ] and we have recently started to see the use of
techniques like style transfer to improve image memorability [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
thus illustrating that we have now moved from just measuring
memorability, to using memorability as an evaluation metric.
      </p>
      <p>
        In contrast, research on visual memorability (VM) from a
computer science point of view is in its early stage. Recently we have
seen other work on video memorability [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] with a particular focus
on short term, but the scarcity of studies on VM can be explained
by several reasons. Firstly, there is no publicly available data set
to train and test models, though the VideoMem [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and the
Memento10k [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] datasets are recent additions. The second point,
closely related to the previous one, is the lack of a common
definition for VM. Regarding modelling, previous attempts at predicting
VM [
        <xref ref-type="bibr" rid="ref12 ref3">3, 12</xref>
        ] have highlighted several features which contribute to
the prediction of VM, such as semantic, saliency and colour
features, but the work is far from complete and our capacity to propose
efective computational models will help to meet the challenge of
VM prediction.
      </p>
      <p>The goal of this task is to participate in the harmonisation and
the advancement of this emerging multimedia field. Furthermore, in
contrast to previous work on image memorability prediction, where
memorability was measured a few minutes after memorisation, we
propose a dataset with longer term memorability annotations. We
expect the predictions of the models trained on this data to be more
representative of long-term memory, which is used preferably in
numerous applications.
3</p>
    </sec>
    <sec id="sec-3">
      <title>TASK DESCRIPTION</title>
      <p>
        The Predicting Media Memorability task requires participants to
automatically predict memorability scores for short form videos,
that reflect the probability for a video to be remembered.
Participants were provided with a dataset of videos with short-term and
long-term memorability annotations, related information, and
preextracted state-of-the-art visual features. Therefore, two subtasks
were proposed to participants:
● Short-term VM prediction - scores were measured a few
minutes after the memorisation process;
● Long-term VM prediction - scores were measured 24-72
hours after the memorisation process.
The dataset is composed of a subset of short videos selected from
the TRECVid 2019 Video-to-Text dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] (see Figure 1). These
videos are shared under Creative Commons licenses that allow
their redistribution. The TRECVid videos have much more action
happening in them compared with those in the 2019 VM task, and
thus they correspond to more generic use cases.
      </p>
      <p>
        Each video consists of a coherent unit in terms of meaning and is
associated with two scores of memorability that refer to its
probability to be remembered after two diferent time durations of memory
retention. A set of pre-extracted features are also distributed:
● image-level features: AlexNetFC7 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], HOG [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], HSVHist,
      </p>
      <p>
        RGBHist, LBP [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], VGGFC7 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ];
● video-level feature: C3D [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>The image-level features were extracted from 3 frames for each
video: the first, the middle and the last frame. In addition, each
TRECVid video is accompanied by two textual captions
describing the activity. Additional information on the annotation was
also provided to allow further investigation of the user interaction
for memorability. Hence, the annotations collected from
participants were provided including the first appearance position and
the second appearance position of each target video along with the
response time of the user and the key pressed when watching each
video.</p>
      <p>
        The TRECVid 2019 Video-to-Text dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] contains 6,000
videos. In 2020, three subsets were distributed as part of the
MediaEval Predicting Media Memorability task. The training set
contained 590 videos, the development set 410 videos and the test set
500 videos. Each video was annotated by at least 16 annotators for
their short term memorability. However, there are fewer long term
annotations.
      </p>
      <p>
        Similar to previous editions of the task [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ], memorability has
been measured using recognition tests, i.e., through an objective
measure, a few minutes after the memorisation of the videos (short
term), and then 24 to 72 hours later (long term). The ground truth
dataset was collected by using a video memorability game protocol
proposed by Cohendet et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Two versions of the memorability
game were published. One was published on Amazon
Mechanical Turk (AMT) and another one was issued for general use with
following three language options: English, Spanish and Turkish.
      </p>
      <p>
        In the game of video memorability, participants are expected to
watch 180 and 120 videos in short-term and long-term memorisation
steps, respectively. The task is basically to press the space bar once
the participants recognise a previously seen video, which enables
to determine videos recognised and not recognised by them. In
the first step of the game, 40 target videos are repeated after a
few minutes to collect short-term memorability labels. As for filler
videos in the first step, 60 non-vigilance filler videos are displayed
once. 20 vigilance filler videos are repeated after a few seconds
to check participants’ attention to the task. After 24 hours to 72
hours, the same participants are expected to attend the second
step for collecting long-term memorability labels. This time, 40
target videos chosen randomly from among non-vigilance fillers
of the first step and 80 fillers selected randomly from new videos
are displayed to measure long-term memorability scores for those
target videos. Both short-term and long-term memorability scores
are calculated as the percentage of correct recognition for each
video by the participants. Relevant screenshots and label collection
procedures are demonstrated on the MediaEval task web page [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>SUBMISSION AND EVALUATION</title>
      <p>As in previous editions of the task, each team is required to predict
both short and long term memorability. In total, 10 runs can be
submitted, 5 for each. For the two required runs, all information
can be used in the development of the system, meaning provided
features, ground truth data, video sample titles, features extracted
from the visual content and even external data. The only exception,
in this case, is that the required short-term and long-term
memorability runs must not use each other’s score annotations. For the rest
of the runs, a maximum of 4 per subtask, everything is permitted,
including using cross-annotations between the subtasks.</p>
      <p>The outputs of the prediction models – i.e., the predicted
memorability scores for the videos – will be compared with ground
truth memorability scores using classic evaluation metrics (e.g.,
Spearman’s rank correlation).
6</p>
    </sec>
    <sec id="sec-5">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>In this paper we presented the third edition of the Predicting
Media Memorability at the MediaEval 2020 Benchmarking initiative.
The task provides a framework that allows a comparative study of
diferent state of the art Machine Learning approaches aiming to
predict short and long-term memorability. A collection of videos
is provided as well as memorability annotations and a common
evaluation metric. In addition, related information has been
provided to help participants in developing their approaches. Details
regarding the methods employed by participants and their results
can be found in the proceedings of the 2020 MediaEval workshop1.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was part-funded by NIST Award No. 60NANB19D155, by
Science Foundation Ireland under grant number SFI/12/RC/2289_P2
and under project AI4Media, A European Excellence Centre for
Media, Society and Democracy, H2020 ICT-48-2020, grant 951911.
1See CEUR Workshop Proceedings (CEUR-WS.org).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>George</given-names>
            <surname>Awad</surname>
          </string-name>
          , Asad A Butt, Keith Curtis,
          <string-name>
            <given-names>Yooyoung</given-names>
            <surname>Lee</surname>
          </string-name>
          , Jonathan Fiscus, Afzal Godil, Andrew Delgado, Jesse Zhang, Eliot Godard, Lukas Diduch, and others.
          <source>2019. TRECVID</source>
          <year>2019</year>
          :
          <article-title>An Evaluation Campaign to Benchmark Video Activity Detection, Video Captioning and Matching, and</article-title>
          <string-name>
            <given-names>Video</given-names>
            <surname>Search</surname>
          </string-name>
          &amp; Retrieval. (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Ngoc Duong, Mats Sjöberg, Bogdan Ionescu, and
          <string-name>
            <surname>Thanh-Toan Do</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>MediaEval 2018: Predicting media memorability task</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2018 Workshop</source>
          . Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          ,
          <source>Ngoc QK Duong, and Martin Engilberge</source>
          .
          <year>2019</year>
          .
          <article-title>VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          . 2531-
          <fpage>2540</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Gabriel</surname>
          </string-name>
          <string-name>
            <given-names>Constantin</given-names>
            , Bogdan Ionescu,
            <surname>Claire-Hélène</surname>
          </string-name>
          <string-name>
            <surname>Demarty</surname>
          </string-name>
          , Ngoc QK Duong,
          <article-title>Xavier Alameda-Pineda, and</article-title>
          <string-name>
            <given-names>Mats</given-names>
            <surname>Sjöberg</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Predicting Media Memorability Task at MediaEval 2019</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2019 Workshop</source>
          . Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Navneet</given-names>
            <surname>Dalal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bill</given-names>
            <surname>Triggs</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Histograms of oriented gradients for human detection</article-title>
          .
          <source>In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05)</source>
          , Vol.
          <volume>1</volume>
          . IEEE,
          <fpage>886</fpage>
          -
          <lpage>893</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Dong-Chen He</surname>
            and
            <given-names>Li</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Texture unit, texture spectrum, and texture analysis</article-title>
          .
          <source>IEEE Transactions on Geoscience and Remote Sensing</source>
          <volume>28</volume>
          ,
          <issue>4</issue>
          (
          <year>1990</year>
          ),
          <fpage>509</fpage>
          -
          <lpage>512</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Phillip</given-names>
            <surname>Isola</surname>
          </string-name>
          , Jianxiong Xiao, Devi Parikh, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>What makes a photograph memorable?</article-title>
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>36</volume>
          ,
          <issue>7</issue>
          (
          <year>2013</year>
          ),
          <fpage>1469</fpage>
          -
          <lpage>1482</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Khosla</surname>
          </string-name>
          , Akhil S Raju, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding and predicting image memorability at a large scale</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          . 2390-
          <fpage>2398</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>1097</volume>
          -
          <fpage>1105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>MediaEval</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>MediaEval 2020: Predicting Media Memorability</article-title>
          . (
          <year>2020</year>
          ). https://multimediaeval.github.io/editions/2020/tasks/ memorability/ Accessed:
          <fpage>2020</fpage>
          -11-26.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Anelise</surname>
            <given-names>Newman</given-names>
          </string-name>
          , Camilo Fosco, Vincent Casser,
          <string-name>
            <given-names>Allen</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barry McNamara</surname>
            ,
            <given-names>and Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multimodal Memorability: Modeling Efects of Semantics and Decay on Video Memorability</article-title>
          . In Computer Vision - ECCV
          <year>2020</year>
          ,
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Vedaldi</surname>
          </string-name>
          , Horst Bischof, Thomas Brox, and
          <string-name>
            <surname>Jan-Michael Frahm</surname>
          </string-name>
          (Eds.). Springer International Publishing, Cham,
          <fpage>223</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sumit</surname>
            <given-names>Shekhar</given-names>
          </string-name>
          , Dhruv Singal, Harvineet Singh,
          <string-name>
            <given-names>Manav</given-names>
            <surname>Kedia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Akhil</given-names>
            <surname>Shetty</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Show and recall: Learning what makes videos memorable</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision Workshops</source>
          .
          <fpage>2730</fpage>
          -
          <lpage>2739</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Aliaksandr</surname>
            <given-names>Siarohin</given-names>
          </string-name>
          , Gloria Zen, Cveta Majtanovic,
          <string-name>
            <surname>Xavier</surname>
            <given-names>AlamedaPineda</given-names>
          </string-name>
          , Elisa Ricci, and
          <string-name>
            <given-names>Nicu</given-names>
            <surname>Sebe</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Increasing Image Memorability with Neural Style Transfer</article-title>
          .
          <source>ACM Trans. Multimedia Comput. Commun. Appl</source>
          .
          <volume>15</volume>
          ,
          <issue>2</issue>
          ,
          <string-name>
            <surname>Article 42</surname>
          </string-name>
          (
          <year>June 2019</year>
          ),
          <volume>22</volume>
          pages. https: //doi.org/10.1145/3311781
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Very Deep Convolutional Networks for Large-Scale Image Recognition</article-title>
          .
          <source>In International Conference on Learning Representations.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Hammad</given-names>
            <surname>Squalli-Houssaini</surname>
          </string-name>
          ,
          <source>Ngoc QK Duong</source>
          , Marquant Gwenaëlle, and
          <string-name>
            <surname>Claire-Hélène Demarty</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep learning for predicting image memorability</article-title>
          .
          <source>In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . IEEE,
          <fpage>2371</fpage>
          -
          <lpage>2375</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Du</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and
          <string-name>
            <given-names>Manohar</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning spatiotemporal features with 3d convolutional networks</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          . 4489-
          <fpage>4497</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>