<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Show and Recall @ MediaEval 2018 ViMemNet: Predicting Video Memorability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ritwick Chaudhry</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manoj Kilaru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sumit Shekhar Adobe Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rchaudhr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>kilaru</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>sushekha}@adobe.com</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>In the current age of expanding access to the Internet, there has been a flood of videos on the web. Studying the human cognitive factors that afect the consumption of these videos is becoming increasingly important, to be able to efectively organize and curate them. One such important cognitive factor is Video Memorability, which is the ability to recall a video's content after watching it. In this paper, we present our approach to solving the MediaEval 2018 Predicting Media Memorability Task. We develop a 3-forked pipeline for predicting Memorability Scores, which leverages the visual image features (both low-level and high-level), the image saliency in diferent video frames, and the information present in the captions. We also explore the relevance of other features such as image memorability scores of the diferent frames in the video, and present a detailed analysis of the results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        With the explosion of visual content on the Internet, it is becoming
increasingly important to discover new cognitive metrics to analyze
the content. Memorability of visual content is one such metric.
Previous studies on memorability [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] suggest that even though
we come across a plethora of photos and images each day, our
long-term memory is capable of storing massive number of objects
with details, from images that we have come across. Although
memorability of visual content is afected by personal context and
subjective consumption [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], it has been shown [
        <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
        ] that there is
a high degree of consistency amongst people in the ability to retain
information. This makes memorability an objective target.
      </p>
      <p>
        Recent eforts in trying to predict the memorability of images
have been successful, with the development of a large scale dataset
on image memorability [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], near human consistency rank
correlation for image memorability is achieved, thereby establishing
that human cognitive abilities are within reach for the field of
computer vision. Despite these eforts in the realm of images, there
has been limited work in predicting memorability of videos, given
the added complexities that videos bring in.
      </p>
      <p>Therefore, we seek to analyze the task of predicting memorability
scores for videos in the context of MediaEval 2018.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        The concept of memorability has been studied in psychology and
neuroscience studies. They mostly focused on visual memory,
studying for instance the human capacity of remembering object
details [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], efect of stimuli on encoding and later retrieval from
memory [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], memory systems of the brain [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] etc. Broadly, prior
work on recall of information about viewed visual content can be
divided into the following categories:
      </p>
      <p>
        Image Memorability: Isola et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] started out the
computational study revolving around the cognitive metric, memorability
of images. The authors showed that across various subjects and
under wide range of contexts, memorability of an image is
consistent, which indicates that image memorability is an intrinsic
property of images. Since then many prior works have explored
this problem [
        <xref ref-type="bibr" rid="ref10 ref11 ref14 ref15 ref17">10, 11, 14, 15, 17</xref>
        ]. Khosla et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] introduced largest
annotated image memorability dataset (containing 60,000 images
from diverse sources) and showed that fine-tuned deep features
outperform all other features by a large margin. Fajtl et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] used
a visual attention mechanism and designed an end-to-end
trainable deep neural network for estimating memorability. Siarohin et
al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] adopted a deep architecture for generating a memorable
picture from a given input image and a style seed.
      </p>
      <p>
        Video Memorability: Han et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] commenced computational
studies on memorability of videos by learning from brain functional
magnetic resonance imaging. As the method used fMRI
measurements of the users for learning the model, it would be dificult to
generalize. The authors in [
        <xref ref-type="bibr" rid="ref20 ref5">5, 20</xref>
        ] used spatio-temporal features
to represent video dynamics and used a regression framework for
predicting memorability.
      </p>
      <p>We extend their work by proposing a trainable deep learning
framework for predicting Video Memorability scores.
3</p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>In this section, we discuss the task of predicting Video Memorability.
The feature extraction from videos is described in Section 3.1 and
an analysis of features for memorability prediction is discussed in
Section 3.2
Ritwick Chaudhry, Manoj Kilaru, Sumit Shekhar
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Feature Extraction</title>
      <p>
        • C3D: C3D features are outputs of the final classification
layer of deep 3D convolutional networks trained on a large
scale supervised video dataset;
• Color Histogram: This is computed in the HSV space
using 64 bins in each color space for 3 key-frames (first,
middle and last frames) for each video;
• InceptionV3: This corresponds to the final class
activations of the InceptionV3 deep network for object detection,
trained on the ImageNet dataset;
• Saliency: The aspect of visual content which grabs
human attention has shown to be useful in predicting
memorability [
        <xref ref-type="bibr" rid="ref11 ref6">6, 11</xref>
        ]. We used the highest ranking saliency
prediction model in the MIT Saliency Benchmark on the
MIT300 dataset (AUC, sAUC), DeepGaze II [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and
generated saliency maps for all 3 key frames;
• Captions: We used textual captions present in the dataset
which were generated manually for describing the videos.
Captions can be a compact source of representing the
video content, and thus can be useful for predictions.
• Image Memorability: We divided the video into 10 frames
and used image memorability scores for each frame
predicted using a pre-trained model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
• HMP: Histogram of motion patterns is computed for each
video and Principal Component Analysis (PCA) (with 128
principal components) is applied on them to obtain a
reduced dimensional encoding;
• HoG: HoG descriptors (Histograms of Oriented Gradients)
are calculated on 32x32 windows of each key frame and
256 principal components are extracted from each feature.
      </p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Model Description and Prediction Analysis</title>
      <p>
        Here, we describe our proposed model and provide the
training details. The dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] consists of 8000 training videos
and 2000 test videos, with each video being 7 seconds long.
The train data is randomly split into 80:20 split for
training the model and validation respectively. We describe our
3-forked pipeline architecture (see Figure 1) for predicting
Memorability Scores, which leverages the aforementioned
visual features, image saliency and captions.
      </p>
      <p>Saliency: The saliency maps extracted from the video
frames are down scaled to 120 by 68. A 2 layer CNN is applied
on these maps with each layer consisting of a 2D
convolution, batch normalization, relu activation and a max pool
operation. Finally they are vectorized and a fully connected
linear is applied on it.</p>
      <p>Captions: Each word in the captions is represented using
pre-trained 100 dimensional Glove embeddings and each
embedding is passed through single layered LSTM of hidden
dimension 100. The final representation of this caption is
appended with rest of features as shown in the Figure 1.</p>
      <p>Other Visual Features: C3D, Color Histogram, HMP,
HoG, InceptionV3, Image Memorability features are
concatenated and then combined with saliency and caption
representations. Then a five layered dense fully connected linear
neural is applied to obtain a single number representing the
memorability score. The model is trained using Stochastic
Gradient Descent with Mean Squared Error Loss function.</p>
      <p>We trained diferent models for both Long Term and Short
Term memorability scores using the aforementioned
architecture. Results are presented in Table 1 and Table 2</p>
      <p>We believe that all higher level G1 features are required
for memorability prediction. To test whether low level
features (Group G3) will help in prediction, we ran experiments,
including and excluding the G3 features (Table 1). We also
experimented with using the Image Memorability scores of
sampled frames from the video.</p>
    </sec>
    <sec id="sec-6">
      <title>4 DISCUSSION AND OUTLOOK</title>
      <p>In this work, we have described a robust way to model and
compute Video Memorability. It is empirically clear that
using G3 features (low level image features of keyframes) help.
Also, including Image memorability scores of key frames
didn’t lead to any improvement in performance, hinting to
the fact that videos are much more than just a set of frames,
and that temporal features matter. In future, we plan to
conduct the Video Memorability experiment with improved
features like Dense Optical Flow features, Action based features
representing the sequence of actions in the video, and also
aim to leverage the audio in the videos.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Wilma</surname>
            <given-names>A Bainbridge</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daniel D Dilks</surname>
            , and
            <given-names>Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Memorability: A stimulus-driven perceptual neural signature distinctive from memory</article-title>
          .
          <source>NeuroImage</source>
          <volume>149</volume>
          (
          <year>2017</year>
          ),
          <fpage>141</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Wilma</surname>
            <given-names>A Bainbridge</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Phillip</given-names>
            <surname>Isola</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The intrinsic memorability of face photographs</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          <volume>142</volume>
          ,
          <issue>4</issue>
          (
          <year>2013</year>
          ),
          <fpage>1323</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Timothy</surname>
            <given-names>F Brady</given-names>
          </string-name>
          ,
          <article-title>Talia Konkle, George A Alvarez,</article-title>
          and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Visual long-term memory has a massive storage capacity for object details</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>105</volume>
          ,
          <issue>38</issue>
          (
          <year>2008</year>
          ),
          <fpage>14325</fpage>
          -
          <lpage>14329</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngoc Q. K. Duong</surname>
          </string-name>
          , Mats Sjöberg, Bogdan Ionescu, and
          <string-name>
            <surname>Thanh-Toan Do</surname>
          </string-name>
          .
          <source>MediaEval</source>
          <year>2018</year>
          :
          <article-title>Predicting Media Memorability Task</article-title>
          .
          <source>In The Proceedings of MediaEval 2018 Workshop</source>
          ,
          <fpage>29</fpage>
          -31
          <source>October</source>
          <year>2018</year>
          ,
          <string-name>
            <given-names>Sophia</given-names>
            <surname>Antipolis</surname>
          </string-name>
          , France.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          , Karthik Yadati, Ngoc QK Duong, and
          <string-name>
            <surname>Claire-Hélène Demarty</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Annotating, Understanding, and Predicting Long-term Video Memorability</article-title>
          .
          <source>In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval. ACM</source>
          ,
          <volume>178</volume>
          -
          <fpage>186</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Rachit</given-names>
            <surname>Dubey</surname>
          </string-name>
          , Joshua Peterson, Aditya Khosla,
          <string-name>
            <surname>Ming-Hsuan Yang</surname>
            , and
            <given-names>Bernard</given-names>
          </string-name>
          <string-name>
            <surname>Ghanem</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>What makes an object memorable?</article-title>
          .
          <source>In Proceedings of the ieee international conference on computer vision</source>
          . 1089-
          <fpage>1097</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Jiri</given-names>
            <surname>Fajtl</surname>
          </string-name>
          , Vasileios Argyriou, Dorothy Monekosso, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Remagnino</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>AMNet: Memorability Estimation with Attention</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>6363</fpage>
          -
          <lpage>6372</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Junwei</given-names>
            <surname>Han</surname>
          </string-name>
          , Changyuan Chen, Ling Shao, Xintao Hu, Jungong Han, and Tianming Liu.
          <year>2015</year>
          .
          <article-title>Learning computational models of video memorability from fMRI brain imaging</article-title>
          .
          <source>IEEE transactions on Cybernetics 45</source>
          ,
          <issue>8</issue>
          (
          <year>2015</year>
          ),
          <fpage>1692</fpage>
          -
          <lpage>1703</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R Reed</given-names>
            <surname>Hunt and James B Worthen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Distinctiveness and memory</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Devi Parikh, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Understanding the intrinsic memorability of images</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>2429</volume>
          -
          <fpage>2437</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Jianxiong Xiao, Devi Parikh, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>What makes a photograph memorable?</article-title>
          <source>IEEE Transactions on Pattern Analysis &amp; Machine Intelligence</source>
          <volume>1</volume>
          (
          <year>2013</year>
          ),
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Jianxiong Xiao, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>What makes an image memorable? (</article-title>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Aditya</surname>
            <given-names>Khosla</given-names>
          </string-name>
          , Akhil S. Raju, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding and Predicting Image Memorability at a Large Scale</article-title>
          .
          <source>In International Conference on Computer Vision</source>
          (ICCV).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Aditya</surname>
            <given-names>Khosla</given-names>
          </string-name>
          , Jianxiong Xiao, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Memorability of image regions</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          .
          <volume>296</volume>
          -
          <fpage>304</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Jongpil</surname>
            <given-names>Kim</given-names>
          </string-name>
          , Sejong Yoon, and
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Pavlovic</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Relative spatial features for image memorability</article-title>
          .
          <source>In Proceedings of the 21st ACM international conference on Multimedia. ACM</source>
          ,
          <volume>761</volume>
          -
          <fpage>764</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Matthias</surname>
            <given-names>Kummerer</given-names>
          </string-name>
          , Thomas S. A.
          <string-name>
            <surname>Wallis</surname>
          </string-name>
          , Leon A.
          <string-name>
            <surname>Gatys</surname>
            , and
            <given-names>Matthias</given-names>
          </string-name>
          <string-name>
            <surname>Bethge</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Understanding Low-</article-title>
          and
          <string-name>
            <surname>High-Level Contributions</surname>
          </string-name>
          to Fixation Prediction.
          <source>In The IEEE International Conference on Computer Vision</source>
          (ICCV).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Matei</given-names>
            <surname>Mancas and Olivier Le Meur</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Memorability of natural scenes: The role of attention</article-title>
          .
          <source>In Image Processing (ICIP)</source>
          ,
          <year>2013</year>
          20th IEEE International Conference on. IEEE,
          <fpage>196</fpage>
          -
          <lpage>200</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>James</surname>
            <given-names>L McGaugh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larry Cahill</surname>
            , and
            <given-names>Benno</given-names>
          </string-name>
          <string-name>
            <surname>Roozendaal</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Involvement of the amygdala in memory storage: interaction with other brain systems</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>93</volume>
          ,
          <issue>24</issue>
          (
          <year>1996</year>
          ),
          <fpage>13508</fpage>
          -
          <lpage>13514</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>James</surname>
            <given-names>L McGaugh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ines B Introini-Collison</surname>
          </string-name>
          , Larry F Cahill,
          <string-name>
            <surname>Claudio Castellano</surname>
          </string-name>
          , Carla Dalmaz, Marise B Parent, and
          <string-name>
            <surname>Cedric L Williams</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Neuromodulatory systems and memory storage: role of the amygdala</article-title>
          .
          <source>Behavioural brain research 58</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>1993</year>
          ),
          <fpage>81</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Sumit</surname>
            <given-names>Shekhar</given-names>
          </string-name>
          , Dhruv Singal, Harvineet Singh,
          <string-name>
            <given-names>Manav</given-names>
            <surname>Kedia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Akhil</given-names>
            <surname>Shetty</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Show and Recall: Learning What Makes Videos Memorable</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>2730</fpage>
          -
          <lpage>2739</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Aliaksandr</surname>
            <given-names>Siarohin</given-names>
          </string-name>
          , Gloria Zen, Cveta Majtanovic,
          <string-name>
            <surname>Xavier</surname>
            <given-names>AlamedaPineda</given-names>
          </string-name>
          , Elisa Ricci, and
          <string-name>
            <given-names>Nicu</given-names>
            <surname>Sebe</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>How to Make an Image More Memorable?: A Deep Style Transfer Approach</article-title>
          .
          <source>In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval. ACM</source>
          ,
          <volume>322</volume>
          -
          <fpage>329</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>