<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting Media Memorability with Audio, Video, and Text representations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alison Reboud</string-name>
          <email>alison.reboud@eurecom.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ismail Harrando</string-name>
          <email>ismail.harrando@eurecom.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorma Laaksonen</string-name>
          <email>jorma.laaksonen@aalto.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raphaël Troncy</string-name>
          <email>raphael.troncy@eurecom.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aalto University</institution>
          ,
          <addr-line>Espoo</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>EURECOM</institution>
          ,
          <addr-line>Sophia Antipolis</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper describes a multimodal approach proposed by the MeMAD team for the MediaEval 2020 “Predicting Media Memorability” task. Our best approach is a weighted average method combining predictions made separately from visual, audio, textual and visiolinguistic representations of videos. Our best model achieves Spearman scores of 0.101 and 0.078, respectively, for the short and long term predictions tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Considering video memorability as a useful tool for digital content
retrieval as well as for sorting and recommending an ever growing
number of videos, the Predicting Media Memorability task aims
at fostering the research in the field by asking its participants to
automatically predict both a short and a long term memorability
score for a given set of annotated videos. The full description for
this task is provided in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Last year’s best approaches for both
the long term [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and short term tasks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] rely on multimodal
features. Our method is inspired from last year’s best approaches
but also acknowledges the specifics of the 2020’s edition dataset.
More specifically, because in comparison to last year’s set of videos,
the TRECVid videos contain more actions, our model uses video
features and image features for multiple frames. In addition, because
this year sound was included in the videos, our model includes
audio features. Finally, a key contribution of our approach is to test the
relevance of visiolinguistic representation for the Media
Memorability task. Our final model 1 is a multimodal weighted average with
visual and audio deep features extracted from the videos, textual
features from the provided captions and visiolinguistic features.
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>We trained separate models for the short and long term predictions
using originally a 6-fold cross-validation of the training set, which
means that we typically had 492 samples for training and 98 samples
for testing each model.
1https://github.com/MeMAD-project/media-memorability
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Audio-Visual Approach</title>
      <p>
        Our audio-visual memorability prediction scores are based on
using a feed-forward neural network with a concatenation of video
and audio features in the input, one hidden layer of units and
one unit in the output layer. The best performance was obtained
with 2575-dimensional features consisting of the concatenation of
2048-dimensional I3D [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] video features and 527-dimensional audio
features. Our audio features encode the occurrence probabilities
of the 527 classes of the Google AudioSet Ontology [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] in each
video clip. The hidden layer uses ReLU activations and dropout
during the training phase, while the output unit is sigmoidal. The
training of the network used the Adam optimizer. The features, the
number of training epochs and the number of units in the hidden
layer were selected with the 6-fold cross-validation. For short term
memorability prediction, the optimal number of epochs was 750
and the optimal hidden layer size 80 units, whereas for the long
term prediction these figures were 260 and 160, respectively.
      </p>
      <p>
        We also experimented with other types of features and their
combinations. These include the ResNet [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] features extracted just
from the middle frames of the clips as this approach worked very
well last year. The contents of this year’s videos are, however, such
that genuine video features I3D and C3D [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] work better than still
image features. When I3D and AudioSet features are used, C3D
features do not bring any additional advantage.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Textual Approach</title>
      <p>Our textual approach leverages the video descriptions provided by
the organizers. First, all the provided descriptions are concatenated
by video identifier to get one string per video. To generate the
textual representation of the video content, we used the following
methods:
•
•
•
•</p>
      <p>Computing TF-IDF, removing rare (less than 4 occurrences)
and stopwords and accounting for frequent 2-grams.</p>
      <p>
        Averaging GloVe embeddings for all non-stopwords words
using the pre-trained 300d version [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Averaging BERT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] token representations (keeping all the
words in the descriptions up to 250 words per sentence).
Using Sentence-BERT [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] sentence representations. We
use the distilled version that is fine-tuned for the STS
Textual Similarity Benchmark2.
      </p>
      <p>For each representation, we experimented with multiple
regression models and finetuned the hyper-parameters for each model
2https://huggingface.co/sentence-transformers/distilbert-base-nli-stsb-mean-tokens
using the 6-fold cross-validation on the training set. For our
submission, we used the Averaging GloVe embeddings with a Support
Machine Regressor with an RBF kernel and a regulation parameter
 = 1 − 5.</p>
      <p>We also attempted enhancing the provided descriptions with
additional captions automatically generated using the DeepCaption3
software. We did not see an improvement in the results, which
is probably due to the nature of the clips provided for this year’s
edition (as DeepCaption is trained on static stock images from MS
COCO and TGIF datasets).
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Visiolinguistic Approach</title>
      <p>
        ViLBERT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is a task-agnostic extension of BERT that aims to learn
the associations and links between visual and linguistic properties
of a concept. It has a two-stream architecture, first modelling each
modality (i.e. visual and textual) separately, and then fusing them
through a set of attention-based interactions (co-attention).
ViLBERT is pre-trained using the Conceptual Captions data set (3.3M
image-caption pairs) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] on masked multi modal learning and
multi-modal alignment prediction. We used a frozen pre-trained
model which was fine-tuned twice, first on the task of
VideoQuestion Answering (VQA) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and then on the 2019 MediaEval
Memorability task and dataset.
      </p>
      <p>The 1024-dimensional features extracted for the two modalities
can be combined in diferent ways.In our experiment, multiplying
textual and visual feature vectors performed the best for short term
memorability prediction but using the sole visual feature vectors
worked better for long term memorability prediction. Averaging
the features extracted from 6 frames performed better than only
using only the middle frame. We experimented with the same set
of regression models as for the textual approach. In our submission,
we used a Support Machine Regressor with a regulation parameter
 = 1 − 5 and an RBF or Poly kernel respectively for short and
long term scores prediction.
3</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>We have prepared 5 diferent runs following the task description
defined as follows:
• run1 = Audio-Visual Score
• run2 = Visiolinguistic Score
• run3 = Textual Score
• run4 = 0.5 * run1 + 0.2 * run2 + 0.3 * run3
• run5 = run4 with LT scores for LT task
For the Long Term task, all models except run5 use exclusively
shortterm scores. For runs 4 and 5, we normalise the scores obtained
from runs 1, 2 and 3 before combining them.</p>
      <p>Table 1 provides the Spearman score obtained for each run when
performing a 6-folds cross-validation on the training set. We
observe that our models use only the training set, as the annotations
on the later-provided development set did not yield better results.
We hypothesize that this is due to the fewer number of annotations
per video available as many videos had a score for 1, for instance,
which we do not observe on the training set.
3https://github.com/aalto-cbir/DeepCaption</p>
      <p>We present in Table 2 the final results obtained on the test set
using models trained on the full training set composed of 590 videos.
We observe that the weighted average method which uses short
term scores works the best for both short and long term prediction,
obtaining results which are approximately double the mean
Spearman score obtained across the teams. Our best results (Spearman
scores) on the test set are however significantly worse than the
ones we obtained on average over the 6-folds of the training set
suggesting that the test set is quite diferent from the training set.
The results for Long Term prediction are always worse than the
ones for Short Term prediction. Finally, both our scores and the
mean score across team are below the ones obtained for the 2018
and 2019 videos.
4</p>
    </sec>
    <sec id="sec-7">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>This paper describes a multimodal weighted average method
proposed for the 2020 Predicting Media Memorability task of
MediaEval. One of the key contribution of this paper is to have shown that
based on our experiments during the model construction or testing
phase, in comparison to image, audio and text, video features
performed the best. Similarly to last year, short term scores predictions
correlated better with long term scores than the predictions made
when training directly on long term scores. Finally considering the
diference of results obtained between the training and test set, it
would be interesting to investigate further the diferences between
these datasets in terms of content (video, audio and text) and
annotation. We conclude that generalizing this type of task to diferent
video genres and characteristics remain a scientific challenge.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This work has been partially supported by the European Union’s
Horizon 2020 research and innovation programme via the project
MeMAD (GA 780069).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Stanislaw</given-names>
            <surname>Antol</surname>
          </string-name>
          , Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and
          <string-name>
            <given-names>Devi</given-names>
            <surname>Parikh</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>VQA: Visual Question Answering</article-title>
          .
          <source>In IEEE International Conference on Computer Vision</source>
          (ICCV). IEEE, Santiago, Chile.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>David</given-names>
            <surname>Azcona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Enric</given-names>
            <surname>Moreu</surname>
          </string-name>
          , Feiyan Hu, Tomás E Ward, and Alan F Smeaton.
          <year>2019</year>
          .
          <article-title>Predicting media memorability using ensemble models</article-title>
          .
          <source>In MediaEval 2019: Multimedia Benchmark Workshop</source>
          . Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>João</given-names>
            <surname>Carreira</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset</article-title>
          .
          <source>In IEEE Conference on Computer Vision</source>
          and
          <article-title>Pattern Recognition (CVPR)</article-title>
          . IEEE,
          <fpage>4724</fpage>
          -
          <lpage>4733</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
          <article-title>In Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)</article-title>
          . ACL, Minneapolis, Minnesota, USA,
          <fpage>4171</fpage>
          --
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Alba</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , Rukiye Savran Kiziltepe, Jon Chamberlain, Mihai Gabriel Constantin,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Faiyaz Doctor, Bogdan Ionescu,
          <string-name>
            <given-names>and Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of MediaEval 2020 Predicting Media Memorability task: What Makes a Video Memorable?</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2020 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Jort</surname>
            <given-names>F Gemmeke</given-names>
          </string-name>
          , Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence,
          <string-name>
            <given-names>R Channing</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Manoj</given-names>
            <surname>Plakal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Marvin</given-names>
            <surname>Ritter</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Audio set: An ontology and human-labeled dataset for audio events</article-title>
          .
          <source>In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . New Orleans, Louisiana, USA,
          <fpage>776</fpage>
          -
          <lpage>780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In IEEE Conference on Computer Vision</source>
          and
          <article-title>Pattern Recognition (CVPR)</article-title>
          . IEEE,
          <string-name>
            <surname>Las</surname>
            <given-names>Vegas</given-names>
          </string-name>
          , Nevada, USA,
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jiasen</given-names>
            <surname>Lu</surname>
          </string-name>
          , Dhruv Batra, Devi Parikh, and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Visionand-Language Tasks</article-title>
          . In 33
          <source>Conference on Neural Information Processing Systems (NeurIPS)</source>
          . Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Pennington</surname>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In International Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . ACL, Melbourne, Australia,
          <fpage>1532</fpage>
          --
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Alison</surname>
            <given-names>Reboud</given-names>
          </string-name>
          , Ismail Harrando, Jorma Laaksonen, Danny Francis, Raphaël Troncy, and Héctor Laria Mantecón.
          <year>2019</year>
          .
          <article-title>Combining Textual and Visual Modeling for Predicting Media Memorability</article-title>
          . In MediaEval 2019: Multimedia Benchmark Workshop. Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks</article-title>
          .
          <source>In International Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . ACL,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China,
          <fpage>3982</fpage>
          --
          <lpage>3992</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Piyush</surname>
            <given-names>Sharma</given-names>
          </string-name>
          , Nan Ding,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Goodman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Radu</given-names>
            <surname>Soricut</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</article-title>
          . ACL, Melbourne, Australia,
          <fpage>2556</fpage>
          -
          <lpage>2565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Du</surname>
            <given-names>Tran</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Lubomir D.</given-names>
            <surname>Bourdev</surname>
          </string-name>
          , Rob Fergus, Lorenzo Torresani, and
          <string-name>
            <given-names>Manohar</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning Spatiotemporal Features with 3D Convolutional Networks</article-title>
          .
          <source>In International Conference on Computer Vision</source>
          (ICCV). IEEE, Santiago, Chile,
          <fpage>4489</fpage>
          -
          <lpage>4497</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>