<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multi-Modal Ensemble Models for Predicting Video Memorability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tony Zhao</string-name>
          <email>tonyzhao@berkeley.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irving Fang</string-name>
          <email>irvingf7@berkeley.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jefrey Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerald Friedland</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of California</institution>
          ,
          <addr-line>Berkeley</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>2283</volume>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Modeling media memorability has been a consistent challenge in the field of machine learning. The Predicting Media Memorability task in MediaEval2020 is the latest benchmark among similar challenges addressing this topic. Building upon techniques developed in previous iterations of the challenge, we developed ensemble methods with the use of extracted video, image, text, and audio features. Critically, in this work we introduce and demonstrate the eficacy and high generalizability of extracted audio embeddings as a feature for the task of predicting media memorability.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION AND RELATED WORK</title>
      <p>
        The MediaEval2020 Predicting Media Memorability Task aims to
predict how memorable videos are. The dataset consists of videos
with short-term and long-term memorability scores, pre-extracted
images and video features, and other information described in detail
in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Participants predict the probability that each video will be
remembered in both the short (after a few minutes) and long (after
24-72 hours) term.
      </p>
      <p>
        Image-based features extracted from pre-trained convolutional
neural networks (CNNs) have been demonstrated as useful in
predicting the memorability of videos in the dataset. These features in
combination with semantic-embedding and image captioning
models have also proved efective for predicting memorability scores
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Video-based features, such as C3D and I3D, have also recently
been considered in the study of video memorability [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Finally,
the best performing models of the 2019 iteration of the Predicting
Media Memorability challenge utilized ensemble models with the
above-mentioned features [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        The training set comprises of 590 short videos (1-8 seconds long)
with 2-5 human-annotated captions each. Another development set
of 410 additional videos was provided later, but was not used as we
discovered the quality of annotated memorability scores to be worse
than that of the original set. Each video has corresponding
shortterm and long-term memorability scores, which were adjusted using
methods outlined in previous work [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], namely the following
update functions to calculate decay rate  and memorability  at
duration 
 ←
Í=1 1() Í=(1) log( () ) [ () − () ]
      </p>
      <p>Í=1 1() Í=(1) [log( () )]2
1 Õ()   ()
() ←  () =1 [ ( ) −  log(  )]
where we have  () observations for video  given by  () ∈ {0, 1}
and  () where   = 1 implies that the repeated video was correctly
detected when shown after time   . The average  was 74.96, so we
adjusted the memorability score for each video to be 75 after 10
iterations with  converging to -0.0264.</p>
      <p>Models were then trained on a 80-20 training-validation split on
video id. Since concatenating multiple multi-modal features resulted
in extremely high-dimensional feature vectors which were dificult
to train with, we trained various models (Support Vector Regressor,
Bayesian Ridge Regressor, Linear Models) on each individual feature
independently.</p>
      <p>We also explored several prediction aggregation functions for
when a video has multiple embeddings of a specific feature, such
as several diferent model predictions due to multiple captions for
a given video. While previous work generally defaulted to a simple
average, we discovered that taking the median value performed
consistently better than either mean, max, or min. In the opposite
scenario, where a video has no feature such as a soundless video
for audio-based models, we defaulted to using the average of all
predicted memorability scores.</p>
      <p>We noticed high variance in the performance of the models,
which we attributed to the relatively small dataset. Therefore we
ran the feature models over 5 random seeds and took the best
performing features of each modality (VGGish for audio, ResNet152
for image, C3D for video, GloVe for text). The predictions of these
models were then used to perform grid-search over permutations of
buckets of 5% to calculate a weighted average as our final ensemble
model.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Audio Features</title>
      <p>
        Using VGGish [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a pre-trained CNN model (trained on AudioSet),
we extracted 128-dimensional embeddings for each second of video
audio. Each video had on average had 5.6 embeddings extracted,
with one soundless video having none. Bayesian Ridge Regressor
provided the best performance on these features.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Image/Video Features</title>
      <p>
        Local Binary Patterns (LBP), VGG [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and Convolutional 3D (C3D)
proved to have notable performance among the provided features.
We discovered that Support Vector Regressor (SVR) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] models
produced the best results with these features.
      </p>
      <p>
        We also extracted 8 equally spaced frames from each video, which
were used for our own image-based feature extraction. The
penultimate layer of a pre-trained ResNet-152 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (trained on ImageNet)
was used to extract a 2048-dimensional feature vector. Similarly,
Modality
video
image
image
image
audio
text
      </p>
      <p>Feature</p>
      <p>C3D</p>
      <sec id="sec-4-1">
        <title>ResNet152</title>
        <p>VGG
LBP</p>
      </sec>
      <sec id="sec-4-2">
        <title>VGGish</title>
      </sec>
      <sec id="sec-4-3">
        <title>GloVe</title>
        <p>Model
SVR
SVR
SVR</p>
        <p>SVR
Bayesian Ridge</p>
        <p>GRU
we found that the best performance came from SVR models trained
on the extracted features.</p>
        <p>
          Following the methods outlined in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], pre-trained facial-emotion
models were used to extract emotion-based features on the video
frames. Ultimately, these features were not used in any of the final
models, as only approximately a third of the videos had detectable
faces as well as overall poor performance from all emotion-based
models.
2.3
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Text Features</title>
      <p>
        Provided with several human-annotated captions for each video,
we explored several diferent methods to extract semantic features.
Past work suggested that simpler methods like bag-of-words could
outperform more sophisticated methods in terms of both
efectiveness and eficiency [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. We vectorized the captions using
bag-ofwords before training with ordinary least squares, ridge, and lasso
regression models. Bag-of-words vectorization and linear models
performed the worst among our text-based approaches and were
not used in the final model.
      </p>
      <p>
        We tokenized the captions before extracting 300-dimensional
GloVe [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] word embeddings (trained on Wikipedia 2014 and
Gigaword 5) to vectorize the caption tokens. These vectors were used as
input to train a recurrent neural network with gated recurrent units
(GRU) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our final text-based model had 64 units in the initial
GRU layer with a dropout of 0.8, followed by 4 dense layers with a
dropout of 0.25 and ReLU activation, trained for 150 epochs with
early stopping, a learning rate of 0.001, batch size of 64, and an
Adam optimizer.
      </p>
      <p>
        We also explored machine-generated captions [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to augment
our text-based approaches. After generating captions for each video
based on the first, middle, and last frames and mixing the generated
captions with the human-annotated captions, we discovered that
the performance improvement was insignificant. These automatic
captions were not included in our final models.
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>The notable features are seen in Table 1, with the bolded features
being selected for our final ensemble models, seen in Table 2 and
Table 3. Given the small dataset, we observe relatively high
variance in performance. VGGish had the best performance, smallest
variance, and thus presumably the best generalization for
shortterm memorability. However, VGGish features tended to perform
poorly for predicting long-term memorability despite maintaining
low variance.</p>
      <p>The diference in validation and test Spearman’s rank correlation
in our final ensemble models are likely due in part to overfitting after
performing grid-search over a relatively small dataset. In addition,
we noticed that the quality of annotations between datasets were
diferent, which may cause distribution diferences in memorability
scores and consequently poorer performance at test time.
4</p>
    </sec>
    <sec id="sec-7">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>Our main contribution is the demonstration that audio-based
models perform well for predicting short-term memorability and
generalize much more readily than other methods with the dataset.
This may be in part due to the low dimensionality of the extracted
audio embeddings. We reiterate that ensembling models of diferent
modalities achieve the best performance as each model represents
a diferent high-level abstraction of the data.</p>
      <p>The size of the dataset was notably smaller than similar datasets
for the task of predicting media memorability. Future work would
ideally iterate on our findings for much larger datasets.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>David</given-names>
            <surname>Azcona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Enric</given-names>
            <surname>Moreu</surname>
          </string-name>
          , Feiyan Hu,
          <string-name>
            <given-names>Tomás E.</given-names>
            <surname>Ward</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Predicting Media Memorability Using Ensemble Models</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2019 Workshop</source>
          , Sophia Antipolis, France,
          <fpage>27</fpage>
          -30
          <source>October 2019 (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          .
          <article-title>CEUR-WS.org</article-title>
          . http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          / MediaEval_19_paper_15.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chih-Chung Chang</surname>
          </string-name>
          and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          <volume>2</volume>
          (
          <year>2011</year>
          ),
          <volume>27</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>27</fpage>
          . Issue 3. Software available at http://www.csie.ntu.edu.tw/~cjlin/libsvm.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Kyunghyun</given-names>
            <surname>Cho</surname>
          </string-name>
          , Bart van Merriënboer,
          <string-name>
            <surname>Caglar Gulcehre</surname>
            , Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and
            <given-names>Yoshua</given-names>
          </string-name>
          <string-name>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Doha, Qatar,
          <fpage>1724</fpage>
          -
          <lpage>1734</lpage>
          . https://doi.org/10.3115/v1/
          <fpage>D14</fpage>
          -1179
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Helene</surname>
            <given-names>Demarty</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngoc Q. K. Duong</surname>
            , and
            <given-names>Martin</given-names>
          </string-name>
          <string-name>
            <surname>Engilberge</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video Memorability</article-title>
          .
          <source>In Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          (ICCV).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Alba</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , Rukiye Savran Kiziltepe, Jon Chamberlain, Mihai Gabriel Constantin,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Faiyaz Doctor, Bogdan Ionescu,
          <string-name>
            <given-names>and Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of MediaEval 2020 Predicting Media Memorability task: What Makes a Video Memorable?</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2020 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          .
          <source>In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <volume>770</volume>
          -
          <fpage>778</fpage>
          . https://doi.org/10.1109/CVPR.
          <year>2016</year>
          .90
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Shawn</given-names>
            <surname>Hershey</surname>
          </string-name>
          , Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, Channing Moore, Manoj Plakal, Devin Platt, Rif A.
          <string-name>
            <surname>Saurous</surname>
            , Bryan Seybold, Malcolm Slaney,
            <given-names>Ron</given-names>
          </string-name>
          <string-name>
            <surname>Weiss</surname>
          </string-name>
          , and Kevin Wilson.
          <year>2017</year>
          .
          <article-title>CNN Architectures for Large-Scale Audio Classification</article-title>
          . In International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          . https://arxiv.org/abs/1609.09430
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Khosla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Raju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding and Predicting Image Memorability at a Large Scale</article-title>
          .
          <source>In 2015 IEEE International Conference on Computer Vision</source>
          (ICCV).
          <volume>2390</volume>
          -
          <fpage>2398</fpage>
          . https: //doi.org/10.1109/ICCV.
          <year>2015</year>
          .275
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Deng</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Very deep convolutional neural network based image classification using small training sample size</article-title>
          .
          <source>In 2015 3rd IAPR Asian Conference on Pattern Recognition (ACPR)</source>
          .
          <volume>730</volume>
          -
          <fpage>734</fpage>
          . https://doi.org/10.1109/ACPR.
          <year>2015</year>
          .7486599
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Luo</surname>
          </string-name>
          , G. Shakhnarovich,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cohen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Price</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Discriminability Objective for Training Descriptive Captions</article-title>
          .
          <source>In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>6964</fpage>
          -
          <lpage>6974</lpage>
          . https://doi.org/10.1109/CVPR.
          <year>2018</year>
          .00728
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Jefrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>GloVe: Global Vectors for Word Representation</article-title>
          .
          <source>In Empirical Methods in Natural Language Processing (EMNLP)</source>
          .
          <volume>1532</volume>
          -
          <fpage>1543</fpage>
          . http: //www.aclweb.org/anthology/D14-1162
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Alison</surname>
            <given-names>Reboud</given-names>
          </string-name>
          , Ismail Harrando, Jorma Laaksonen, Danny Francis, Raphaël Troncy, and Hector Laria Mantecon.
          <year>2019</year>
          .
          <article-title>Combining textual and visual modeling for predicting media memorability</article-title>
          .
          <source>In MediaEval</source>
          <year>2019</year>
          , 10th MediaEval Benchmarking Initiative for Multimedia Evaluation Workshop,
          <fpage>27</fpage>
          -29
          <source>October</source>
          <year>2019</year>
          ,
          <string-name>
            <given-names>Sophia</given-names>
            <surname>Antipolis</surname>
          </string-name>
          , France. Sophia Antipolis, FRANCE. http://www.eurecom.fr/publication/6062
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Ricardo</given-names>
            <surname>Manhães</surname>
          </string-name>
          <string-name>
            <surname>Savii</surname>
          </string-name>
          , Samuel Felipe dos Santos, and
          <string-name>
            <given-names>Jurandy</given-names>
            <surname>Almeida</surname>
          </string-name>
          .
          <year>2018</year>
          . GIBIS at MediaEval 2018:
          <article-title>Predicting Media Memorability Task</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2018</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>