<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigating Memorability of Dynamic Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>APPROACH Motivation</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Phuc H. Le-Khac, Ayush K. Rai</institution>
          ,
          <addr-line>Graham Healy, Alan F. Smeaton</addr-line>
          ,
          <institution>Noel E. O'Connor Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The Predicting Media Memorability task in MediaEval'20 has some challenging aspects compared to previous years. In this paper we identify the high-dynamic content in videos and dataset of limited size as the core challenges for the task, we propose directions to overcome some of these challenges and we present our initial result in these directions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Since the seminal paper on image memorability by Isola et. al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
there has been growing interest in building computational models
to understand and predict the intrinsic memorability of media, as
well as other subjective perceptions of media [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The Predicting
Media Memorability task at MediaEval is designed to facilitate and
promote research in this area by requiring task participants to
develop computational approaches to generate measures of media
memorability, which in turn also helps to better understand the
subjective memorability of human cognition. Having a prediction
model of a media’s memorability also allows for interesting
applications to be developed, such as techniques to enhance memorability
of images through the use of neural style transfer [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Continuing from the success of the previous two years [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ],
the MediaEval’20 Predicting Media Memorability challenge [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
remains the same, where teams need to build prediction models
for the memorability scores from a dataset of videos with captions.
There are two types of memorability scores: short-term scores for
videos that are shown for a second time within a short timespan of
the initial viewing, and long-term scores for videos that are shown
again two to three days later. The videos making up this year’s
training dataset contain 590 videos with sound from the TRECVid
2019 Video-to-text dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], compared to 8,000 soundless videos
from VideoMem [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] dataset.
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        In previous challenges, a variety of methods that utilised
diferent data modalities were explored [
        <xref ref-type="bibr" rid="ref11 ref17 ref2">2, 11, 17</xref>
        ]. As previous works
have shown [
        <xref ref-type="bibr" rid="ref13 ref6">6, 13</xref>
        ], utilising high-level semantic features, either
extracted using deep networks or provided by human annotators
via text captions, are among the most efective methods to predict
memorability. However, the modalities and features that are most
predictive of memorability remain unclear i.e. the best approach
on last year’s challenge [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] used an ensemble of models trained on
a variety of modalities.
      </p>
      <p>
        High dynamic video. The short videos used in previous years’
memorability challenges were extracted from raw professional
footage that is used for example in creating high-quality film and
commercials [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Consequently, the majority of videos therefore
contain only one scene and are mostly static. Conversely, the data
for this year’s 2020 challenge is extracted from the TRECVid 2019
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Video-to-Text dataset. These videos were collected from social
media posts and TV shows. They contain multiple scene changes
in one video, and thus are more dynamic and with more complex
movement and activity levels, as illustrated in Fig.1.
      </p>
      <p>While previous approaches to memorability computation use
higher-semantic features, such as those extracted from deep neural
networks, text captions provided by human-annotators and
humancentric features such as emotions or aesthetics, most of these
features were extracted using frame-based models. In contrast, this
year’s dataset provided an unexplored challenge for methods that
can capture semantic features from highly dynamic videos.</p>
      <p>Small size dataset. Another challenge for this year’s data set is
the limited number of annotated videos with memorability scores.
Compared to previous years, this is an order of magnitude smaller
in size, with 590 videos in the training set and 500 videos in the test
set. A development set with an additional 500 videos was released
in the later stage of the benchmark.</p>
      <p>Challenging benchmark. The changes to this year’s dataset made
the task considerably more challenging, as the videos used were
more dynamic thus making the extraction of high-level semantic
features more dificult, and further compounded by the limited size
of the dataset.</p>
      <p>Motivated by these insights, we focus our work on using
videolevel features to capture the dynamic in videos, and attempt to
pre-train a memorability model on a larger dataset before
finetuning it on this year MediaEval’s data.
3.2</p>
    </sec>
    <sec id="sec-3">
      <title>Spatiotemporal baseline</title>
      <p>While the provided captions for each video can be used as high-level
semantic-rich features for predicting media memorability. However,
in our approach we focus on a method to extract high-level features
from raw video input, without any additional annotation.</p>
      <p>
        Motivated from the fact that this year’s video is highly dynamic
as discussed above, we chose to use features extracted from C3D
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], the only video-based method instead of frame-based model. We
used the C3D pre-extracted features provided by the task organisers
as the input for our memorability regression model.
      </p>
      <p>
        C3D model learned spatio-temporal representation. C3D [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], short
for 3D Convolution, is among the earliest approaches to learning
generic representation from videos. Extending the 2D Convolution
operations common in most image-processing deep learning
models, C3D’s homogeneous architecture composed of small 3 × 3 × 3
convolution kernels, expanding in all height, width and time
dimension of videos. Trained on a generic action recognition dataset, each
video passed through C3D returns a 4096-dimension feature vector
that extracts high-level semantic features from a video segment.
      </p>
      <p>
        Memorability Regression Model. With the extracted features via
C3D, we used a multi-layer perceptron (MLP) with two hidden
layers of 512 units followed by a Rectified Linear Unit (ReLU)
nonlinearity layer. To mitigate overfitting, we used Dropout [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] with
a probability of 0.1. Batch Normalisation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] was used to normalise
the features in each batch after each layer to help optimisation. The
ifnal model is a sequential application 3 building blocks composed
from BatchNorm, Dropout, fully-connected and ReLU activation
layers. For the final output layer, the ReLU activation is replaced
with a Sigmoid to obtain the final memorability score in the range
from 0 to 1.
      </p>
      <p>There was a total of more than 2 million parameters in the model,
most belonging to the first fully-connected layer due to the large
C3D’s feature size of 4096. To minimise the efect of outliers in
ranking score, we use L1 loss instead of the standard L2
MeanSquared Error (MSE) and saw a slight improvement in convergence
speed. The model is trained for 100 epochs on the training set
and the final model was picked based on the performance on the
validation set.
3.3</p>
    </sec>
    <sec id="sec-4">
      <title>Large scale Memorability Pre-training</title>
      <p>
        To overcome the limited size of this year’s dataset, we used the
recently released Memento10K [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] for pre-training our model
P.H. Le-Khac et al.
before fine tuning on the MediaEval challenge’s dataset.
Containing 10,000 videos, the Memento10K dataset is roughly the same
size as the VideoMem [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] dataset. Moreover, videos in the
Memento10K contains more action and are more similar to the new
data of this year’s challenge. Therefore we decided to focus our
large-scale pre-training approach on the Memento10K dataset and
replicate the accompanying SemNet model [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] instead of
pretraining on VideoMem [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] from previous years. The SemNet model
contains three separate sub-networks to process three diferent
input streams: image, optical flow and video stream.
      </p>
      <p>After pre-training SemNet on Memento10k data, we only retain
the video stream sub-network to fine tune on the MediaEval dataset
and discarded all other components.
4</p>
    </sec>
    <sec id="sec-5">
      <title>RESULTS AND DISCUSSION</title>
      <p>
        The results of our runs together with this year’s mean and variance
are reported in Table 1. Due to technical issues and time constraints,
we only submitted the baseline regression model based on C3D
features and did not have the result for fine tuning SemNet on the
Memento10K [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] dataset for this year’s challenge.
Even with the simplest method possible, the result from our
regression baseline with pre-extracted C3D features is not very far from
the mean. This support our hypothesis that spatio-temporal
representation is important for this year’s dataset and the performance
can be increased with other complementary high-level features.
Future work can explore this hypothesis further by fine tuning
the entire C3D model instead of just using the pre-extracted
features, or by using diferent models that also capture spatio-temporal
features.
      </p>
      <p>
        On the other hand, compared to the state-of-the-art result of
0.528 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] in last year’s challenge, the mean of the Spearman rank
correlation for all runs this year (0.058) is an order of magnitude
lower. This highlights the importance of large scale pre-training
and transfer learning techniques, and methods that can extract
high-level features from dynamic videos efectively.
      </p>
      <p>We believe the direction of our solution, given the challenges
in this year’s challenge, not only shows promise but also indicates
the importance of spatio-temporal models to capture high-level
semantics of videos for memorability prediction.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was co-funded by Science Foundation Ireland through the
SFI Centre for Research Training in Machine Learning (18/CRT/6183)
and the Insight Centre for Data Analytics (SFI/12/RC/2289_P2),
cofunded by the European Regional Development Fund. A.K. Rai also
acknowledges support from FotoNation Ltd.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>George</given-names>
            <surname>Awad</surname>
          </string-name>
          , Asad A.
          <string-name>
            <surname>Butt</surname>
            , Keith Curtis,
            <given-names>Yooyoung</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            , Jonathan Fiscus, Afzal Godil, Andrew Delgado, Jesse Zhang, Eliot Godard, Lukas Diduch, Alan F. Smeaton, Yvette Graham, Wessel Kraaij, and
            <given-names>Georges</given-names>
          </string-name>
          <string-name>
            <surname>Quenot</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>TRECVID 2019: An Evaluation Campaign to Benchmark Video Activity Detection, Video Captioning and Matching, and Video Search &amp; Retrieval. In TREC Video Retrieval Evaluation Notebook Papers and Slides</article-title>
          . https://www-nlpir.nist.gov/projects/tvpubs/tv.pubs.19.org. html
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>David</given-names>
            <surname>Azcona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Enric</given-names>
            <surname>Moreu</surname>
          </string-name>
          , Feiyan Hu, Tomás E Ward, and Alan F Smeaton.
          <year>2019</year>
          .
          <article-title>Predicting Media Memorability Using Ensemble Models</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          ,
          <source>France. CEUR Workshop Proceedings</source>
          , 3. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Helene</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Ngoc Duong, and Martin Engilberge.
          <article-title>VideoMem: Constructing, Analyzing, Predicting ShortTerm and Long-Term Video Memorability</article-title>
          .
          <source>In 2019 IEEE/CVF International Conference on Computer Vision</source>
          (ICCV) (
          <year>2019</year>
          ). IEEE,
          <fpage>2531</fpage>
          -
          <lpage>2540</lpage>
          . https://doi.org/10.1109/ICCV.
          <year>2019</year>
          .00262
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngoc Q K Duong</surname>
            ,
            <given-names>Mats Sjöberg</given-names>
          </string-name>
          , Bogdan Ionescu, and
          <string-name>
            <surname>Thanh-Toan Do</surname>
          </string-name>
          .
          <source>MediaEval</source>
          <year>2018</year>
          :
          <article-title>Predicting Media Memorability</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings (
          <year>2018</year>
          ). 3. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Gabriel</surname>
          </string-name>
          <string-name>
            <given-names>Constantin</given-names>
            , Bogdan Ionescu,
            <surname>Claire-Hélène</surname>
          </string-name>
          <string-name>
            <given-names>Demarty</given-names>
            ,
            <surname>Ngoc Q K Duong</surname>
            , Xavier
          </string-name>
          Alameda-Pineda, and
          <string-name>
            <given-names>Mats</given-names>
            <surname>Sjöberg</surname>
          </string-name>
          .
          <source>The Predicting Media Memorability Task at MediaEval</source>
          <year>2019</year>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France (
          <year>2019</year>
          ).
          <source>CEUR Workshop Proceedings</source>
          , 3. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Gabriel</surname>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
          </string-name>
          , Chen Kang, G. Dinu, Frédéric Dufaux, Giuseppe Valenzise, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Ionescu</surname>
          </string-name>
          .
          <article-title>Using Aesthetics and Action Recognition-Based Networks for the Prediction of Media Memorability</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France (
          <year>2019</year>
          ). CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Gabriel</surname>
          </string-name>
          <string-name>
            <surname>Constantin</surname>
          </string-name>
          , Miriam Redi, Gloria Zen, and
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Ionescu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Computational Understanding of Visual Interestingness Beyond Semantics: Literature Survey and Analysis of Covariates</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>52</volume>
          ,
          <issue>2</issue>
          ,
          <string-name>
            <surname>Article 25</surname>
          </string-name>
          (
          <year>2019</year>
          ),
          <volume>37</volume>
          pages. https://doi.org/10.1145/3301299
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Alba</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , Rukiye Savran Kiziltepe, Jon Chamberlain, Mihai Gabriel Constantin,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Faiyaz Doctor, Bogdan Ionescu,
          <string-name>
            <given-names>and Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of MediaEval 2020 Predicting Media Memorability task: What Makes a Video Memorable?</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2020 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Iofe</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift</article-title>
          .
          <source>In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 (ICML'15)</source>
          . JMLR.org,
          <volume>448</volume>
          -
          <fpage>456</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Devi Parikh, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Understanding the Intrinsic Memorability of Images</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , J.
          <string-name>
            <surname>Shawe-Taylor</surname>
            , R. Zemel,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bartlett</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
            , and
            <given-names>K. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          (Eds.), Vol.
          <volume>24</volume>
          . Curran Associates, Inc.,
          <fpage>2429</fpage>
          -
          <lpage>2437</lpage>
          . https://proceedings.neurips.cc/paper/ 2011/file/286674e3082feb7e5afb92777e48821f-Paper.pdf
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Leyva</surname>
          </string-name>
          and
          <string-name>
            <given-names>Faiyaz</given-names>
            <surname>Doctor</surname>
          </string-name>
          .
          <article-title>Multimodal Deep Features Fusion For Video Memorability Prediction</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France (
          <year>2019</year>
          ).
          <source>CEUR Workshop Proceedings</source>
          , 3. http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Anelise</surname>
            <given-names>Newman</given-names>
          </string-name>
          , Camilo Fosco, Vincent Casser,
          <string-name>
            <given-names>Allen</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barry McNamara</surname>
            ,
            <given-names>and Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multimodal Memorability: Modeling Efects of Semantics and Decay on Video Memorability</article-title>
          . In Computer Vision - ECCV
          <year>2020</year>
          ,
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Vedaldi</surname>
          </string-name>
          , Horst Bischof, Thomas Brox, and
          <string-name>
            <surname>Jan-Michael Frahm</surname>
          </string-name>
          (Eds.). Springer International Publishing, Cham,
          <fpage>223</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Alison</surname>
            <given-names>Reboud</given-names>
          </string-name>
          , Ismail Harrando,
          <string-name>
            <surname>Jorma</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Laaksonen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Francis</surname>
            , Raphaël Troncy, and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mantecón</surname>
          </string-name>
          .
          <article-title>Combining Textual and Visual Modeling for Predicting Media Memorability</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France (
          <year>2019</year>
          ). CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Aliaksandr</surname>
            <given-names>Siarohin</given-names>
          </string-name>
          , Gloria Zen, Cveta Majtanovic,
          <string-name>
            <surname>Xavier</surname>
            <given-names>AlamedaPineda</given-names>
          </string-name>
          , Elisa Ricci, and
          <string-name>
            <given-names>Nicu</given-names>
            <surname>Sebe</surname>
          </string-name>
          .
          <year>2019</year>
          -
          <volume>06</volume>
          -
          <fpage>14</fpage>
          .
          <article-title>Increasing Image Memorability with Neural Style Transfer</article-title>
          .
          <source>ACM Transactions on Multimedia Computing, Communications, and Applications 15</source>
          ,
          <issue>2</issue>
          (
          <fpage>2019</fpage>
          -06-14),
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          . https://doi.org/10.1145/3311781
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Nitish</surname>
            <given-names>Srivastava</given-names>
          </string-name>
          , Geofrey Hinton, Alex Krizhevsky, Ilya Sutskever, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Dropout: A Simple Way to Prevent Neural Networks from Overfitting</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>15</volume>
          ,
          <issue>56</issue>
          (
          <year>2014</year>
          ),
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          . http://jmlr.org/papers/v15/ srivastava14a.html
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bourdev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Torresani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <article-title>Learning Spatiotemporal Features with 3D Convolutional Networks</article-title>
          .
          <source>In 2015 IEEE International Conference on Computer Vision</source>
          (ICCV)
          <article-title>(</article-title>
          <year>2015</year>
          -
          <fpage>12</fpage>
          ).
          <fpage>4489</fpage>
          -
          <lpage>4497</lpage>
          . https://doi.org/10.1109/ICCV.
          <year>2015</year>
          .510
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Le-Vu</surname>
            <given-names>Tran</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinh-Loc Huynh</surname>
          </string-name>
          , and
          <string-name>
            <surname>Minh-Triet Tran</surname>
          </string-name>
          .
          <article-title>Predicting Media Memorability Using Deep Features with Attention and Recurrent Network</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France (
          <year>2019</year>
          ).
          <source>CEUR Workshop Proceedings</source>
          , 3. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>