<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Leveraging Audio Gestalt to Predict Media Memorability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorin Sweeney</string-name>
          <email>lorin.sweeney8@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Graham Healy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan F. Smeaton</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics, Dublin City University</institution>
          ,
          <addr-line>Glasnevin, Dublin 9</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Memorability determines what evanesces into emptiness, and what worms its way into the deepest furrows of our minds. It is the key to curating more meaningful media content as we wade through daily digital torrents. The Predicting Media Memorability task in MediaEval 2020 aims to address the question of media memorability by setting the task of automatically predicting video memorability. Our approach is a multimodal deep learning-based late fusion that combines visual, semantic, and auditory features. We used audio gestalt to estimate the influence of the audio modality on overall video memorability, and accordingly inform which combination of features would best predict a given video's memorability scores.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION AND RELATED WORK</title>
      <p>
        Our memories make us who we are—holding together the very
fabric of our being. The vast majority of the population predominantly
rely on visual information to remember and identify people, places,
and things [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. It is known that people generally tend to remember
and forget the same images [
        <xref ref-type="bibr" rid="ref14 ref3">3, 14</xref>
        ], which implies that there are
intrinsic qualities or characteristics that make visual content more
or less memorable. While there is some evidence to suggest that
sounds similarly have such intrinsic properties [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], it is
generally accepted that auditory memory is inferior to visual memory,
and decays more quickly [
        <xref ref-type="bibr" rid="ref18 ref4 ref6">4, 6, 18</xref>
        ]. However, it is important not to
resultantly resort to dismissing the role of the audio modality in
memorability—as multisensory experiences exhibit increased recall
accuracy compared to unisensory ones [
        <xref ref-type="bibr" rid="ref26 ref27">26, 27</xref>
        ]—but to avail of the
potential contextual priming information sounds provide [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In
our work, we seek to explore the influence of the audio modality
on overall video memorability, and employ [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]’s concept of audio
gestalt in order to do so.
      </p>
      <p>
        Image memorability is commonly defined as the probability of
an observer detecting a repeated image in a stream of images a
few minutes after exposition [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref3">3, 13–15</xref>
        ]. This paper outlines our
participation in the 2020 MediaEval Predicting Media Memorability
Task [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] where memorability is sub-categorised into short-term
(after minutes) and long-term (after 24-72 hours) memorability. The
task requires participants to create systems that can predict the
short-term and long-term memorability of a set of viral videos. The
dataset, annotation protocol, pre-computed features, and
groundtruth data are described in the overview paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Previous
attempts at computing memorability have consistently demonstrated
that of-the-shelf pre-computed features, such as C3D; HMP; LBP;
etc., have mostly been unfruitful [
        <xref ref-type="bibr" rid="ref11 ref24 ref5 ref7">5, 7, 11, 24</xref>
        ], with captions being
the exception [
        <xref ref-type="bibr" rid="ref25 ref29">25, 29</xref>
        ]. The high usefulness of captions is likely
due to the semantically rich nature of text—the medium with the
highest cued recall and free recall for narrative [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]—capturing the
most information about the elements in the videos with the smallest
semantic gap. Recent attempts highlighted the efectiveness of deep
features in conjunction with other semantically rich features, such
as emotions or actions [
        <xref ref-type="bibr" rid="ref2 ref28">2, 28</xref>
        ]. To our knowledge, the influence of
audio on overall video memorability has yet to be explored.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>OUR APPROACH</title>
      <p>The dataset is comprised of three sets—an initial development set,
a supplemental development set, and a test set. Both development
sets contain ground-truth (GT) scores and anonymised user
annotation data for 590 and 410 videos respectively, while the test
set contains 500 videos without GT scores. Upon inspection of the
supplemental development set, we noticed that the short-term GT
scores were true-negative rates (probabilities that unseen videos
would not be falsely remembered), rather than true-positive rates,
like the rest of the scores. We accordingly decided to generate our
own short-term GT scores by employing collaborative filtering with
the provided reaction time annotations. The resulting matrix of
predicted reaction times, for each user-video combination, was used to
calculate a short-term memorability score for each video. Predicted
reaction times that were more than two standard deviations from
the mean reaction time were counted as misses. We then divided
this new development set into a training (800 videos) and testing
set (200 videos). Our approach is predicated on the conditional
exclusion/inclusion of audio related features depending on audio
gestalt levels.</p>
      <p>
        Audio Gestalt: Defined in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] as high-level conceptual audio
features, imageability; human causal uncertainty (Hcu); arousal;
and familiarity, were found to be strongly correlated with audio
memorability. We predict audio gestalt by doing a weighted sum
of these four features. Imageability is based on whether the audio
is music or not using the PANNs [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] network. Hcu and arousal
scores are predicted with an xResNet34 pre-trained on ImageNet
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and fine-tuned on HCU400 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For familiarity we chose to use
the top audio-tag confidence score of the PANNs [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] network,
as we had observed a correlation between the two scores. These
scores were then weighted and summed to produce an audio gestalt
score for a video. Depending on a threshold, one of two pathways—
without audio, using captions and frames, and with audio, using
audio augmented captions, frames, and audio spectrograms—was
used to predict memorability scores.
      </p>
      <p>
        Without Audio: Deep Neural Network (DNN) frame-based and
caption-based models were chosen, as they had proven to be quite
effective in previous video memorability prediction attempts [
        <xref ref-type="bibr" rid="ref25 ref29">25, 29</xref>
        ].
For our caption model, given that overfitting was a primary concern,
we decided to use the AWD-LSTM (ASGD Weight-Dropped LSTM)
architecture [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], as it is highly regularised, and is used in
state-ofthe-art language modelling. In order to fully take advantage of the
high level representations that a language model ofers, we used
transfer learning. The specific transfer learning method employed
was UMLFiT [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], a method that uses discriminative fine-tuning,
slanted triangular learning rates, and gradual unfreezing to avoid
catastrophic forgetting. A language model was pre-trained on the
Wiki-103 dataset, and fine-tuned on the first 300,000 captions from
Google’s Conceptual Captions dataset [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. The encoder from that
ifne-tuned language model was then re-used in another model of
the same architecture, but trained on 800 of the development set
captions to predict short-term and long-term memorability scores
rather than the next word in a sentence. For our frame based model,
we transfer trained an xResNet50 model pre-trained on ImageNet
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. It was first fine-tuned on Memento10k [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], then fine-tuned on
800 of the development set videos.
      </p>
      <p>
        With Audio: We incorporated audio features by augmenting
captions with audio tags, and training a CNN on audio spectrograms.
We kept the same frame based model as a control. We opted to
augment captions with audio tags, rather than relying on audio
tags alone, so that the context provided by captions was not lost.
We used the PANNs model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] to extract audio tags, append them
to their corresponding development set captions, and re-trained the
aforementioned caption model. For the audio spectrogram model,
we extracted Mel-frequency cepstral coeficients from the audio,
and stacked them with their delta coeficients in order to create
three channel spectrogram images. We then used the same transfer
training procedure as our frame based model.
      </p>
      <p>For each stream, final predictions were the result of a weighted
sum of their constituent model predictions.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>Table 1 shows the scores for our runs when tested on 200 dev set
videos kept for validation. Runs 1-3 were intended as controls for
our 4th run in which we tested our proposed framework. Run 1
was the “audio only” control, using augmented captions and audio
L. Sweeney, G. Healy, A. Smeaton
spectrograms; run 2 was the “no audio” control, using captions
and video frames; and run 3 was the “everything” control, using
all features irrespective of audio gestalt scores. These preliminary
result suggest that the audio gestalt-based conditional inclusion of
audio features does indeed improve memorability prediction (run 3
vs run 4).</p>
      <p>
        Table 2 shows the scores achieved by each of our submitted
runs. Unfortunately, due to a mistake in the submission process,
we do not have oficial short-term results for our 3 rd run, and are
therefore unable to properly compare it with run 4 and validate
our preliminary results. The stark diference between the oficial
results and our preliminary results highlight the inherent dificulty
of achieving generalisability when learning from limited data.
Unfortunately, there is no way yet of knowing if these diferences are
simply due the lack of distributional overlap between training and
testing videos, or if it is a case of overfitting. Surprisingly, our
highest short-term score was achieved by a model not trained on any
data from the task—an xResNet50 model pre-trained on ImageNet
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and fine-tuned on Memento10k [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. We believe that this result
is most likely due to the better generalisability of models trained on
larger and more diverse datasets. Our highest long-term score was
achieved by our “audio only” run, which indicates that the audio
modality does indeed provide some contextual information during
video memorability recognition tasks.
      </p>
      <p>Further investigation into the influence of the audio modality on
overall video memorability is clearly required. Testing our proposed
framework on a much larger video memorability dataset would be
an interesting next step towards that goal.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This publication has emanated from research supported by Science
Foundation Ireland (SFI) under Grant Number SFI/12/RC/2289_P2,
co-funded by the European Regional Development Fund.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ishwarya</given-names>
            <surname>Ananthabhotla</surname>
          </string-name>
          , David B Ramsay, and Joseph A Paradiso.
          <year>2019</year>
          .
          <article-title>HCU400: An Annotated Dataset for Exploring Aural Phenomenology Through Causal Uncertainty</article-title>
          .
          <source>In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . IEEE,
          <fpage>920</fpage>
          -
          <lpage>924</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>David</given-names>
            <surname>Azcona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Enric</given-names>
            <surname>Moreu</surname>
          </string-name>
          , Feiyan Hu,
          <string-name>
            <given-names>Tomás</given-names>
            <surname>Ward</surname>
          </string-name>
          , and Alan F Smeaton.
          <year>2019</year>
          .
          <article-title>Predicting media memorability using ensemble models</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2019</year>
          ,
          <string-name>
            <surname>Sophia</surname>
            <given-names>Antipolis</given-names>
          </string-name>
          , France. CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2670</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Wilma</surname>
            <given-names>A Bainbridge</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Phillip</given-names>
            <surname>Isola</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The intrinsic memorability of face photographs</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          <volume>142</volume>
          ,
          <issue>4</issue>
          (
          <year>2013</year>
          ),
          <fpage>1323</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>James</given-names>
            <surname>Bigelow</surname>
          </string-name>
          and
          <string-name>
            <given-names>Amy</given-names>
            <surname>Poremba</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Achilles' ear? Inferior human short-term and recognition memory in the auditory modality</article-title>
          .
          <source>PloS One 9</source>
          ,
          <issue>2</issue>
          (
          <year>2014</year>
          ),
          <year>e89914</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Ritwick</given-names>
            <surname>Chaudhry</surname>
          </string-name>
          , Manoj Kilaru, and
          <string-name>
            <given-names>Sumit</given-names>
            <surname>Shekhar</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Show and Recall@ MediaEval 2018 ViMemNet: Predicting Video Memorability</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Michael</surname>
            <given-names>A Cohen</given-names>
          </string-name>
          ,
          <article-title>Todd S Horowitz,</article-title>
          and
          <string-name>
            <surname>Jeremy M Wolfe</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Auditory recognition memory is inferior to visual recognition memory</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>106</volume>
          ,
          <issue>14</issue>
          (
          <year>2009</year>
          ),
          <fpage>6008</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Romain</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>Claire-Hélène Demarty</surname>
          </string-name>
          , and Ngoc QK Duong.
          <year>2018</year>
          .
          <article-title>Transfer Learning for Video Memorability Prediction.</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>ImageNet: A large-scale hierarchical image database</article-title>
          .
          <source>In 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE</source>
          ,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Furnham</surname>
          </string-name>
          , Barrie Gunter, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Green</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Remembering science: The recall of factual information as a function of the presentation mode</article-title>
          .
          <source>Applied Cognitive Psychology 4</source>
          ,
          <issue>3</issue>
          (
          <year>1990</year>
          ),
          <fpage>203</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Alba</given-names>
            <surname>García Seco de Herrera</surname>
          </string-name>
          , Rukiye Savran Kiziltepe, Jon Chamberlain, Mihai Gabriel Constantin,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Faiyaz Doctor, Bogdan Ionescu,
          <string-name>
            <given-names>and Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Overview of MediaEval 2020 Predicting Media Memorability task: What Makes a Video Memorable?</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2020 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Rohit</given-names>
            <surname>Gupta</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kush</given-names>
            <surname>Motwani</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Linear Models for Video Memorability Prediction Using Visual and Semantic Features</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Jeremy</given-names>
            <surname>Howard</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Ruder</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Universal Language Model Fine-tuning for Text Classification</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          .
          <fpage>328</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Jianxiong Xiao, Devi Parikh, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>What makes a photograph memorable</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>36</volume>
          ,
          <issue>7</issue>
          (
          <year>2013</year>
          ),
          <fpage>1469</fpage>
          -
          <lpage>1482</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Phillip</surname>
            <given-names>Isola</given-names>
          </string-name>
          , Jianxiong Xiao, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>What makes an image memorable</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE</source>
          ,
          <fpage>145</fpage>
          -
          <lpage>152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Aditya</surname>
            <given-names>Khosla</given-names>
          </string-name>
          , Akhil S Raju, Antonio Torralba, and
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Understanding and predicting image memorability at a large scale</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          . 2390-
          <fpage>2398</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Edwin</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Kirkpatrick</surname>
          </string-name>
          .
          <year>1894</year>
          .
          <article-title>An experimental study of memory</article-title>
          .
          <source>Psychological Review</source>
          <volume>1</volume>
          ,
          <issue>6</issue>
          (
          <year>1894</year>
          ),
          <fpage>602</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Qiuqiang</surname>
            <given-names>Kong</given-names>
          </string-name>
          , Yin Cao, Turab Iqbal, Yuxuan Wang,
          <string-name>
            <given-names>Wenwu</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mark D</given-names>
            <surname>Plumbley</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>PANNs: Large-scale pretrained audio neural networks for audio pattern recognition</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          <volume>28</volume>
          (
          <year>2020</year>
          ),
          <fpage>2880</fpage>
          -
          <lpage>2894</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Larsson</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lars</given-names>
            <surname>Bäckman</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Modality memory across the adult life span: evidence for selective age-related olfactory deficits</article-title>
          .
          <source>Experimental Aging Research</source>
          <volume>24</volume>
          ,
          <issue>1</issue>
          (
          <year>1998</year>
          ),
          <fpage>63</fpage>
          -
          <lpage>82</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Stephen</surname>
            <given-names>Merity</given-names>
          </string-name>
          , Nitish Shirish Keskar, and Richard Socher.
          <year>2018</year>
          .
          <article-title>Regularizing and Optimizing LSTM Language Models</article-title>
          . In International Conference on Learning Representations.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Anelise</surname>
            <given-names>Newman</given-names>
          </string-name>
          , Camilo Fosco, Vincent Casser,
          <string-name>
            <given-names>Allen</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barry McNamara</surname>
            ,
            <given-names>and Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multimodal Memorability: Modeling Efects of Semantics and Decay on Video Memorability</article-title>
          . In Computer Vision - ECCV
          <year>2020</year>
          ,
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Vedaldi</surname>
          </string-name>
          , Horst Bischof, Thomas Brox, and
          <string-name>
            <surname>Jan-Michael Frahm</surname>
          </string-name>
          (Eds.). Springer International Publishing, Cham,
          <fpage>223</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>David</given-names>
            <surname>Ramsay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ishwarya</given-names>
            <surname>Ananthabhotla</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Paradiso</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>The Intrinsic Memorability of Everyday Sounds</article-title>
          . In Audio Engineering Society Conference: 2019
          <source>AES International Conference on Immersive and Interactive Audio.</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Annett</surname>
            <given-names>Schirmer</given-names>
          </string-name>
          , Yong Hao Soh, Trevor B Penney, and
          <string-name>
            <given-names>Lonce</given-names>
            <surname>Wyse</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Perceptual and conceptual priming of environmental sounds</article-title>
          .
          <source>Journal of Cognitive Neuroscience</source>
          <volume>23</volume>
          ,
          <issue>11</issue>
          (
          <year>2011</year>
          ),
          <fpage>3241</fpage>
          -
          <lpage>3253</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Piyush</surname>
            <given-names>Sharma</given-names>
          </string-name>
          , Nan Ding,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Goodman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Radu</given-names>
            <surname>Soricut</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          .
          <fpage>2556</fpage>
          -
          <lpage>2565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Alan</surname>
            <given-names>F Smeaton</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Owen Corrigan</surname>
          </string-name>
          , Paul Dockree, Cathal Gurrin, Graham Healy, Feiyan Hu,
          <string-name>
            <surname>Kevin</surname>
            <given-names>McGuinness</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Eva</given-names>
            <surname>Mohedano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Tomás E</given-names>
            <surname>Ward</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Dublin's participation in the predicting media memorability task at MediaEval 2018</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Wensheng</given-names>
            <surname>Sun</surname>
          </string-name>
          and
          <string-name>
            <given-names>Xu</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Video Memorability Prediction with Recurrent Neural Networks and Video Titles at the 2018 MediaEval Predicting Media Memorability Task.</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Antonia</given-names>
            <surname>Thelen and Micah M Murray</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The eficacy of single-trial multisensory memories</article-title>
          .
          <source>Multisensory Research</source>
          <volume>26</volume>
          ,
          <issue>5</issue>
          (
          <year>2013</year>
          ),
          <fpage>483</fpage>
          -
          <lpage>502</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Antonia</surname>
            <given-names>Thelen</given-names>
          </string-name>
          , Durk Talsma, and
          <string-name>
            <surname>Micah M Murray</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Singletrial multisensory memories afect later auditory and visual object discrimination</article-title>
          .
          <source>Cognition</source>
          <volume>138</volume>
          (
          <year>2015</year>
          ),
          <fpage>148</fpage>
          -
          <lpage>160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Duy-Tue Tran-Van</surname>
            ,
            <given-names>Le-Vu</given-names>
          </string-name>
          <string-name>
            <surname>Tran</surname>
          </string-name>
          , and
          <string-name>
            <surname>Minh-Triet Tran</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Predicting Media Memorability Using Deep Features and Recurrent Network.</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Shuai</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiying</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shizhe Chen</surname>
            , and
            <given-names>Qin</given-names>
          </string-name>
          <string-name>
            <surname>Jin</surname>
          </string-name>
          .
          <year>2018</year>
          . RUC at MediaEval 2018:
          <article-title>Visual and Textual Features Exploration for Predicting Media Memorability.</article-title>
          .
          <source>In Proceedings of MediaEval</source>
          <year>2018</year>
          , CEUR Workshop Proceedings. http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2283</volume>
          /
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>