<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Vision Transformers and Memorable Moments for the Prediction of Video Memorability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mihai Gabriel Constantin</string-name>
          <email>mihai.constantin84@upb.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogdan Ionescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University Politehnica of Bucharest</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper describes the approach taken by the AI Multimedia Lab team for the MediaEval 2021 Predicting Media Memorability task. Our approach is based on a Vision Transformer-based learning method, which is optimized by filtering the training sets for the two proposed datasets. We attempt to train the methods we propose with video segments that are more representative for the videos they are part of. We test several types of filtering architectures, and submit and test the architectures that best performed in our preliminary studies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Media Memorability has attracted the attention of researchers from
diferent domains for a long time. This included studies that revealed
that humans have an uncanny ability of memorizing large quantities
of images, going so far as to correctly encode details from those
images. Generally speaking, there is a certain discrepancy between
the study of image and video memorability, with more attention
given in the current literature to the former. In this context, the
MediaEval Predicting Media Memorability task [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], now at its fourth
edition, creates a common benchmarking task for predicting the
short- and long-term memorability of videos. This task ofers data
extracted and annotated from two datasets – TRECVid [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and
Memento10k [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]; and proposes two opened task, related to direct
memorability prediction and generalization prediction between the
two datasets.
      </p>
      <p>Our proposed method for video memorability prediction relies
on the use of Vision Transformer networks for feature extraction,
a dense network ending for sample regression and an important
frame filtering method that attempts to use the most memorable
moments from the video samples in the training process. The rest of
the paper is organized as follows: Section 2 presents the works most
related to our proposed approach, while our method is presented
in Section 3. Section 4 presents the results both in our training and
development process and on the final testing set. Finally, the main
conclusions are presented in Section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Deep Neural Networks have come a long way in addressing many
machine learning problems, and for a long time, starting with the
success of AlexNet [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], when it came to visual data processing
in general, convolutional neural networks were the norm in
getting the best results. Interestingly, some domains related to the
human perception of media data did not adhere to this general
trend, as concepts like fusion, data manipulation and traditional
feature extractors were sometimes more important in getting good
results than deep neural networks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], indicating a need for deeply
understanding the data and the way it influences human subjects.
      </p>
      <p>
        Recently, Vision Transformers shown their usefulness for image
processing, surpassing convolutional approaches in image
recognition tasks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. To the best of our knowledge, this approach is
relatively untested in the domain of media memorability. This is
perhaps to be expected, as the rise of Vision Transformers is in
itself a novelty at this point in time.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>The general outline of our memorability prediction method is
presented in Figure 1. We propose creating a three stage system. In
the first stage, we theorize that not all frames might be valuable
for memorability calculation and therefore propose a frame
filtering method. Following this, we extract visual features by using a
Vision Transformer architecture, and, in a final step, we perform
regression with a dense MLP architecture.</p>
      <p>Frame filtering. We base our frame filtering system on the
assumption that not all frames are equal when trying to determine
the properties of a larger video sequence. In our case, we propose
using the annotations provided by the organizers for selecting the
frames that may best characterize the video from a memorability
standpoint. We call these frames "Memorable Moments", and while
they may not represent the exact moment or the exact process of
human memory retrieval, we theorize that they may represent a
better approach than simply attempting to use the entire video for
processing.</p>
      <p>We test several setups for the frame filtering method as follows.
First of all, we have to take into account the lag time between
human memory recognition and button press. Therefore, given
 , a user’s response time in milliseconds from the start of the
video, we subtract the following values: 500, 1000, 1500 milliseconds
from the  value in order to define the actual time of retrieval
from memory. Of course we cap the resulting value at zero in case
retrieval occurred very close to the start of the video. Furthermore,
we take a variable number of frames, namely 15, 30, 60 from the
resulting location and use them for analysis. We will compare our
ifltering method (which we call 2) against a default method, where
all the video is taken into consideration (called 1).</p>
      <p>
        Visual Features. For visual feature extractors, we test two popular
Vision Transformer architectures, namely the DeiT [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and the
BEiT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. No special fusion will be employed with these two features,
as at this stage we will test them separately and choose the best
performing one for the final set of experiments.
.
.
.
      </p>
      <sec id="sec-3-1">
        <title>Frame Selection</title>
      </sec>
      <sec id="sec-3-2">
        <title>Vision</title>
      </sec>
      <sec id="sec-3-3">
        <title>Transformer</title>
      </sec>
      <sec id="sec-3-4">
        <title>Dense</title>
      </sec>
      <sec id="sec-3-5">
        <title>Memorability</title>
      </sec>
      <sec id="sec-3-6">
        <title>Score</title>
        <p>Dense MLP. The final stage consists of classifying the chosen
features extracted from the selected frames and outputting the final
memorability score. This is done via a simple dense architecture
with 3 hidden layers of size 1024, 512, and 256.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND ANALYSIS</title>
      <p>In the first stage of development, we test the setups proposed in
the previous Section, by training on the Memento training set and
testing on its development set. With regards to the frame filtering
method, we find that a setup of 1000 milliseconds delay in response
time and 30 frames analyzed are the best setups, though not by a
significant margin. For the Transformer architecture, we select the
DeiT architecture as the best performer in these preliminary tests,
though again not by a large margin.</p>
      <p>The final results computed on the testset are presented in Table
1. It is interesting to notice that, in five out of the six (1, 2)
comparison pairs, the results were better for the variant of the
system which employed filtered training, via Memorable Moments,
while at times even being so with a significant margin.</p>
      <p>
        For the prediction subtask (subtask 1) we find that results for the
Memento10K prediction are much better than the ones for TRECVid.
This may be a result of many factors, but one of them may be
represented by the lower number of video samples in the latter
dataset. Also, continuing the trend recorded at the previous version
of the Predicting Media Memorability task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we observe lower
performance for long-term memorability prediction compared to
short term.
      </p>
      <p>Finally, regarding the generalization subtask (subtask 2), we find
a significant drop in performance when compared to subtask 1. This
may be due to diferences in the types of movies in the dataset, but
methods that reduce this issue must definitely be studied.
5</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS</title>
      <p>In this paper we present a media memorability prediction method
that is based on the use of Vision Transformer architectures and
frame filtering method we call Memorable Moments. Our
experiments show good results for both these components and, for future
developments we propose improving this framework by testing
more feature extraction architectures, performing tests against
convolutional architectures, predicting Memorable Moments on the
testset, and testing this type of approach on other subjective
multimedia concepts and properties.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded under project AI4Media “A European
Excellence Centre for Media, Society and Democracy”, grant 951911,
H2020 ICT-48-2020.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>George</given-names>
            <surname>Awad</surname>
          </string-name>
          , Asad A Butt, Keith Curtis,
          <string-name>
            <given-names>Yooyoung</given-names>
            <surname>Lee</surname>
          </string-name>
          , Jonathan Fiscus, Afzal Godil, Andrew Delgado, Jesse Zhang, Eliot Godard, Lukas Diduch, and others.
          <source>2020. Trecvid</source>
          <year>2019</year>
          :
          <article-title>An evaluation campaign to benchmark video activity detection, video captioning and matching, and video search &amp; retrieval</article-title>
          . arXiv preprint arXiv:
          <year>2009</year>
          .
          <volume>09984</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Hangbo</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Li</given-names>
            <surname>Dong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Furu</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>BEiT: BERT Pre-Training of Image Transformers</article-title>
          .
          <source>arXiv preprint arXiv:2106.08254</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Mihai</given-names>
            <surname>Gabriel</surname>
          </string-name>
          <string-name>
            <given-names>Constantin</given-names>
            ,
            <surname>Liviu-Daniel</surname>
          </string-name>
          <string-name>
            <given-names>Ştefan</given-names>
            , Bogdan Ionescu, Ngoc QK Duong,
            <surname>Claire-Hélène Demarty</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mats</given-names>
            <surname>Sjöberg</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Visual Interestingness Prediction: A Benchmark Framework</article-title>
          and
          <string-name>
            <given-names>Literature</given-names>
            <surname>Review</surname>
          </string-name>
          .
          <source>International Journal of Computer Vision</source>
          (
          <year>2021</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Alba</given-names>
            <surname>García Seco De Herrera</surname>
          </string-name>
          , Rukiye Savran Kiziltepe, Jon Chamberlain, Mihai Gabriel Constantin,
          <string-name>
            <surname>Claire-Hélène</surname>
            <given-names>Demarty</given-names>
          </string-name>
          , Faiyaz Doctor, Bogdan Ionescu, and Alan F Smeaton.
          <year>2020</year>
          .
          <article-title>Overview of MediaEval 2020 Predicting Media Memorability Task: What Makes a Video Memorable?</article-title>
          <source>Proceedings of MediaEval'20</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Alexey</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          , Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, and others.
          <year>2020</year>
          .
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11929</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Rukiye</given-names>
            <surname>Savran</surname>
          </string-name>
          <string-name>
            <given-names>Kiziltepe</given-names>
            , Mihai Gabriel Constantin,
            <surname>Claire-Hélène</surname>
          </string-name>
          <string-name>
            <surname>Demarty</surname>
          </string-name>
          , Graham Healy, Camilo Fosco, Alba García Seco de Herrera, Sebastian Halder, Bogdan Ionescu, Ana Matran-Fernandez,
          <string-name>
            <given-names>Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lorin</given-names>
            <surname>Sweeney</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Overview of The MediaEval 2021 Predicting Media Memorability Task</article-title>
          .
          <source>In Proceedings of MediaEval'21.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          .
          <source>Advances in neural information processing systems</source>
          <volume>25</volume>
          (
          <year>2012</year>
          ),
          <fpage>1097</fpage>
          -
          <lpage>1105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Anelise</given-names>
            <surname>Newman</surname>
          </string-name>
          , Camilo Fosco, Vincent Casser,
          <string-name>
            <given-names>Allen</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barry McNamara</surname>
            ,
            <given-names>and Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multimodal memorability: Modeling efects of semantics and decay on video memorability</article-title>
          .
          <source>In Computer Vision-ECCV</source>
          <year>2020</year>
          : 16th European Conference, Glasgow, UK,
          <year>August</year>
          23-
          <issue>28</issue>
          ,
          <year>2020</year>
          , Proceedings,
          <source>Part XVI 16</source>
          . Springer,
          <fpage>223</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Hugo</given-names>
            <surname>Touvron</surname>
          </string-name>
          , Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and
          <string-name>
            <given-names>Hervé</given-names>
            <surname>Jégou</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Training data-eficient image transformers &amp; distillation through attention</article-title>
          .
          <source>In International Conference on Machine Learning. PMLR</source>
          ,
          <fpage>10347</fpage>
          -
          <lpage>10357</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>