<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Bergen, Norway and Online
* Corresponding author.
$ m.agarla@campus.unimib.it (M. Agarla); luigi.celona@unimib.it (L. Celona); raimondo.schettini@unimib.it
(R. Schettini)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Predicting Video Memorability Using a Model Pretrained with Natural Language Supervision</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mirko Agarla</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luigi Celona</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raimondo Schettini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics</institution>
          ,
          <addr-line>Systems and Communication</addr-line>
          ,
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Milano</addr-line>
          ,
          <country country="IT">ITALY</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Video memorability prediction aims to quantify how much a given video content will be remembered over time. The main attributes afecting the prediction of memorability are not yet fully understood and many of the methods in the literature are based on features extracted from content recognition models. In this paper we demonstrate that features extracted from a model trained with natural language supervision are efective for estimating video memorability. The proposed method exploits a Vision Transformer pretrained using Contrastive Language-Image Pretraining (CLIP) for encoding video frames. A temporal attention mechanism is then used to select and aggregate relevant frame representations into a video-level feature vector. Finally, a multi-layer perceptron maps the video-level features into a score. We test several types of encoding and temporal aggregation modules and submit our best solution to the MediaEval 2022 Predicting Media Memorability task. We achieve a correlation of 0.707 in subtask 1 (i.e. the Memento10k dataset). In task 2 we obtain a Pearson correlation of 0.487 by training on Memento10k and testing on videoMem and of 0.529 by training on videoMem and testing on Memento10k.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The exponential growth of images and videos shared on social media platforms require new ways
to organize and retrieve digital contents. Like other video metrics of importance, such as quality
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], aesthetics [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] or interestingness [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], memorability can be regarded as a useful aspect to
help make a choice between competing videos. The Predicting Media Memorability Challenge,
hosted within the MediaEval workshop, focuses on the estimation of video memorability. In
its fourth edition, the task is the same as in previous years, but it involves videos depicting
in-the-wild scenes collected from social media. More information can be found in the challenge
description document [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Most image and video memorability methods based on deep learning
are usually built on top of models pre-trained on ImageNet [
        <xref ref-type="bibr" rid="ref10 ref7 ref8 ref9">7, 8, 9, 10</xref>
        ]. However, we argue
that pre-trained models for semantic content classification may not model important factors
for estimating memorability such as aesthetics and interestingness. Conversely to semantic
categories, natural language can provide a complete description of the video. For this reason we
hypothesize that models trained with natural language supervision might provide richer features
useful for characterizing memorability. In this paper, we exploit the Contrastive
LanguageImage Pretraining (CLIP) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] model for encoding video frames. An attention mechanism is
then proposed for selecting relevant frames. Finally a Multi Layer Perceptron maps the video
features into a memorability score. The experimental results for subtask 1 and subtask 2 of the
MediaEval memorability task demonstrate the efectiveness of the proposed method.
      </p>
      <sec id="sec-1-1">
        <title>CLIP-based</title>
        <p>frame encoder</p>
      </sec>
      <sec id="sec-1-2">
        <title>Temporal</title>
        <p>attention
video frames
frame features
video features
video memorability
score</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Approach</title>
      <p>
        CLIP-based frame encoder. Contrastive Language-Image Pretraining (CLIP) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is a
multimodal model that learns to represent images and text jointly in the same vector space. It
consists of image and text transformer-based encoder networks [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. CLIP shows very good
performance on content classification datasets [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], but it also demonstrate to be efective in
perceptual tasks [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. In our method we exploit the CLIP-based Vision Transformer (ViT)
encoder for video frame encoding, as known as ViT L/14. The ViT extracts 14 × 14 patches
from an input image with size 336 × 336 and outputs a 1024-dimensional feature vector1. Given
a video with  frames, we resize each frame at a resolution of 336 × 336 pixels and feed it into
the ViT. We obtain E = (e1, e2, ..., e ) with E ∈ R × 1024 representing the set of frame-level
feature vectors of the input video. Each feature vector e is normalized by its L2-norm before
further processing.
      </p>
      <p>
        Temporal attention module. The temporal attention module aims at weighting the
contribution of each frame feature vector to obtain the video-level representation. A Bidirectional
Gated Recurrent Unit (Bi-GRU) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] is used for modelling temporal information among frames.
The GRU consists of 6 GRU layers, each layer has an hidden state with size equal to 64 and is
followed by a dropout layer with a probability of 0.2. The set of frame-level features extracted by
CLIP, E, is fed to the GRU which outputs H = (h1, h2, ..., h ) with size  × 128. The matrix
H is converted to a scaling factor w ∈ R × 1 per frame throw a linear layer. The video-level
feature vector e ∈ R1024 is finally obtained as follows:
e = 1 ∑︁ e * w.
      </p>
      <p>
        =1
(1)
Multi-Layer Perceptron. The Multi-Layer Perceptron (MLP) estimates the video memorability
score given the 1024-dimensional video-level feature vector e. It consists of a stack of three
linear layers. The first two linear layers reduce the size of the feature vector first to 512 and
then to 128 dimensions. Each linear layer is followed by a ReLU activation function. The last
linear layer outputs a scalar representing the video memorability score. The sigmoid activation
function is exploited to limit the values in the range [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ].
1https://huggingface.co/openai/clip-vit-large-patch14-336
1.0
0.9
ed0.8
t
ic0.7
d
e
r
P0.6
0.5
0.40.4 0.5 0.6 0G.T7 0.8 0.9 1.0
      </p>
      <sec id="sec-2-1">
        <title>Memento10k/Memento10k</title>
      </sec>
      <sec id="sec-2-2">
        <title>Memento10k/VideoMem</title>
      </sec>
      <sec id="sec-2-3">
        <title>VideoMem/VideoMem VideoMem/Memento10k</title>
        <p>
          2.1. Implementation details
The method is implemented using the PyTorch [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] framework. For each video all  -frames
are considered. During the training phase, the frames are shufled for data augmentation, the
optimizer Adam [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is used with an initial learning rate equal to 1 × 10− 4 which is then reduced
every 5 epochs by 0.95. We train using the L1 criterion and a batch size of 8 for a maximum
of 100 epochs. However, the training process stops when there is no improvement after five
consecutive epochs.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results and Analysis</title>
      <p>
        In this section we present the results achieved on the subtask 1 and the subtask 2. For both tasks
we measure the performance in terms of Pearson’s Linear Correlation Coeficient (PLCC) and
Spearman’s Rank Order Correlation Coeficient (SROCC). The results for the development set of
subtask 1 which consists in training and testing on Memento10k [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] are depicted in Table 1. We
highlight that our best approach (whose details are provided in the previous section) achieves a
PLCC of 0.7132 and an SROCC of 0.7100 on the development set. The performance estimated by
the organizers on the Memento10k test set corresponds to 0.707 for both correlation metrics and
0.005 of mean squared error. For subtask2 which consists in a cross-dataset scenario involving
Memento10k and VideoMem [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] datasets, our best method obtains the performance reported
in Table 2. As expected, in cross-dataset scenario performance decreases by about 20% for both
correlations. It can also be noted that the training on VideoMem allows the method to generalize
better on Mement10k with performance about 10% higher than the one obtained by training on
Memento10k and testing on VideoMem. Figure 2 shows scatter plots on the four training/test
combinations for the development set. The distribution of samples in the diferent plots reflects
what was previously stated, i.e. that apart from the combination Memento10k/VideoMem the
other distributions are well fit. Figure 3 shows the samples with the highest prediction errors. It
is possible to notice that the proposed model tends to overestimate the memorability for such
videos. Particular is the case of Memento10k/VideoMem and VideoMem/VideoMem, for which
the worst error was obtained for the same video.
      </p>
      <p>
        Ablation study. Table 1 shows the results for our less efective solutions. For the encoder, in
addition to the ViT variants, we proposed several approaches that include an I3D model [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
pre-trained on Kinetics-400 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] or Charades [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] used as a feature extractor or finetuned on
memorability. We also proposed a method combining frame-level (ViT) and spatio-temporal
(I3D) features. For the temporal aggregation of the features we experimented with the GRU
used in diferent ways, a transformer and a combination of linear to reduce the dimensionality
of the frame-level features followed by the temporal averaging. Finally, a simple linear layer or
an MLP were tested for the memorability predictor.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Outlook</title>
      <p>
        Our solution involving a CLIP-ViT-L/14 + Bi-GRU attention + MLP achieves better results than
the transformer, the I3D network, and the ViT+I3D. This result confirms our hypothesis that
a model trained with natural language supervision can provide richer features than a model
trained for action recognition. The results can also be attributed to the Bi-GRU-based temporal
attention module allowing the selection of the most relevant video frames. Temporal information
modeling treated as the I3D network allows the model to extract the flow relationship between
frames but lacks relevant semantic information features. Moreover, the limited dataset size and
the small length of the videos, approximately 3, make complex architectures (like the I3D and
the transformer) prone to overfitting. From the cross-dataset results, we can conclude that the
VideoMem dataset allows the model to generalize better on the Memento10k dataset. This is
likely due to the wide variation in content, scenes, and memorability score of the Memento10k
dataset. As future works, we will first exploit a frame sampling algorithm to avoid processing
frames containing redundant information [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Secondly, we will investigate the use of spatial
attention mechanisms for estimating video memorability.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Agarla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Celona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schettini</surname>
          </string-name>
          ,
          <article-title>An eficient method for no-reference video quality assessment</article-title>
          ,
          <source>Journal of Imaging</source>
          <volume>7</volume>
          (
          <year>2021</year>
          )
          <fpage>55</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhattacharya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Nojavanasghari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-F.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <article-title>Towards a comprehensive computational model for aesthetic assessment of videos</article-title>
          , in: International Conference on Multimedia, ACM,
          <year>2013</year>
          , pp.
          <fpage>361</fpage>
          -
          <lpage>364</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Nieto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Celona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. F.</given-names>
            <surname>Labrador</surname>
          </string-name>
          ,
          <article-title>Understanding aesthetics with language: A photo critique dataset for aesthetic assessment</article-title>
          ,
          <source>in: Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gygli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Grabner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Riemenschneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nater</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Van Gool</surname>
          </string-name>
          ,
          <article-title>The interestingness of images</article-title>
          , in: ICCV, IEEE,
          <year>2013</year>
          , pp.
          <fpage>1633</fpage>
          -
          <lpage>1640</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-D. Ştefan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>N. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Duong</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sjöberg</surname>
          </string-name>
          ,
          <article-title>Visual interestingness prediction: a benchmark framework and literature review</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          <volume>129</volume>
          (
          <year>2021</year>
          )
          <fpage>1526</fpage>
          -
          <lpage>1550</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sweeney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Fosco</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>García Seco de Herrera</surname>
            , S. Halder,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Healy</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Matran-Fernandez</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Smeaton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Sultana, Overview of the MediaEval 2022 predicting video memorability task</article-title>
          , in: MediaEval Multimedia Benchmark Workshop Working Notes,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Leonardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Celona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Napoletano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bianco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Schettini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Manessi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rozza</surname>
          </string-name>
          ,
          <article-title>Image memorability using diverse visual features and soft attention</article-title>
          ,
          <source>in: ICIAP</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>171</fpage>
          -
          <lpage>180</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Perera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zelnik-Manor</surname>
          </string-name>
          ,
          <article-title>Is image memorability prediction solved?</article-title>
          , in: CVPR Workshops, IEEE/CVF,
          <year>2019</year>
          , pp.
          <fpage>0</fpage>
          -
          <lpage>0</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Casser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McNamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          , Multimodal memorability:
          <article-title>Modeling efects of semantics and decay on video memorability</article-title>
          ,
          <source>in: ECCV</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sweeney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Healy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          ,
          <article-title>The influence of audio on video memorability with an audio gestalt regulated video memorability system</article-title>
          ,
          <source>in: International Conference on Content-Based Multimedia Indexing (CBMI)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          , et al.,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          , et al.,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , in: ICLR,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. C.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Loy</surname>
          </string-name>
          ,
          <article-title>Exploring clip for assessing the look and feel of images</article-title>
          ,
          <source>in: AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hentschel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hotho</surname>
          </string-name>
          ,
          <article-title>Clip knows image aesthetics</article-title>
          ,
          <source>Frontiers in Artificial Intelligence</source>
          <volume>5</volume>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. van Merriënboer</given-names>
            ,
            <surname>Ç. Gulçehre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bougares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Learning phrase representations using rnn encoder-decoder for statistical machine translation</article-title>
          ,
          <source>in: EMNLP</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1724</fpage>
          -
          <lpage>1734</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paszke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chintala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>DeVito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Desmaison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerer</surname>
          </string-name>
          ,
          <article-title>Automatic diferentiation in pytorch (</article-title>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>in: ICLR</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Cohendet</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>N. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Duong</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Engilberge</surname>
          </string-name>
          ,
          <article-title>Videomem: Constructing, analyzing, predicting short-term and long-term video memorability</article-title>
          , in: ICCV, IEEE/CVF,
          <year>2019</year>
          , pp.
          <fpage>2531</fpage>
          -
          <lpage>2540</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Carreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Quo vadis, action recognition? a new model and the kinetics dataset</article-title>
          ,
          <source>in: proceedings of the IEEE CVPR</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>6299</fpage>
          -
          <lpage>6308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>W.</given-names>
            <surname>Kay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hillier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vijayanarasimhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Viola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Green</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Back</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Natsev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Suleyman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>The kinetics human action video dataset</article-title>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Sigurdsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schmid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Farhadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Alahari</surname>
          </string-name>
          ,
          <article-title>Charades-ego: A large-scale dataset of paired third and first person videos</article-title>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>