<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SELAB-HCMUS at MediaEval 2023: A cross-domain and subject-centric approach towards the memorability prediction task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Quang Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Huy Trinh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Huy-Giap Bui</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khac-Trieu Vo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Triet Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thien-Phuc Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Information Technology, University of Science - VNU-HCM</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Software Engineering Lab, University of Science - VNU-HCM</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Viet Nam National University</institution>
          ,
          <addr-line>Ho Chi Minh City</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The captivating field of Human Memorability extends across various dimensions, each ofering distinctive insights. In this paper, we undertake a comprehensive exploration, examining the unique impact of individual domains on the complex fabric of human memorability. Within this inquiry, we introduce two innovative yet straightforward methodologies crafted not only to spark the reader's interest but also to achieve remarkable results in our analytical pursuits. The intricacies of human memorability are delved into, presenting fresh perspectives and inventive approaches to enrich the discourse on this compelling subject.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Due to the explosive quantities of data from social media and short content platforms in recent
years, Media Memorability has attracted more research on the retention of users and their
cognitive reactions towards the content. Whereas multiple previous works emphasized the
visual domain of the media, we reckon that humans perceive the media in a multimodal fashion
with a deeper understanding of the content.</p>
      <p>Event-Related Potentials (ERPs) and Event-Related Spectral Perturbations (ERSPs) are crucial
tools in neuroscience, ofering precise insights into brain processing timing and location. They
are pivotal in various applications, such as understanding learning disorders and enabling
thought-controlled devices. In educational and clinical settings, ERPs and ERSPs provide
invaluable insights into the intricate workings of the mind, advancing our understanding of
brain function, one brainwave at a time.</p>
      <p>
        The study of video memorability has diverse applications, including education, content
retrieval, summarization, storytelling, and advertising. This motivates the organization of
Video Memorability Prediction in MediaEval [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Participants in this task will need to predict
memorability scores for videos using a dataset containing annotations, visual features, and EEG
recordings to gauge their memorability over both short and long periods.
      </p>
      <p>The authors introduce eficient and promising methods spanning diverse domains for
predicting video memorability (for Subtasks 1 and 2), emphasizing their lightweight nature and
straightforward structure. Additionally, a novel technique in Subtask 2 is presented and
thoroughly assessed for its robustness.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Various factors, especially media features, influence memorability. Models using visual or texture
features have been proposed, and combining diverse models often improves performance. For
example, Guinaudeau et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] merged visual models (ResNet and DenseNet) with a texture
model (Sentence-BERT), achieving commendable results. Insights from this study [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] suggest
that highly memorable videos tend to have saturated colors, people and faces, manipulable
objects, and man-made environments, while less memorable videos often feature darker settings,
clutter, or inanimate scenes.
      </p>
      <p>
        Contemporary computer vision faces challenges in generality and usability, requiring
additional labeled data. An alternative involves learning directly from raw text associated with
images. Radford et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] show caption prediction’s eficacy, learning image representations
from a dataset of 400 million image-text pairs. This enables natural language use in
downstream tasks. Alec et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] verify CLIP’s [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] transferability across 30+ datasets, showcasing its
performance without specific data training.
3. Experimental setup and results
Motivations
For Subtask 1, we are convinced that an ordinary subject would perceive the video as a
multimodal object – the video is embedded with content in various forms, including visual, semantic,
and temporal content. The user may require multiple content domains to assemble an
‘impression’ of the video to memorize it thoroughly. Therefore, we utilize the nature of the videos to
extract their corresponding features and use them to predict the memorability of each video.
      </p>
      <p>For Subtask 2, we recognize that memorability is personalized and varies among individuals
based on their preferences for diferent aspects of video content. Instead of focusing on the
content, we aim to understand and identify the most memorable moments for each person in
each corresponding video, emphasizing a user-centric approach.
3.1. Subtask 1: Using cross-domain features to estimate memorability score</p>
      <sec id="sec-2-1">
        <title>Method</title>
      </sec>
      <sec id="sec-2-2">
        <title>Spearman</title>
      </sec>
      <sec id="sec-2-3">
        <title>Pearson</title>
        <p>Resnet + EficientNet + CLIP Text</p>
        <p>CLIP (Text + Vision)</p>
        <p>NGram
0.313
0.445
0.336
0.326
0.452
0.350</p>
        <p>MSE
0.006
0.008
0.008</p>
        <p>N-gram, ResNet, and EficientNet</p>
        <p>
          Initially, we used a simple yet eficient statistical method to analyze words from captions in
the Memento10k dataset. We followed the N-gram approach, aligning with previous research.
This method served as our starting point. Our results show how words are expressed plays a
crucial role in making video content more memorable. Significantly, our approach outperformed
the use of isolated features from ResNet [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and EficientNet [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          In the provided Memento10k dataset, each video is paired with five lowercase,
stop-wordremoved, punctuation-free, and lemmatized captions. The text underwent N-gram processing
(unigrams, bigrams, trigrams), with memorability scores assigned based on the methodology in
[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. This method involves computing mean memorability scores, scaled by N-gram frequency.
The predicted memorability score for each video in the testing set is the mean score of its
N-gram. The final predicted memorability score is derived from a weighted sum of N-gram
components: 80% from unigrams, 12% from bigrams, and 8% from trigrams.
        </p>
        <sec id="sec-2-3-1">
          <title>Video</title>
          <p>Caption1
Caption2
Caption3
Caption4
Caption4</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>First frame</title>
        </sec>
        <sec id="sec-2-3-3">
          <title>Middle frame</title>
        </sec>
        <sec id="sec-2-3-4">
          <title>Last frame</title>
          <p>CLIP text
ResNet
EfficientNet</p>
          <p>FC layer</p>
          <p>FC layer</p>
          <p>FC layer
BN ReLU</p>
          <p>BN ReLU
FC layer
2048x768
FC layer</p>
          <p>BN ReLU</p>
          <p>BN ReLU
FC layer
768x256
FC layer</p>
          <p>Dropout</p>
          <p>p=0.2
Dropout</p>
          <p>p=0.2
FC layer
FC layer
256x256 Dropout</p>
          <p>p=0.2
BN ReLU</p>
          <p>BN ReLU
1536x768
768x256
256x256</p>
          <p>FC layer</p>
          <p>Sigmoid</p>
          <p>Score</p>
          <p>Comparing CLIP and N-gram, CLIP outperformed as a text feature extractor, leading us
to select CLIP for the next experiment. For CLIP, five captions yield five feature vectors
concatenated into a single 2560-length vector. Visual-based features are extracted from three
frames (first, middle, last), with ResNet and EficientNet generating vectors of lengths 2048 and
1536, respectively. Integration involves concatenating the three vectors from each model.</p>
          <p>Fusing CLIP text and image encoding</p>
          <p>Previous results using CLIP for textual feature extraction have proven efective for the task.
This motivates us to use CLIP for visual feature extraction. Because CLIP is trained on text-image
pairs, it can learn the nuanced connections between them. And because both the text and image
are encoded into the same representation space, we can use a much simpler architecture for
fusing them, which might reduce over-fitting and generalize better to the test dataset.</p>
          <p>C
F
t
a
c
n
o
C
C
F
Reduction layer</p>
          <p>Merge layer
Regression layer
m
r
o
N
h
c
t
a
B
C
F
m
r
o
N
h
c
t
a
B</p>
          <p>U
L
e
R
m
r
o
N
h
c
t
a
B
d
i
o
m
g
i
S</p>
          <p>U
L
e
R</p>
          <p>Video</p>
          <p>Text
First frame
Middle frame
Last frame
Caption 1
Caption 2
Caption 3
Caption 4
Caption 5
Image
r
e
d
o
c
n
e
e
g
a
m
iI
P
L
C
r
e
d
o
c
n
e
ttI
x
e
P
L
C</p>
          <p>CLIP embedding</p>
          <p>Feature 1
Feature 2
Feature 3
Feature 4
Feature 5
Feature 6
Feature 7
Feature 8
r
e
y
a
l
n
o
it
c
u
d
e
R
r
e
y
a
l
e
g
r
e
M
r
e
y
a
l
n
o
i
s
s
e
r
g
e
R</p>
          <p>Result</p>
          <p>Instead of training two separate networks for text and visual features, we trained a single
network on the combined features. We reduced dimensionality on the features before
concatenating them, as shown in Figure 2. Since CLIP features are noticeably smaller than ResNet’s
and EficientNet’s features, we only used a single linear layer to reduce the dimension of each
feature from 512 to 128, which is shared across both textual and visual features of CLIP. This
drastically reduces the number of parameters of the model. We suspect the model might have
dificulty learning but can generalize better.</p>
          <p>Similarly to the previous approach, the first, last, and middle frames are extracted and encoded
into CLIP’s embedding, along with five captions. Batch normalization is applied to help the
model converge faster.
3.2. Subtask 2: Regression on neural signals from ERP and ERSP</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Method AOC ERP one sensor (FC6) signals 0.536 ERP all 28 sensors signals 0.540</title>
        <p>ERP all 28 sensors signals (with subject one-hot encoding) 0.657</p>
        <p>We introduce two regression methodologies applied to the provided dataset within this
specific subtask. The initial approach involves employing a straightforward linear regression
step to assimilate the features inherent in ERPs (28 x 30 input features to 1 output tensor) and
ERSPs (28 x 30 x 30 input features to 1 output tensor) on the collected data from diferent sensors.
For ERSPs, an additional normalization step is undertaken before the commencement of the
training process. The results of the method are shown in the table 2.</p>
        <p>In the second method, we adopt a personalized approach, training models for each individual
due to unique neural signals observed in the dataset. We believe that individual cognitive
responses vary when interacting with video content. Using one-hot encoding, we create
individualized vectors for each subject, adding 12 more features (for 12 test subjects) to the
original 28 × 30 ERP features. The results in Table 2 highlight the method’s high eficiency.
4. Discussion and Outlook
The main findings from Subtask 1 highlight that text-based features are more efective than
visual-based features in determining a video’s impact on human-memorability. This emphasizes
the importance of providing diverse captions for each video to enhance predictability, as well as
the need to consider multiple factors from various perspectives. It’s crucial to acknowledge that
the disparities in backgrounds and content between the training set and test set raise questions
about the robustness of the findings.</p>
        <p>In Subtask 2, distinct patterns in ERP and ERSP diagrams suggest that video memorization is
encoded in neural signals. These unique patterns call for individualized investigation to achieve
optimal results, highlighting the importance of conducting a detailed analysis of the neural
responses of each participant to improve predictive performance.</p>
        <p>Acknowledgments
This research was funded by Vingroup and supported by Vingroup Innovation Foundation
(VINIF) under project code VINIF.2019.DA19.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Fosco</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          <string-name>
            <surname>Seco de Herrera</surname>
            , S. Halder,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Healy</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Matran-Fernandez</surname>
            ,
            <given-names>R. S.</given-names>
          </string-name>
          <string-name>
            <surname>Kiziltepe</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Smeaton</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Sweeney</surname>
          </string-name>
          ,
          <article-title>Overview of The MediaEval 2023 Predicting Video Memorability Task</article-title>
          ,
          <source>in: Proceedings of the MediaEval 2023 Workshop</source>
          , Amsterdam, The Netherlands,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Guinaudeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Girbau</surname>
          </string-name>
          <string-name>
            <surname>Xalabarder</surname>
          </string-name>
          ,
          <article-title>Textual Analysis for Video Memorability Prediction</article-title>
          ,
          <source>in: the 13th MediaEval Multimedia Benchmark Workshop</source>
          , Bergen, Norway,
          <year>2023</year>
          . URL: https: //universite-paris-saclay.
          <source>hal.science/hal-04091024.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fosco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Casser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. A.</given-names>
            <surname>McNamara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          , Multimodal memorability:
          <article-title>Modeling efects of semantics and decay on video memorability</article-title>
          , CoRR abs/
          <year>2009</year>
          .02568 (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2009</year>
          .02568. arXiv:
          <year>2009</year>
          .02568.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: Proceedings of the 38th International Conference on Machine Learning, PMLR139</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Maji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kalogerakis</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Learned-Miller, Multi-view convolutional neural networks for 3d shape recognition</article-title>
          ,
          <source>in: 2015 IEEE International Conference on Computer Vision</source>
          (ICCV),
          <source>IEEE Computer Society</source>
          , Los Alamitos, CA, USA,
          <year>2015</year>
          , pp.
          <fpage>945</fpage>
          -
          <lpage>953</lpage>
          . URL: https://doi.ieeecomputersociety.
          <source>org/10</source>
          .1109/ICCV.
          <year>2015</year>
          .
          <volume>114</volume>
          . doi:
          <volume>10</volume>
          .1109/ICCV.
          <year>2015</year>
          .
          <volume>114</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G. G.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: Proceedings of the 38th International Conference on Machine Learning, PMLR139</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          , Eficientnet:
          <article-title>Rethinking model scaling for convolutional neural networks</article-title>
          ,
          <source>in: Proceedings of the 36th International Conference on Machine Learning</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6105</fpage>
          -
          <lpage>6114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>M. M. A. Usmani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Zahid</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Tahir</surname>
          </string-name>
          ,
          <article-title>Quest for insight: Predicting memorability based on frequency of n-grams (</article-title>
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>