<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HCMUS at MediaEval2021: Atention-based Hierarchical Fusion Network for Predicting Media Memorability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>E-Ro Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Huynh-Lam</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Triet Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>John von Neumann Institute</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Science</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh city</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Predicting Media Memorability is a task ofered by The Benchmarking Initiative for Multimedia Evaluation in the set of challenges for the MediaEval 2021 Workshop. This task aims at predicting the memorability of visual media to explore the possibility of automated supporting systems in multiple areas of application such as advertisement, recommendations, education, and more. To approximate the memorability score of media, we employ an attention-based fusion network with a hierarchical structure that resembles binary computation trees with the embedding of root nodes used to compute the final memorability score.</p>
      </abstract>
      <kwd-group>
        <kwd>Share weights</kwd>
        <kwd>Share weights</kwd>
        <kwd>Share weights</kwd>
        <kwd>Share weights</kwd>
        <kwd>Memorability score</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The task of Predicting Media Memorability at MediaEval 2021 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
requires participants to automatically predict the probability that a
human may remember a specific visual media of type video after
a specified time period. This task ofered us two datasets for the
training and evaluation of our methods, namely the Memento10k
dataset [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] with short term memorability and the TRECVid dataset
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] with both short term and long term scores.
      </p>
      <p>To aid readers in understanding our approach, we organize our
paper as follow: Section 2 visits some prior works with concepts
related to our approach that might help readers gain preliminary
knowledge; Section 3 introduces the proposed architecture as well
as elaborating details about our network; Section 4 provides detailed
results of our runs together with multiple insights that guided us
through our experiments; Section 5 discuss about the conclusion of
our research and possible future approaches based on our method.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        In Predicting Media Memorability task, participants need to
approximate the probability of each video sample being memorized
by human and hence, this task may be categorized as video
regression with input being 4D features sampled from each video [
        <xref ref-type="bibr" rid="ref5 ref7">5, 7</xref>
        ].
Regression and classification on video has long been studied in
academic literature [
        <xref ref-type="bibr" rid="ref1 ref13 ref14">1, 13, 14</xref>
        ] with many achievements recently
when Transformer-based architecture of neural networks [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] being
applied on this category [
        <xref ref-type="bibr" rid="ref11 ref9">9, 11</xref>
        ].
a man cracks
an egg
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>Figure 1 show an overview of our proposed method. The core
of our method is the attention-based hierarchical fusion network
(AHFNet). Its mechanism is to repeat the process of pairing and
fusing two consecutive visual features into a high-level semantic
feature by a fusion module according to the hierarchical structure
as a binary tree from the leaf to the root node.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Fusion Module</title>
      <p>
        We propose a Fusion Module to fuse two visual features into one
based on the attention mechanism, given those two features are
computed using similar method on diferent inputs. As illustrated
in figure 2, for any two visual features 1, 2 ∈ R × × , we first
add a positional encoding   ∈ R × to each of two features and
employ a multi-headed self-attention [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] which is responsible for
learning the association or correlation among the targets within
each current frame:
      </p>
      <p>Position</p>
      <p>Encoding
 ∈ ℝ$ ×% ×&amp;"
 ∈ ℝ ×
…
 ∈ ℝ×  ∈ ℝ ·
 ∈ ℝ</p>
      <p>% ∈ ℝ$ ×% ×&amp;"
where  ∈ {1, 2}.</p>
      <p>For consistency, 1 =  (ˆ1) and  =  (ˆ2) where ˆ1, ˆ2 ∈
R × × × are inputs while  is the computation acts on subset
of  used to compute 1 and 2, respectively.</p>
      <p>The cross-attention between two visual features are then taken
before fusing both of them into a single feature, helping the network
learn the relationship of those two:</p>
      <p>( ,  ) = (( ), ( ), ( )) (2)
where  ,  ∈ {1, 2}.</p>
      <p>The fusion operator is then applied to merge two visual features
into single one:</p>
      <p>′ =  ((1, 2), (2, 1))
where  may be any reduction operator.</p>
      <p>In our approach, we adopt the summation as Fusion operator:
 ′ = (1, 2) + (2, 1)
(3)
(4)
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Hierarchical Fusion Network</title>
      <p>
        We extracted 8 frames from each video, which were used for our
image-based feature extraction. The ResNet-50 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (pre-trained on
ImageNet) was used to extract a 2048-dimensional feature vector
for each frame. And then, we make pair of the features of the frame
(1, 2), (3, 4), (5, 6), (7, 8) and then fuse each pair with a fusion
module to achieve a higher semantic feature from each pair. Then,
the number of features is reduced by half. We continue doing the
same process until the final feature is fused. The final feature has
the video’s high-level information that can now be used to predict
the memorability score.
      </p>
      <p>In the figure 1 we show only the short version with only 4 frames
with 2-levels of fusion module. However, our work uses 8 frames
with 3-levels of fusion module.
3.3</p>
    </sec>
    <sec id="sec-6">
      <title>Cross-Modal With Text Captions</title>
      <p>
        We use a pre-trained BERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to extract the linguistic features of
each video’s text caption. These features are inserted into Fusion
Module to highlight the visual features that are matched with
corresponding linguistic clues by the CMEM module (Cross-Modal
Excitation Modulation) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The CMEM module is illustrated as a
violet component in Figure 2.
      </p>
    </sec>
    <sec id="sec-7">
      <title>RESULTS AND ANALYSIS</title>
      <p>We have 2 diferent runs for each dataset (TRECVid, Memento10k)
with each type of score (short raw, short normalised, long raw)
in Subtask 01. The first run of each is the AHFNet without the
text captions (AHFNetWTC), and the second is the full version
of AHFNet. Table 1 and 2 show our results on the TRECVid and
Memento10k, respectively.</p>
      <p>With our experiments, we observe that the raw short term
is almost better than the normalised one for both datasets. Our
AHFNetWTC is better on the Memento10k test set. On the TRECVid
test set, however, our AHFNet achieves higher results in all
metrics and score types. These better results can be explained that
our network extracts the TRECVid’s text captions better than the
Memento10k.
5</p>
    </sec>
    <sec id="sec-8">
      <title>CONCLUSION</title>
      <p>This paper describes a hierarchical fusion network with the
attentionbased proposed for the 2021 Predicting Media Memorability task of
MediaEval. The main contributions of this paper are to propose a
fusion module to capture the high-level semantics of two consecutive
frames, leverage the binary hierarchical structure to fuse the video’s
features and highlight the visual features by the corresponding text
caption.</p>
      <p>In the future, we plan to conduct this task with additional features
like audio in videos and a more robust feature extractor. So that
can extract high-level features from dynamic videos efectively.</p>
    </sec>
    <sec id="sec-9">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded by Gia Lam Urban Development and
Investment Company Limited, Vingroup and supported by Vingroup
Innovation Foundation (VINIF) under project code VINIF.2019.DA19.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Sami</given-names>
            <surname>Abu-El-Haija</surname>
          </string-name>
          , Nisarg Kothari,
          <string-name>
            <given-names>Joonseok</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Apostol</surname>
            (Paul) Natsev, George Toderici, Balakrishnan Varadarajan, and
            <given-names>Sudheendra</given-names>
          </string-name>
          <string-name>
            <surname>Vijayanarasimhan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>YouTube-8M: A Large-Scale Video Classification Benchmark</article-title>
          . In arXiv:
          <volume>1609</volume>
          .08675. https://arxiv.org/pdf/1609. 08675v1.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>George</given-names>
            <surname>Awad</surname>
          </string-name>
          , Asad A.
          <string-name>
            <surname>Butt</surname>
            , Keith Curtis,
            <given-names>Yooyoung</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            , Jonathan Fiscus, Afzal Godil, Andrew Delgado, Jesse Zhang, Eliot Godard, Lukas Diduch, Alan F. Smeaton, Yvette Graham, Wessel Kraaij, and
            <given-names>Georges</given-names>
          </string-name>
          <string-name>
            <surname>Quenot</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>TRECVID 2019: An Evaluation Campaign to Benchmark Video Activity Detection, Video Captioning and Matching, and Video Search Retrieval</article-title>
          . (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>2009</year>
          .09984
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . CoRR abs/
          <year>1810</year>
          .04805 (
          <year>2018</year>
          ). arXiv:
          <year>1810</year>
          .04805 http://arxiv.org/abs/
          <year>1810</year>
          .04805
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Zihan</given-names>
            <surname>Ding</surname>
          </string-name>
          , Tianrui Hui, Shaofei Huang, Si Liu, Xuan Luo, Junshi Huang, and
          <string-name>
            <given-names>Xiaoming</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Progressive Multimodal Interaction Network for Referring Video Object Segmentation</article-title>
          .
          <source>The 3rd Largescale Video Object Segmentation Challenge, Workshop in conjunction with CVPR</source>
          <year>2021</year>
          <article-title>(virtual</article-title>
          ).
          <source>(June</source>
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Feichtenhofer</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>X3D: Expanding Architectures for Eficient Video Recognition</article-title>
          . (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>2004</year>
          .04730
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          .
          <source>arXiv preprint arXiv:1512.03385</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Hirokatsu</given-names>
            <surname>Kataoka</surname>
          </string-name>
          , Tenga Wakamiya, Kensho Hara, and
          <string-name>
            <given-names>Yutaka</given-names>
            <surname>Satoh</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Would Mega-scale Datasets Further Enhance Spatiotemporal 3D CNNs? CoRR abs/</article-title>
          <year>2004</year>
          .04968 (
          <year>2020</year>
          ). arXiv:
          <year>2004</year>
          .04968 https://arxiv. org/abs/
          <year>2004</year>
          .04968
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Rukiye</given-names>
            <surname>Savran</surname>
          </string-name>
          <string-name>
            <given-names>Kiziltepe</given-names>
            , Mihai Gabriel Constantin,
            <surname>Claire-Hélène</surname>
          </string-name>
          <string-name>
            <surname>Demarty</surname>
          </string-name>
          , Graham Healy, Camilo Fosco, Alba García Seco de Herrera, Sebastian Halder, Bogdan Ionescu, Ana Matran-Fernandez,
          <string-name>
            <given-names>Alan F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lorin</given-names>
            <surname>Sweeney</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Overview of The MediaEval 2021 Predicting Media Memorability Task</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2021 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Linjie</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Yen-Chun</surname>
            <given-names>Chen</given-names>
          </string-name>
          , Yu Cheng, Zhe Gan,
          <string-name>
            <given-names>Licheng</given-names>
            <surname>Yu</surname>
          </string-name>
          , and Jingjing Liu.
          <year>2020</year>
          .
          <article-title>HERO: Hierarchical Encoder for Video+ Language Omnirepresentation Pre-training</article-title>
          .
          <source>In EMNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Anelise</surname>
            <given-names>Newman</given-names>
          </string-name>
          , Camilo Fosco, Vincent Casser,
          <string-name>
            <given-names>Allen</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barry McNamara</surname>
            ,
            <given-names>and Aude</given-names>
          </string-name>
          <string-name>
            <surname>Oliva</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multimodal Memorability: Modeling Efects of Semantics and Decay on Video Memorability</article-title>
          . (
          <year>2020</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>2009</year>
          .02568
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Chen</surname>
            <given-names>Sun</given-names>
          </string-name>
          , Austin Myers, Carl Vondrick, Kevin Murphy, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>VideoBERT: A Joint Model for Video and Language Representation Learning</article-title>
          . (
          <year>2019</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CV/
          <year>1904</year>
          .01766
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Ashish</surname>
            <given-names>Vaswani</given-names>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
          <string-name>
            <given-names>Aidan N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Łukasz Kaiser, and
          <string-name>
            <given-names>Illia</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is All You Need</article-title>
          .
          <source>In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS'17)</source>
          . Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Limin</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Yuanjun Xiong,
          <string-name>
            <surname>Zhe</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            <given-names>Qiao</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Dahua</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Xiaoou</given-names>
            <surname>Tang</surname>
          </string-name>
          , and Luc Van Gool.
          <year>2019</year>
          .
          <article-title>Temporal Segment Networks for Action Recognition in Videos</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>41</volume>
          ,
          <issue>11</issue>
          (Nov
          <year>2019</year>
          ),
          <fpage>2740</fpage>
          -
          <lpage>2755</lpage>
          . https://doi.org/ 10.1109/TPAMI.
          <year>2018</year>
          .2868668
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Chao-Yuan</surname>
            <given-names>Wu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ross B. Girshick</surname>
          </string-name>
          , Kaiming He,
          <string-name>
            <surname>Christoph Feichtenhofer</surname>
            , and
            <given-names>Philipp</given-names>
          </string-name>
          <string-name>
            <surname>Krahenbuhl</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>A Multigrid Method for Eficiently Training Video Models</article-title>
          .
          <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2020</year>
          ),
          <fpage>150</fpage>
          -
          <lpage>159</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>