<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modelling of Video Memorability using Ensemble Learning and Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Muhammad Mustafa Ali Usmani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sumaiyah Zahid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muhammad Atif Tahir</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National University of Computer Emerging Sciences, (NUCES-FAST), Karachi Campus</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The modeling of video memorability is still an open challenge for researchers in machine learning. This paper presents our methodology for the MediaEval 2022: Predicting Media Memorability Challenge. The proposed approach investigated ensemble learning methods using the pre-extracted image features: Alexnet, Resnet, and Densenet. In addition to that Transformers and TF-IDF modeling were done on text features. Further, image and text features were ensembled using late fusion that helped in predicting the video memorability on the Memento10k dataset with an accuracy of 0.661.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        The study [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] investigated the efect and dependency of video image color, brightness, and hue
on predicting media memorability, in comparison to complex data-driven representations such
as image classification, composition, and recognition. It is observed that high-level
representations (image composition, recognition, and classification) are best suited to predict media
memorability.
      </p>
      <p>
        The studies [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ] provide knowledge of the most commonly remembered media are war-like
scenes, nature, and open spaces [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The brain’s capability to remember any piece of media is
highly dependent on the abstraction of the same level of scene and object representations. This
helps in understanding the remembered media would include the cinematic and high-level object
attributes. The best-suited models are the Transformers models providing better results [
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7, 8, 9</xref>
        ]
due in-depth representation of input features. These models are proposed as an alternative to
the other neural architectures.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Approach</title>
      <p>The models were applied to the textual features provided as the video description in the
Memento10k dataset. Several machine learning techniques are applied based on text inputs
as well as visual information. The video frames were modeled and only the best-performed
models were used to predict the media memorability. The multi-modal approach is applied to
encode textual and visual features. These embeddings are provided to linear regression models
to predict the best media memorability.</p>
      <p>The pre-extracted features are provided as image features from video frames of the beginning,
middle and last of each video dataset. The machine learning model is applied where the
topscoring pre-extracted features are selected. The top resulting features include ALEX-NET,
RES-NET, DENSE-NET, EFFICIENT-NET, and VGG. Further analysis is carried out by applying
XG Boost, ADA Boost, Random Forest, KNN, and MLP (Multi-layer Perceptron) where ADA
Boost outperformed and provide the best results. Figure 1 shows the complete understanding
of the approach taken to accurately predict the media memorability.</p>
      <sec id="sec-3-1">
        <title>3.1. Textual Features</title>
        <p>
          The video caption or information helps in understanding the context of the video frame using
semantic techniques. These features help in providing a baseline understanding of the
video/image frame. TF-IDF[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is a statistical formulation to understand the similarity of words in a
document. This depends on the number of words and frequency of each word appearing in the
same document. The TF-IDF is used to predict memorability as it provides the best among other
techniques using textual information.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Text Transformers</title>
        <p>
          Text Transformers provides another approach that helps in understanding the semantics of a
video. There are diferent text transformers that are applied over the provided Memento10k
dataset. The applied transformers help in getting the similarities and dissimilarities among
the sentences which helped in identifying topics in text data. We got the best result using
DistilBERT[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], so for further analysis and ensemble, we consider results from only DistilBERT
Transformer. Since the text transformers are able to generate synthetic texts, therefore, this
helped in automatic text generation. The extracted textual features are used to build language
representation. The text transformers help in the classification of images when the image
features and textual features are combined whereas the text encoders help to encode sentences
to understand the context of the video.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Image Features</title>
        <p>
          Videos are nothing but a continuous transition of images. For image-related features, 3 frames
of every video were considered i.e. first, last and middle. Our approach consisted of learning the
best pre-extracted feature that is already provided in the Memento10k dataset. After training
on several machine learning models, only the top 3 image features were selected which were
giving best accuracies on validation data. Namely, Alexnet[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], Resnet[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], and Densenet [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
outperformed the remaining image features. We ran multiple machine-learning models on
these 3 image features, and Adaboost outperform all of them. Therefore for further analysis, we
consider only the result of Adaboost.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Ensemble Learning</title>
        <p>The ensemble combines more than one machine learning model that outperforms in predicting
more accurately. This technique also provides a concept of a multi-classification model where the
resulting factor to predict is based on combining multiple models. After running the base model
on diferent features, late fusion was applied. Linear regression and average-out techniques
were used in the ensemble step.
3.4.1. Runs Details
Run 1: Only image features were considered and then an average was taken of all 3 of them wto
visualize the efect of only image features in predicting media memorability. Run 2: Average
of Transformers and TF-IDF was taken to see if there was any efect of the image features
as captions were the depiction of the same video. Run 3: Image features and TF-IDF scoring
were considered and averag was taken. Run 4: All the best models were considered, from the
image/video baseline Ada Boost was picked for all the features, and for text, Transformers and
TF-IDF both were considerd. And then late fusion was done by taking the average of them, this
gave us an improvement on the baseline methods. Run 5: Results from image features together
with TF-IDF were passed to a Linear Regression Model.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results and Analysis</title>
      <p>The results elaborate on Spearman’s correlation coeficient, Pearson’s correlation coeficient,
and Mean Square Error (MSE) values over the validation and testing set of data. The results in
Table 1 show the obtained correlation against the validation and the results in Table 2 shows
the obtained correlation on testing of the dataset of Memento10k for each submitted execution.</p>
      <p>It is observed that Spearman’s correlation coeficient started predicting with 0.421 accuracies
and suddenly improved in the prediction accuracy in its second execution and till the fifth
execution. The maximum accuracy gained from Spearmen’s correlation coeficient is 0.661.
On the other side, Pearson’s correlation coeficient provides a little improved accuracy when
compared with Spearman’s correlation coeficient. Pearson’s correlation coeficient obtained
an accuracy of 0.439 on the test dataset improved quickly in next following executions. The
maximum obtained accuracy from Pearson’s correlation coeficient is 0.667. Furthermore, the
MSE also reduced from 0.021 to 0.006 during several executions on the testing dataset.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>The approach of multi-classification through ensemble provides a better approach to predicting
media memorability. The textual features and visual features predict well when modeled through
ensemble learning. The textual features provide the semantic context while the visual features
provide the understanding of each image through pre-extracted features from the Memento10k
dataset. The accuracy to predict the media memorability increased during the testing dataset in
its second execution by changing the hyperparameters of the machine learning models. TF-IDF
is already a better approach to predict media memorability but our approach showed that TF-IDF
performed way better while modeling through Linear Regression.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgements</title>
      <p>This work was supported in part by the Higher Education Commission (HEC) Pakistan, and in
part by the Ministry of Planning Development and Reforms under the National Center in Big
Data and Cloud Computing.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Arnheim</surname>
          </string-name>
          ,
          <article-title>Art and Visual Perception: A Psychology of the Creative Eye</article-title>
          , University of California Press,
          <year>1974</year>
          . URL: http://www.amazon.com/exec/obidos/redirect?tag=
          <fpage>citeulike07</fpage>
          -
          <lpage>20</lpage>
          &amp;path=ASIN/ 0520243838.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sweeney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Constantin</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Demarty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Fosco</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>García Seco de Herrera</surname>
            , S. Halder,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Healy</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Matran-Fernandez</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Smeaton</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Sultana, Overview of the MediaEval 2022 predicting video memorability task</article-title>
          , in: MediaEval Multimedia Benchmark Workshop Working Notes,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Isola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <article-title>What makes a photograph memorable?</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>36</volume>
          (
          <year>2014</year>
          )
          <fpage>1469</fpage>
          -
          <lpage>1482</lpage>
          . doi:
          <volume>10</volume>
          .1109/ TPAMI.
          <year>2013</year>
          .
          <volume>200</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaegle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mehrpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mohsenzadeh</surname>
          </string-name>
          , T. Meyer, A. Oliva,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rust</surname>
          </string-name>
          ,
          <article-title>Population response magnitude variation in inferotemporal cortex predicts image memorability</article-title>
          ,
          <source>eLife</source>
          <volume>8</volume>
          (
          <year>2019</year>
          )
          <article-title>e47596</article-title>
          . URL: https: //doi.org/10.7554/eLife.47596. doi:
          <volume>10</volume>
          .7554/eLife.47596.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Konkle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Brady</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <article-title>Conceptual distinctiveness supports detailed visual long-term memory for real-world objects</article-title>
          ,
          <source>J Exp Psychol Gen</source>
          <volume>139</volume>
          (
          <year>2010</year>
          )
          <fpage>558</fpage>
          -
          <lpage>578</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Konkle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Brady</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Alvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          ,
          <article-title>Scene memory is more detailed than you think: the role of categories in visual long-term memory</article-title>
          ,
          <source>Psychol Sci</source>
          <volume>21</volume>
          (
          <year>2010</year>
          )
          <fpage>1551</fpage>
          -
          <lpage>1556</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          ,
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          , CoRR abs/
          <year>1810</year>
          .04805 (
          <year>2018</year>
          ). URL: http://arxiv.org/abs/
          <year>1810</year>
          .04805. arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Houlsby</surname>
          </string-name>
          ,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , CoRR abs/
          <year>2010</year>
          .11929 (
          <year>2020</year>
          ). URL: https://arxiv.org/ abs/
          <year>2010</year>
          .11929. arXiv:
          <year>2010</year>
          .11929.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Krueger</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          ,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          , in: M.
          <string-name>
            <surname>Meila</surname>
          </string-name>
          , T. Zhang (Eds.),
          <source>Proceedings of the 38th International Conference on Machine Learning</source>
          ,
          <string-name>
            <surname>ICML</surname>
          </string-name>
          <year>2021</year>
          ,
          <volume>18</volume>
          -
          <issue>24</issue>
          <year>July 2021</year>
          ,
          <string-name>
            <given-names>Virtual</given-names>
            <surname>Event</surname>
          </string-name>
          , volume
          <volume>139</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          . URL: http://proceedings.mlr.press/v139/radford21a. html.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sammut</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. I.</surname>
          </string-name>
          Webb (Eds.),
          <source>TF-IDF</source>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>US</given-names>
          </string-name>
          , Boston, MA,
          <year>2010</year>
          , pp.
          <fpage>986</fpage>
          -
          <lpage>987</lpage>
          . URL: https: //doi.org/10.1007/978-0-
          <fpage>387</fpage>
          -30164-8_
          <fpage>832</fpage>
          . doi:
          <volume>10</volume>
          .1007/978-0-
          <fpage>387</fpage>
          -30164-8_
          <fpage>832</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Debut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chaumond</surname>
          </string-name>
          , T. Wolf,
          <article-title>Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>
          , CoRR abs/
          <year>1910</year>
          .01108 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1910</year>
          .01108. arXiv:
          <year>1910</year>
          .01108.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          , in: F. Pereira,
          <string-name>
            <given-names>C.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bottou</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          Weinberger (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>25</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2012</year>
          . URL: https://proceedings.neurips.cc/ paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>CoRR abs/1512</source>
          .03385 (
          <year>2015</year>
          ). URL: http://arxiv.org/abs/1512.03385. arXiv:
          <volume>1512</volume>
          .
          <fpage>03385</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <article-title>Densely connected convolutional networks</article-title>
          ,
          <source>CoRR abs/1608</source>
          .06993 (
          <year>2016</year>
          ). URL: http://arxiv.org/abs/1608.06993. arXiv:
          <volume>1608</volume>
          .
          <fpage>06993</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>