<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HCMUS at MediaEval 2021: Fine-tuning CLIP for Automatic News-Images Re-Matching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thien-Tri Cao</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nhat-Khang Ngo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thanh-Danh Le</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tuan-Luc Huynh</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ngoc-Thien Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hai-Dang Nguyen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minh-Triet Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>John von Neumann Institute</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Science</institution>
          ,
          <addr-line>VNU-HCM</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vietnam National University</institution>
          ,
          <addr-line>Ho Chi Minh city</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Matching text and images based on their semantics have an essential role in cross-media retrieval. The NewsImages task, MediaEval2021, explores the challenge of building accurate and high-performance algorithms. We proposed diferent approaches leveraging the advantages of fine-tuning CLIP for the multi-class retrieval task. With our approach, the best-performed method reaches a recall@100 score of 0,77441.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        In the context of journalism, authors often use images to represent
the main content of a particular article. A study in 2020 indicates
that the textual content and accompany images might not be related
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Many previous studies in multimedia and recommendation
system domains mostly investigate image-text pairs with simple
relationships, e.g., [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The MediaEval 2021 NewsImages Task calls for
researchers to investigate the real-world relationship of news text
and images in more depth, in order to understand its implications
for journalism and news recommendation system [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
      </p>
      <p>The HCMUS-team participates in the Image-Text-Re-Matching
task. Particularly, given a set of image-text pairs in the wild, the
task requires us to correctly re-assign images to their decoupled
articles, with the aim to understand the implication of journalism
in choosing illustrative images.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Learning correspondences between images and texts are
challenging because of their representation discrepancies. A majority of
studies focus on connecting objects with corresponding
semantic words in sentences. Lee et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed Stack-Cross
attention mechanism to find correspondence scores between objects
and words. As an improvement, Liu et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] introduced a
graphstructured network to capture both image-sentence level relations
and object-word level correspondences. On the other hand, Wang et
al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] combines early and late fusion strategies. The incorporation
helps models to learn both intra-modal and inter-modal information
eficiently.
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>
        CLIP(Contrastive Language–Image Pre-training)[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is proposed by
Radford et al. It is a powerful pretrain-model for text-image
matching tasks. In our survey, CLIP is the best choice as the baseline for
ifne-tuning. The CLIP model has been trained with more than 400
million text-image pairs, and the dataset domain is huge, covering
the dataset portion of the NewsImage task. The inference result of
the CLIP model (without training with NewsImage dataset) for the
dataset provided by the organizers is out-performance compared to
the models that we built ourselves or using CLIP as the backbone
and train it with NewsImage dataset. In addition, the number of
text-image pairs in the NewsImage dataset is relatively small and
does not represent the specificity of the dataset. Therefore, we
decided not to retrain the model with the NewsImage Task dataset
but just use it as an evaluation dataset and fine-tune the number
of words, the preprocessing step based on the performance of the
model on this dataset. Our fine-tuning takes place at a step that
determines how many words to include in the model as well as
which words should be kept or discarded. Basically, our approach
consists of 4 steps:(1) translation, (2) text preprocessing, (3) image
and text vectorization, (4) feed to CLIP, and (5) evaluation.
3.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Translation</title>
      <p>The language used in articles in NewsImage dataset is German, but
the language in CLIP is English, so we need to translate all articles
into English. Google translate is a useful API to help us do this as
it is free and highly accurate.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Text preprocessing</title>
      <p>
        The conventional preprocess includes dropping NA instances,
converting categorical labels into numerical labels, converting all text
into lowercase, etc. Additionally, we also expand contractions such
as "He’s", "She’s" and remove some words like "an","a","the". We
believe this extra preprocessing works will help extract even more
useful information for our embedding features. Finally, Ekphrasis
library [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] helps us segment words that are intentionally or
unintentionally written and correct misspellings or typos for cleaner
text. After the preprocessing step, we determined the number of
words to be fed into the model as we realized it greatly afected the
model’s performance. Basically, we gradually adjust the number of
words fed into the model and observe the change of performance
of model. Experimental results on its influence will be described in
more detail in the Experiments and Experimental results sections.
4.1
      </p>
    </sec>
    <sec id="sec-6">
      <title>EXPERIMENTS AND EXPERIMENTAL</title>
    </sec>
    <sec id="sec-7">
      <title>RESULTS</title>
    </sec>
    <sec id="sec-8">
      <title>Experiments</title>
      <p>We have submitted five runs for this task. Basically, they are all
generated from CLIP model but difer in the number of words
of the article included in the model, resulting in diferent results.
From run 1 to run 4, corresponding to the number of words of
each article that we feed into CLIP is 10, 20, 30, and 40 words.
We fine-tuned the word count of each article because during our
experiments on the NewsImage set, we noticed that, as we gradually
increased the number of words feed into the CLIP, the recall@1
gradually decreased while the recall@100 increased. That said, the
influence of the word count in an article on the model’s performance
is significant. The last run is the average ensemble submission
combines all the results of run 1 to run 4 methods. In this method,
all runs have the same weights of 0.25.
4.2</p>
    </sec>
    <sec id="sec-9">
      <title>Experimental Results</title>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION AND FUTURE WORKS</title>
      <p>News Images is a dificult task when it requires exactly matching
the image with the text for nearly 2000 pairs, but we have obtained
relatively satisfactory results with 0,77441 for the MeanRecall@100
scale. This demonstrates the eficiency of the model structure, as
well as the benefits that the pretrain-model brings, when the dataset
used to train in the NewsImage task is not too large. In the future,
we wish to investigate more methods and delve into this topic as it
is a potential field that still has many problems to be solved.</p>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded by Gia Lam Urban Development and
Investment Company Limited, Vingroup and supported by Vingroup
Innovation Foundation (VINIF) under project code VINIF.2019.DA19</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Özlem</given-names>
            <surname>Özgöbek Duc Tien Dang Nguyen adn Mehdi Elahi Andreas Lommatzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Benjamin</given-names>
            <surname>Kille</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>News Images in MediaEval 2021</article-title>
          .
          <source>In Proc. of the MediaEval 2021 Workshop</source>
          . Online. (
          <year>2021</year>
          ). https: //multimediaeval.github.io/editions/2021/tasks/newsimages/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Christos</given-names>
            <surname>Baziotis</surname>
          </string-name>
          , Nikos Pelekis, and
          <string-name>
            <given-names>Christos</given-names>
            <surname>Doulkeridis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>DataStories at SemEval-2017 Task 4: Deep LSTM with Attention for Messagelevel and Topic-based Sentiment Analysis</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-</source>
          <year>2017</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          , Vancouver, Canada,
          <fpage>747</fpage>
          -
          <lpage>754</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>MD. Zakir</given-names>
            <surname>Hossain</surname>
          </string-name>
          , Ferdous Sohel, Mohd Fairuz Shiratuddin, and
          <string-name>
            <given-names>Hamid</given-names>
            <surname>Laga</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A Comprehensive Survey of Deep Learning for Image Captioning</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>51</volume>
          ,
          <issue>6</issue>
          , Article 118 (feb
          <year>2019</year>
          ),
          <volume>36</volume>
          pages. https://doi.org/10.1145/3295748
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Kuang-Huei</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Xi Chen, Gang Hua, Houdong Hu, and
          <string-name>
            <given-names>Xiaodong</given-names>
            <surname>He</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Stacked cross attention for image-text matching</article-title>
          .
          <source>In Proceedings of the European Conference on Computer Vision (ECCV)</source>
          .
          <volume>201</volume>
          -
          <fpage>216</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Chunxiao</given-names>
            <surname>Liu</surname>
          </string-name>
          , Zhendong Mao, Tianzhu Zhang, Hongtao Xie,
          <string-name>
            <given-names>Bin</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Yongdong</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Graph structured network for image-text matching</article-title>
          .
          <source>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>10921</fpage>
          -
          <lpage>10930</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Nelleke</given-names>
            <surname>Oostdijk</surname>
          </string-name>
          , Hans van Halteren, Erkan Bas, ar, and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>The Connection between the Text and Images of News Articles: New Insights for Multimedia Analysis</article-title>
          .
          <source>In Proceedings of the 12th Language Resources and Evaluation Conference. European Language Resources Association</source>
          , Marseille, France,
          <fpage>4343</fpage>
          -
          <lpage>4351</lpage>
          . https: //aclanthology.org/
          <year>2020</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>535</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Alec</given-names>
            <surname>Radford</surname>
          </string-name>
          , Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Learning Transferable Visual Models From Natural Language Supervision</article-title>
          .
          <source>CoRR abs/2103</source>
          .00020 (
          <year>2021</year>
          ). arXiv:
          <volume>2103</volume>
          .00020 https: //arxiv.org/abs/2103.00020
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Yifan</given-names>
            <surname>Wang</surname>
          </string-name>
          , Xing Xu,
          <string-name>
            <surname>Wei</surname>
            <given-names>Yu</given-names>
          </string-name>
          , Ruicong Xu,
          <string-name>
            <given-names>Zuo</given-names>
            <surname>Cao</surname>
          </string-name>
          , and Heng Tao Shen.
          <year>2021</year>
          .
          <article-title>Combine Early and Late Fusion Together: A Hybrid Fusion Framework for Image-Text Matching</article-title>
          .
          <source>In 2021 IEEE International Conference on Multimedia and Expo (ICME)</source>
          .
          <source>IEEE</source>
          , 1-
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>