<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>De-Factify</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>INO at Factify 2: Structure Coherence based Multi-Modal Fact Verification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yinuo Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhulin Tao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xi Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tongyue Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Communication University of China</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Information Engineering, Chinese Academy of Sciences</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>2</volume>
      <issue>2</issue>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper describes our approach to the multi-modal fact verification (FACTIFY) challenge at AAAI2023. In recent years, with the widespread use of social media, fake news can spread rapidly and negatively impact social security. Automatic claim verification becomes more and more crucial to combat fake news. In fact verification involving multiple modal data, there should be a structural coherence between claim and document. Therefore, we proposed a structure coherence-based multi-modal fact verification scheme to classify fake news. Our structure coherence includes the following four aspects: sentence length, vocabulary similarity, semantic similarity, and image similarity. Specifically, CLIP and Sentence BERT are combined to extract text features, and ResNet50 is used to extract image features. In addition, we also extract the length of the text as well as the lexical similarity. Then the features were concatenated and passed through the random forest classifier. Finally, our weighted average F1 score has reached 0.8079, achieving 2nd place in FACTIFY2. The code is available at https://github.com/Catrin-baze/INO-of-factify.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Fact Checking</kwd>
        <kwd>Fake News Detection</kwd>
        <kwd>Multi-modal</kwd>
        <kwd>De-Factify</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, the number of fake news has exploded due to the wide application of social
media and the development of new technologies such as Deepfake. For example, research [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
indicates that fake information appeared in the three months before the election, which favoring
Trump shared a total of 30 million times on Facebook, while those favoring Clinton were shared
8 million times. The rapid development of social media has led to the widespread dissemination
of fake news, which not only afects people’s lives, causes public panic, disrupts social order,
afects public opinion, and manipulates the focus of the public but also damages the credibility
of social media platforms. In 2018, the article "The Science of Fake News" [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] published in
Science stated that falsehood difused significantly further, faster, deeper, and more broadly
than the truth in all categories of information. Therefore, efectively detecting fake news on
social media to suppress the spread of fake news is of great significance for maintaining social
stability and cyberspace security.
      </p>
      <p>
        In order to meet the challenges of fake news, artificial intelligence, and deep learning
technology are likely to play an essential role. Since 2017, the U.S. Defense Advanced Research
Projects Agency(DARPA) has held a "Media Forensics Challenge" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]competition to promote
misinformation detection technology developed rapidly. It is worth noting that the fragmented
digital media environment provides a breeding ground for the uncontrolled dissemination of
misinformation. Multi-modal fake news that combines graphics and text is more dificult to
identify, posing a severe challenge for automatic multi-modal fake news detection. At the same
time, combining multiple modalities has also been applied in various fields, such as
recommendation systems[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Many representative methods[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and datasets have been proposed
to address the multi-modal fake news detection problem. Using multi-modal information to
detect fake news has many advantages as diferent modalities capture diferent dimensions of
the news article, and they can complement each other while evaluating the genuineness of the
article.
      </p>
      <p>In our paper, Our core idea is to thoroughly compare the correlation between claim and
document from multiple perspectives as a basis for classification. Therefore, we propose a
structure coherence-based Multi-Modal Fact Verification. We thoroughly compare the consistency
of the claim and document from structure coherence, which contains four aspects: sentence
length, vocabulary similarity, semantic similarity, and image similarity. Experiments prove the
efectiveness of our method.</p>
      <p>In the rest of this paper, we organize the content as follows. Data description and task
definition are introduced in Section 2. Section 3 introduces related work. Our approach is
described in Section 4, and experiments with results are discussed in Section 5. The conclusion
of our work is presented at the end of the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The Task and Dataset</title>
      <sec id="sec-2-1">
        <title>2.1. Task definition</title>
        <p>
          The multimodal fact verification task FACTIFY2[
          <xref ref-type="bibr" rid="ref9">9</xref>
          ][
          <xref ref-type="bibr" rid="ref10">10</xref>
          ][
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]is essentially modeled as a multimodal
entailment. Each data point contains a claim to be detected and a reliable source document.
Both of them are multimodal and contain an image-text pair. Therefore, the task is to give a
textual claim, claim image, document, and document image. The system has to classify the data
sample into one of the five categories. The descriptions of the labels are as follows.
• Support_Text : the claim text is similar or entailed but images of the document and
claim are not similar;
• Support_Multimodal : both the claim text and image are similar to that of the
document;
• Insufficient_Text : both text and images of the claim are neither supported nor
refuted by the document, although it is possible that the text claim has common words
with the document text;
• Insufficient_Multimodal : the claim text is neither supported nor refuted by the
document but images are similar to the document;
• Refute : The images and/or text from the claim and document are completely
contradictory i.e, the claim is false/fake.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data description</title>
        <p>The data collected from Twitter handles of Indian and US news sources: Hindustan Times1,
ANI2 for India and ABC3, CNN4 for US based on accessibility, popularity and posts per day.
Moreover, these Twitter handles are eminent for their objective and disinterested approach.
The dataset has a total of 50000 samples, and each of the five categories has equal samples. The
dataset has a Train-Val-Test split of 70:15:15.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Related Work</title>
      <sec id="sec-3-1">
        <title>3.1. Fake News Detection</title>
        <p>
          Recently, researchers proposed many efective methods to detect fake news. Qian et
al.[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]showed how text content and semantic information could lead to the detection of fake
news. In [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], the authors used a new framework TM to capture lexical and semantic properties
of both true and fake news text for detecting fake news. Besides, some lawbreakers tamper with
news by forging fake images. So Qi et al.[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]proposed a novel model MVNN which utilizes the
visual information of frequency and pixel domains to detect fake news.
        </p>
        <p>
          Though all the above uni-modal techniques have achieved good results, they still ignore
the truth that news events often contain both text and pictures, which can complement each
1https://twitter.com/htTweets
2https://twitter.com/ANI
3https://twitter.com/ABC
4https://twitter.com/CNN
other. This indicates that we need a multi-modal system for fake news detection. The existing
fake news detection methods based on multi-modal information can be divided into three
categories. Some methods combined textual information with visual information. Singhal et
al.[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] proposed a new model SpotFake to concatenate visual features extracted from the VGG-19
model pre-trained on ImageNet with textual features and classified using a fully connected layer.
Wang et al.[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]built an end-to-end model termed EANN(event adversarial neural network)
to learn and concatenate the textual and visual latent representations and then feed them
into two fully connected neural network classifiers. Other works focus on contrasting
multimodal information.SAFE[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], a similarity-aware multi-modal method, compared the relevance
between the extracted features across modalities to detect fake news. In addition, some methods
utilized multi-modal information enhancement to predict fake news. Jin et al.[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]used Recurrent
Neural Network with an attention mechanism to enhance information understanding between
multi-modal features for efective detection.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Fact Verification</title>
        <p>In general, fact verification is the task of assessing the authenticity of a claim supported by
a validated corpus of documents. Fact-checking introduces objective references and external
knowledge to fake news detection, which is more reliable than relying simply on news text
and visual features to make judgments. Fact verification can also increase users’ awareness of
precaution, which helps them try to find fact-checking information when exposed to fake news.</p>
        <p>
          There are many types of fact-verified datasets, and many of them are uni-modal datasets. For
example, Thorne et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] considered Wikipedia as the source of textual evidence and annotated
the sentences that support or refute each claim. Another popular type of unstructured evidence
dataset often considered is metadata which ofers information complementary to textual sources
[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Although most uni-modal datasets focus on unstructured evidence, structured knowledge
has also been used. Some datasets consist of semi-structured data tables with the ability to
convey information concisely and flexibly. Wang et al.[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] extracted tables from scientific
articles and required evidence selection in the form of cells selected from tables.
        </p>
        <p>
          Recently, fact verification has also begun to consider building multi-modal datasets to improve
the accuracy of fake news detection. A dataset named Mocheg[
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] consisted of textual and
visual evidence and used three labels: support, refute, and NEI. Many fact verification methods
focus on claim verification, which can be seen as natural language inference tasks or a form of
Recognizing Textual Entailment(RTE). Nie et al.[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]used ESIM to verify a claim by concatenating
all pieces of evidence as input and using the max pooling to aggregate the information. Typical
retrieval strategies of RTE include search APIs, Lucene indices etc [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. In [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], dense retrievers
using learned representations and fast dot-product indexing have performed well. Fan et al.
[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] proposed another way to retrieve evidence by using question generation and question
answering via search engine results.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>
        In our paper, we design a structure coherence-based fact-checking method in which the structure
coherence between claims and documents is computed. Structure coherence is reflected in the
following four aspects: literal text similarity, text semantic similarity, text length, and image
similarity. Specifically, the ROUGE[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] is used to extract the literal similarity of the text, two
pre-trained models of CLIP[
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and Sentence BERT[
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] are used to compute the semantic
similarity of the text, the length of the claims and documents are calculated. Then ResNet50 is
used to extract the image features. The above features are spliced and input into the random
forest classifier to obtain the final classification result.
      </p>
      <p>The architecture of the multi-modal model is shown in Figure 2.</p>
      <sec id="sec-4-1">
        <title>4.1. Text Feature Extractor</title>
        <p>
          In the feature extraction module, two features of ROUGE and text length are extracted. ROUGE
stands for Recall-Oriented Understudy for Gisting Evaluation. It is an evaluation metric that
includes measures to automatically determine the quality of a summary by comparing it to other
(ideal) summaries created by humans. The measures count the number of overlapping units
such as n-grams, word sequences, and word pairs between the computer-generated summary to
be evaluated and the ideal summaries created by humans. Since we observe a lot of vocabulary
overlap in the category of refute, ROUGE, which counts the number of overlapping units, is
well suited for our task. At the same time, inspired by Gao et al.[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], we also calculated the
length of the text and found a specific correlation between the text length and the category, so
we also extracted the text length feature.
        </p>
        <p>
          In the pre-training model module, we use the Sentence BERT and cosine similarity in the
baseline[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. In addition, we tried pre-training models CLIP, SimCSE[
          <xref ref-type="bibr" rid="ref29">29</xref>
          ], and RoBERTa[
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]
to extract text features.CLIP is a two-stream model that passes text and vision through the
transformer encoder and calculates the similarity of diferent image-text pairs through linear
projection. At the same time, it uses contrastive learning to convert image classification into
image-text matching tasks. CLIP uses 400 million pairs of graphic and text datasets from the
network and uses text as image labels for training. SimCSE is a simple framework for comparing
sentence vector representations, which shows good performance in sentence embedding tasks.
RoBERTa is an enhanced version of BERT and a fine-tuned version of the BERT model improves
BERT[
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] in terms of model size and training methods. In this part, we tried to use CLIP text
encoder, SimCSE and RoBERTa to extract text features and use cosine similarity to calculate
semantic similarity. We also pass the features extracted by the clip through the MLP layer and
output the three-category result. Finally, we found that using the three-category result of CLIP
combined with MLP as a feature is more conducive to improving the final F1 score.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Image Feature Extractor</title>
        <p>
          In this part, we use the pre-trained ResNet50[
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] and cosine similarity. At the same time,
two other solutions were also tried, using CLIP to extract image features and calculate cosine
similarity; inputting CLIP features into MLP to output three-category results.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Classifier</title>
        <p>Finally, the extracted features are then concatenated and passed through a Random Forest
classifier to output five-category results. A Random Forest is a meta-estimator that fits several
decision tree classifiers on various sub-samples of the dataset and uses averaging to improve
the predictive accuracy and control over-fitting. Among the classifiers provided in the baseline,
we selected the Random Forest classifier that performed best on our features.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments and results</title>
      <sec id="sec-5-1">
        <title>5.1. Experiment Setting</title>
        <p>
          To evaluate the performance of our baseline solutions, we use weighted average F1 for
benchmarking on the validation set. All the experiments are implemented on a single NVIDIA Tesla
T4 GPU with up to 15 GiB RAM. For the CLIP module, we use the pre-trained text encoder
and only train the following MLP layer. MLP contains one hidden layer with 100 nodes and
Adam[
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] optimizer. Because the order of magnitude of text length and other features varies
greatly, we normalize the final features. The final Random Forest classifier with a max_depth
value of 40 performs best.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results</title>
        <p>Our model achieved the F1 score of 0.8079 on the test set, and the competition leaderboard is
shown in Table1. Our method achieves 2nd place in this competition.</p>
        <p>In the process of exploring solutions, our attempts are roughly divided into the following
two aspects: 1) the selection of text pre-training models 2) the ways of using the CLIP.</p>
        <p>During the experiment, we tried to use SimCSE, RoBERTa, and the text encoder of CLIP to
replace Sentence BERT. Specifically, we use these pre-trained models as text feature
extractors and then use cosine similarity to calculate the similarity between claim and document
embeddings. The experimental results of only replacing Sentence BERT were compared with
the baseline. The F1 score on the verification set is shown in Table 2.</p>
        <p>It can be seen that using the text pre-training model to extract text features and simply
calculate the cosine similarity, the F1 scores of other methods are lower than Sentence BERT.
At the same time, we also tried using CLIP to extract image features and calculate cosine
similarity to replace ResNet50. The result only reached an F1 score of 0.5103, far lower than the
baseline. Therefore, simply replacing the pre-trained model without fine-tuning and computing
the cosine similarity does not improve the classification performance. Because diferent
pretraining models essentially only change the representation of the vector, no matter how large the
dimension of the vector is, only one score will be obtained after calculating the cosine similarity.
However, the information contained in a score is limited. So we still keep the structure of
Sentence BERT and ResNet50.</p>
        <p>In addition, we experimented with diferent ways of using CLIP. Because CLIP is a multimodal
model, it contains a text encoder and an image encoder, both of which output a 512-dimensional
vector. Therefore, after using CLIP to extract features from the data, we have a claim, document,
claim image, doc image, and four 512-dimensional vectors. Therefore, we tried three feature
combination methods to use the CLIP module, hoping that it can be used as a feature to improve
the final classification efect.1)Concat the text feature vector into the MLP layer for
threecategory;2)Concat the image feature vector into the MLP layer for three-category;3)Concat
all the image and text feature vectors, and input them into the MLP layer for five-category.
The first two are shown in Figure 3, and the third is shown in Figure 4. The F1 scores on the
validation set using diferent clip modules in our model are shown in Table 3.</p>
        <p>It can be seen that the diferent ways of using the clip module have a more significant impact
on the classification results. Other usage methods will not improve the efect of the model or
even reduce the efect. Because the feature obtained by this module is actually just a multi-class
label. Therefore, the inaccuracy of classification may afect other features. Methods that only
use CLIP to extract image features have the worst results. We consider that classification of
images by CLIP and ResNet50 are mutually exclusive, so adding this feature will reduce the
efect.</p>
        <p>Finally, we conducted ablation experiments on diferent modules in the model, and the results
are shown in Table 4. It can be seen that removing the ResNet50 module has the most significant
impact on the model efect. This result is expected since ResNet50 features are the only image
features in the model. In addition, the most significant contribution to the model is the ROUGE
and text length feature. At the same time, removing any text feature will not significantly
impact the results.Figure5 shows the confusion matrix of the final results on the validation and</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This paper introduces a structure coherence-based approach for the multi-modal fact verification
task. We tried diferent pre-training models such as SimCSE, CLIP, RoBERTa, and diferent ways
of using CLIP multi-modal features. Finally, combining Sentence BERT and CLIP to extract text
features and ResNet50 to extract image features achieved the best F1 score.</p>
      <p>The future work can be carried out in four directions:1)Improve data preprocessing to
address issues such as data noise and multilingualism;2)Our method does not fully use image
features, consider introducing other features to complement ResNet50, or try other methods
in the future;3)Explore better ways to use CLIP or adopt other methods to enhance modal
(a) on the validation set
(b) on the testing set
fusion;4)Improve the universality of the fake news detection model and extend it to more types
of datasets.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Acknowledgement</title>
      <p>The work is supported by the National Key Research and Development Program of
China (No.2020YFB1406800), the Fundamental Research Funds for the Central Universities
(No.CUC22WH002), and the National Natural Science Foundation of China (61702502). We
thank the organizers of DE-FACTIFY 2023 for allowing us to work on the dataset. We also
thank Google colab for providing GPU and deep learning code running platform, and thanks to
skit-learn and hugging face for providing an eficient and convenient deep learning library.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Allcott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gentzkow</surname>
          </string-name>
          ,
          <article-title>Social media and fake news in the 2016 election</article-title>
          ,
          <source>Journal of economic perspectives 31</source>
          (
          <year>2017</year>
          )
          <fpage>211</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Lazer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Baum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Benkler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Berinsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Greenhill</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <string-name>
            <surname>Metzger</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Nyhan</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Pennycook</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Rothschild</surname>
          </string-name>
          , et al.,
          <source>The science of fake news, Science</source>
          <volume>359</volume>
          (
          <year>2018</year>
          )
          <fpage>1094</fpage>
          -
          <lpage>1096</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Fiscus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , et al.,
          <article-title>Media forensics challenge image provenance evaluation and state-of-the-art analysis on largescale benchmark datasets (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Self-supervised learning for multimedia recommendation</article-title>
          ,
          <source>IEEE Transactions on Multimedia</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Elimrec: Eliminating single-modal bias in multimedia recommendation</article-title>
          ,
          <source>in: Proceedings of the 30th ACM International Conference on Multimedia, MM '22</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kabra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kumaraguru</surname>
          </string-name>
          , Spotfake+:
          <article-title>A multimodal framework for fake news detection via transfer learning (student abstract)</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>13915</fpage>
          -
          <lpage>13916</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zafarani</surname>
          </string-name>
          ,
          <article-title>Safe: similarity-aware multi-modal fake news detection (</article-title>
          <year>2020</year>
          ),
          <source>Preprint. arXiv 200304981</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Luo,
          <article-title>Multimodal fusion with recurrent neural networks for rumor detection on microblogs</article-title>
          ,
          <source>in: Proceedings of the 25th ACM international conference on Multimedia</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>795</fpage>
          -
          <lpage>816</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chadha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amitava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Manoj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          , Asif Kumar,
          <article-title>Factify 2: A multimodal fake news and satire news dataset</article-title>
          ,
          <source>in: proceedings of defactify 2: second workshop on Multimodal Fact-Checking and Hate Speech Detection, CEUR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chadha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chinnakotla</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Findings of factify 2: multimodal fake news detection</article-title>
          ,
          <source>in: proceedings of defactify 2: second workshop on Multimodal Fact-Checking and Hate Speech Detection, CEUR</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chopra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <article-title>Benchmarking multi-modal entailment for fact verification</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and Hate Speech Detection, ceur,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sharma</surname>
          </string-name>
          , Y. Liu,
          <article-title>Neural user response generator: Fake news detection with collective user intelligence</article-title>
          ., in: IJCAI, volume
          <volume>18</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>3834</fpage>
          -
          <lpage>3840</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bhattarai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.-C.</given-names>
            <surname>Granmo</surname>
          </string-name>
          , L. Jiao,
          <article-title>Explainable tsetlin machine framework for fake news detection with credibility score assessment</article-title>
          ,
          <source>arXiv preprint arXiv:2105.09114</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Exploiting multi-domain visual information for fake news detection, in: 2019 IEEE international conference on data mining (ICDM)</article-title>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>518</fpage>
          -
          <lpage>527</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kumaraguru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Satoh</surname>
          </string-name>
          ,
          <article-title>Spotfake: A multi-modal framework for fake news detection, in: 2019 IEEE fifth international conference on multimedia big data (BigMM)</article-title>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>39</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yuan</surname>
          </string-name>
          , G. Xun,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          , Eann:
          <article-title>Event adversarial neural networks for multi-modal fake news detection</article-title>
          ,
          <source>in: Proceedings of the 24th acm sigkdd international conference on knowledge discovery &amp; data mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>849</fpage>
          -
          <lpage>857</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Christodoulopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <article-title>Fever: a large-scale dataset for fact extraction and verification</article-title>
          , arXiv preprint arXiv:
          <year>1803</year>
          .
          <volume>05355</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kiesel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Reinartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <article-title>A stylometric inquiry into hyperpartisan and fake news</article-title>
          ,
          <source>arXiv preprint arXiv:1702.05638</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N. X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mahajan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Danilevsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          , Semeval
          <article-title>-2021 task 9: Fact verification and evidence finding for tabular data in scientific documents (sem-tab-facts)</article-title>
          ,
          <source>arXiv preprint arXiv:2105.13995</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>End-to-end multimodal fact-checking and explanation generation: A challenging dataset and models</article-title>
          ,
          <source>arXiv preprint arXiv:2205.12487</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <article-title>Combining fact extraction and verification with neural semantic matching networks</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>6859</fpage>
          -
          <lpage>6866</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cocarascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Christodoulopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <article-title>The fact extraction and verification (fever) shared task</article-title>
          , arXiv preprint arXiv:
          <year>1811</year>
          .
          <volume>10971</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Maillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Karpukhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          , W.-t. Yih,
          <string-name>
            <given-names>B.</given-names>
            <surname>Oğuz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Ghosh, Multi-task retrieval for knowledge-intensive tasks</article-title>
          ,
          <source>arXiv preprint arXiv:2101.00117</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piktus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Petroni</surname>
          </string-name>
          , G. Wenzek,
          <string-name>
            <given-names>M.</given-names>
            <surname>Saeidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bordes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Riedel</surname>
          </string-name>
          ,
          <article-title>Generating fact checking briefs</article-title>
          , arXiv preprint arXiv:
          <year>2011</year>
          .
          <volume>05448</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>C.-Y. Lin</surname>
          </string-name>
          ,
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          , in: Text summarization branches out,
          <year>2004</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>10084</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-F.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Oikonomou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiskovski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bandhakavi</surname>
          </string-name>
          , Logically at the factify 2022:
          <article-title>Multimodal fact verification</article-title>
          ,
          <source>arXiv preprint arXiv:2112.09253</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chopra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
          </string-name>
          , et al.,
          <article-title>Factify: A multi-modal fact verification dataset</article-title>
          ,
          <source>in: Proceedings of the First Workshop on Multimodal Fact-Checking and Hate</source>
          Speech
          <string-name>
            <surname>Detection (DE-FACTIFY)</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Simcse:
          <article-title>Simple contrastive learning of sentence embeddings</article-title>
          ,
          <source>arXiv preprint arXiv:2104.08821</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kasnesis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Heartfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Toumanidis</surname>
          </string-name>
          , G. Sakellari,
          <string-name>
            <given-names>C.</given-names>
            <surname>Patrikakis</surname>
          </string-name>
          , G. Loukas,
          <article-title>Transformer-based identification of stochastic information cascades in social networks using text and image similarity</article-title>
          ,
          <source>Applied Soft Computing</source>
          <volume>108</volume>
          (
          <year>2021</year>
          )
          <fpage>107413</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic optimization</article-title>
          ,
          <source>Computer Science</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>