<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RoBERTa-based models for Fact Checking</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yan Zhuang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yanru Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Multimodality, Fake News, Text Similarity</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Shenzhen Institute for Advanced Study</institution>
          ,
          <addr-line>UESTC</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Electronic Science and Technology of China</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>The development of social networks makes it easier and faster to spread news among people, but the spread of some uncertified news can cause great harm. The 'Factify' task of the 'DE-FACTIFY' workshop aims to solve the multi-modal fact verification problem. In this paper, unimodal and bimodal RoBERTabased models for fact checking are proposed. The text-only model integrates disturbance on embedding layer, a new loss function and data augmentation by sequential dropout layers into the vanilla RoBERTa. Based on the text-only model, the text-image model changes the text embedding input into the fusion features of texts and images. The experiment results show that after the introduction of fusion features, the model improves slightly, but the best model is still our text-only model. With the best average F1 score of 75.59%, we improve the baseline (53.10%) by 22% and are finally ranked 2nd.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The development of social media technology allows people to express themselves and receive
information anytime, anywhere. However, in order to attract the attention of users, some
media often publish some eye-catching but unconfirmed news. For example, 77% supporters of
Donald Trump, the former US president, held the opinion that 2020 US presidential election
was manipulated by ”voter fraud” because of the information spread in tweet even though
they don’t have enough evidence [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The situation becomes even worse during the COVID-19
pandemic period. There is an urgent need for a model that can automatically detect whether
the claim is fake or not.
      </p>
      <p>
        Fact checking can be described as, given a claim and some support information, such as
documents, images and other claims, we need to judge whether the claim entails the support
information. Most claims are evidenced-based so that their veracity can be determined by
external knowledge [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It is of great importance to take the evidence, or support information
into consideration since it helps a lot in reasoning in fact checking [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>In this paper, text-only model and text-image model are both proposed. The unimodal model
based on RoBERTa adds disturbance on embedding layer to boost robustness, creates positive
* Corresponding author
samples through sequential dropout layers to augment data and uses a new loss function for
alleviating the dificulty of predicting the image-related labels while the bimodal model based
on RoBERTa fuses the text embedding and image features and promotes the interaction between
two modalities through the self-attention mechanism in transformer. The experiment results
show that both models have good efectiveness. The text-only model performs better and helps
us rank 2nd in the multi-modal fact verification task.</p>
      <p>The rest of the paper is organised as follows. In section 2, related works about fact verification
is provided. Followed by the introduction of the task in section 3. In section 4, the details of
our proposed models are discussed. Section 5 contains the experiments results and analysis of
diferent models. Towards the end, section 6 concludes the paper along with future directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>
        Lots of eforts have been put into fact checking and related research has shifted from single
modal with monolingual texts to with multilingual texts to multi-modal with multi-text and
multi-image [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7 ref8 ref9">4, 5, 6, 7, 8, 9</xref>
        ]. A multi-level inter-sentence attention model shows competitive
performance in ’FEVER’ dataset, which consists of 185k samples with a claim and a supporting
document [
        <xref ref-type="bibr" rid="ref10 ref4">4, 10</xref>
        ]. Multilingual transformer-based models, additional metadata and evidence
from news stories are combined in multilingual dataset ’X-FACT’, which contains 31k short
statements in 25 languages.
      </p>
      <p>
        As for the multimodal situation, [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] shows that augmenting text with image embedding
immediately boosts performance. In Event Adversarial Neural Networks (EANN) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
TextCNN is adapted to extract textual features and pre-trained VGG-19 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] architecture with fully
connected layer is applied to extract visual features. Besides, a fake news detector and a event
discriminator take the concatenated features as input, then predict the label and identify the
event label respectively.
      </p>
      <p>
        Multimodal Variational Autoencoder (MVAE) and EANN have something in common that they
use the same visual feature extractor and take the concatenated features for further prediction
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. However, instead of Text-CNN, MVAE uses recurrent neural networks (RNNs) with
bidirectional Long-Short Term Memory (LSTM) cells to extract textutal features. After sampling
and reconstructing the concatenation of the both features, the model are trained by optimizing
the sum of the reconstruction loss and the Kullback-Leibler (KL) divergence loss.
      </p>
      <p>
        However, both MAVE and EANN ignore the interactions between the textual and visual
features. Vision Transformers (ViT) shows excellent performance in the vision-related tasks
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], based on which, Vision-and-Language Transformer (ViLT) takes fusion of texts and
images into consideration, performs faster and competitive and shows excellent performance in
vision-language classification tasks such as VQA and MSCOCO [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. By introducing
Mixtureof-Modality-Experts (MOME) Transformer to promote deeper modal interactions, Unified
Vision-Language Pre-Training with Mixture-of-Modality-Experts (ULMo) achieves state-of-art
results in the VQA and MSCOCO task [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The ViLT and ULMo focus more on the interactions
between textual and visual features and they may provide a better solution for fact checking.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task setup</title>
      <p>
        For the task, the dataset contains 50k claims with 100k images [
        <xref ref-type="bibr" rid="ref18 ref9">9, 18</xref>
        ]. Given a claim text, claim
image and OCR of the claim image, we need to predict whether they entail the document ones.
According to the diferent entailment, the claims can be classified into 5 categories:
• Support_Text: the claim text entails the document one but claim image not
• Insuficient_Text: both claim text and image are neither entailed nor refuted by the
document ones
• Support_Multimodal: both claim text and image entail document ones
• Insuficient_Multimodal: the claim image entails the document one but claim text not
• Refute: both claim image and text are contradictory with the document ones
In addition, each category accounts for the same percentage with 7k samples for training, 1.5k
samples for validation and 1.5k samples for testing. In Figure. 1, a sample of each category is
provided.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Model</title>
      <p>Models can be divided into text-only ones and text-image ones according to the data they use.</p>
      <sec id="sec-4-1">
        <title>4.1. Text Pre-processing</title>
        <p>There are lots of meaningless words, URLs and diferent language characters in claim texts and
OCRs, so before feeding the text into the model, the following steps are applied:
• URL removal: There are a lot of URL information in the claims, and the information they
contain is worthless and increases the length of the data processed by the model.
• None-English words removal: Many non-English characters are contained in the claims,
especially in OCRs, which rarely helps increase the performance.
• Short words removal: There are a lot of spaces, various characters like “\n, aa” in the
original data. Words with less than 3 characters are removed.</p>
        <p>Since there is less useful information in the OCRs of the images, many of them are “NaN, ANI, BBC”
and the splicing of several words, so the provided OCR data is not used in our model.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Text-only models</title>
        <p>
          Text-only models treat the task as the sentence pairs similarity problem and solve it by classifying
the cosine similarity between the embeddings of the claim and the corresponding document.
SentenceBERT [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] is used as for extracting the text embeddings and serves as text-only model
baseline in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. We use the pre-trained RoBERTa as the backbone and make some modifications
[
          <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
          ]. The models structure can be seen in Figure. 2: After removing URL, non-English words
and short words, the claim text and document text are fed into the transformer, and here we
use vanilla BERT and RoBERTa for comparison. The robustness of the model can be boosted
In vanilla BERT, the final loss can be computed as the sum of the cross entropy loss. However,
since the text-only model only uses text information, it is hard to judge whether the images
entail or not. So we add focal loss to alleviate the dificulty of predicting the image-related
labels [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ]. Here focal loss is defined as Equation 4:
The hyper-parameter  is used to balance the relative importance of positive and negative
samples and  is applied to reduce the weight of easy-to-classify samples, so that the model
focuses more on dificult-to-classify samples during training, which satisfies our need. And
ifnal loss function of our model is defined as Equation
5:
 
 
= −(1 − )
        </p>
        <p>()
 =  
  +  
(1)
(2)
(3)
(4)
(5)
through introducing disturbance on embedding layer, and we use PGD, which iterates several
times to slowly find the optimal perturbation and can be formulated in Equation 1:
|+1</p>
        <p>
          =   /||  ||2
 = ▽  (, ,  )
 
= [ (
1
2
 ( |  )||  ( |  )) +  (
 ( |  )||  ( |  ))]
Here  means the input gradient and is defined as equation 2
There are many other adversarial training methods, such as FGM [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] and FreeAT [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. The
former one can only obtain the locally optimal parameters. Although the latter is also a
stepby-step iterative search for the optimal disturbance, it is updated based on the gradient and
parameters of the previous step, and the parameters found in the current step are suboptimal
and do not maximize the Loss.
        </p>
        <p>
          After we get the last hidden state layer of the CLS in the model, we use a sequencetial network
with two dropout layers to generate another CLS layer so as to generate the positive samples,
and try to minimize the bidirectional KL divergence between the two CLS layers. The above
method is called ’R-Drop’ [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]and can be formulated in Equation 3:


here  denotes the loss weight and we let it equals 4 to compute the loss. The 5 fold
crossvalidation is also adapted in our model and the averaged logits are used for classification.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Text-image models</title>
        <p>
          Existing multimodal models focus more on classification tasks with an image and its description
[
          <xref ref-type="bibr" rid="ref12 ref14 ref15 ref16">12, 14, 16, 15</xref>
          ], and there are relatively few researches dealing with multiple texts and multiple
images. The multimodal baseline model provided in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] computes the cosine similarity between
the text embedding derived from the pre-trained SentenceBERT and between the image features
derived from the pre-trained ResNet50 respectively [27]. The values of the similarity are seen
as the features, as well as the corresponding label as the target, then put into several algorithms
like Random Forest, Decision Tree, and Logistic Regression. The above baseline model ignores
the interactions between diferent modalities.
        </p>
        <p>Our text-image method shares something in common with baseline that they all use
pretrained model to extract embeddings and features. However, the pre-trained RoBERTa instead
of pre-trained SentenceBERT, pre-trained VGG16 instead of pre-trained ResNet50 are applied in
our model for their better representation ability. The diference between our text-image model
and text-only model is the input. In the former model, the image features are concatenated with
the text ones and then put into Multi-Layer Perception for interaction.The whole structure of
our text-image model is shown in Figure 4.3.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments and evaluations</title>
      <p>Here we choose the text-only baseline, pre-trained BERT and RoBERTa as the comparison
with our text-only model, use the multimodal baseline and mixed_input RoBERTa without our
modifications as a comparison with our text-image model. The mixed_input RoBERTa model
takes the fusion of text embeddings and image features, just the way we use in our text_image
model, as the model input. The results of all models are classified after averaging all logits
obtained from 5-fold cross-validation, except for the two baselines. Besides, all hyperparameters
are the same in BERT and RoBERTa models for fair comparison, just as shown in Table 1. The
oficial evaluation for this task is Macro-F1 and the final ranking is based on the weighted
average F1 score. The Macro-F1 scores of the models are shown in Table 2. With the best
average F1 score of 75.59%, we improve the baseline (53.10%) by 22% and are finally ranked
second in this task.</p>
      <p>Noting that the models above our text-image model in Table 2 are all text-only models. And
the column name in the first row of the table except ’Model’ is the first few letters of the
corresponding label, such as ’Sup_Text’ for ’Support_Text’, ’Insufi_Multi’ for
’Insuficient_Multimodal’. The figures in the ’Final’ column denotes the the weighted average F1 scores of the
former 5 categories.</p>
      <p>It can be seen that most model perform better in ’Insuficient_Text’ and ’Support_Multimodal’
label prediction than in ’Support_Text’ and ’Insuficient_Multimodal’ prediction task for that the
judgment basis for the first two labels is that either claim and claim image are both entailed or
both not entailed with the document ones. It shows that the information about the interaction
between two modalities the models learned is not enough. Besides, all models perform perfectly
in predicting the ’Refute’ except baselines because it is relatively easy to distinguish texts with
the opposite meaning.</p>
      <p>The multimodal baseline exceeds text-only baseline over 10% and achieves highest score
in ’Support_Text’ prediction, but compared with our text-only model, it is over 20% less. The
mixed_input RoBERTa model that combines two modalities performs better than the single
modality one. Our text-only model shows the best performance among all models and is
1% higher than vanilla text-only RoBERTa. And our text-image model scores higher than
mixed_input RoBERTa but does not show competitive performance in image-related label
prediction and scores 0.64% less than our text-only model. It is because that the introduction of
the image features in RoBERTa decreases the representation ability and results may be the same
after interacting diferent text embeddings and image features. Besides, the diference of the
magnitudes may cause bias and variances too. In addition, the ensemble model only ensembles
the first 3 models in the Table 2 and performs as well as our text-only model but it costs too
much time.</p>
      <p>Classifying the multiple texts and images is a tough task for it not only involves the entailment
between the texts and texts, images and images, but also between the many texts and images at
the same time. Combining the two modalities improves slightly the understanding of the text
and image pairs. But better interactions and understanding of the two modalities may further
improve the results in future works.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this paper, the unimodal and bimodal RoBERTa-based models are discussed to solve
multimodal fact checking task in De-Factify workshop. The major challenge of fact checking task
derives from the entailment between multiple texts and images, and existing approaches showed
Model</p>
      <p>Sup_Text Insufi_Text Sup_Multi Insufi_Multi
Refute</p>
      <p>Final
unsatisfactory performance. To address the problem, our model integrates the PGD, focal loss
and R-Drop into the RoBERTa model, which shows better efectiveness. Besides, our text-image
model show better performance compared with the vanilla model by fusing the text embedding
and image features, but the efect is still worse than the our text-only model, which helps us
stand 2nd in this task. Better multi-modal feature fusion and interaction strategies are conducive
to the better solving this challenge.
[27] P. Kasnesis, R. Heartfield, X. Liang, L. Toumanidis, G. Sakellari, C. Patrikakis, G. Loukas,
Transformer-based identification of stochastic information cascades in social networks
using text and image similarity, Applied Soft Computing 108 (2021) 107413.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Pennycook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Rand</surname>
          </string-name>
          ,
          <article-title>Examining false beliefs about voter fraud in the wake of the 2020 presidential election</article-title>
          ,
          <source>The Harvard Kennedy School Misinformation Review</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. D. S.</given-names>
            <surname>Martino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babulkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <article-title>That is a known lie: Detecting previously fact-checked claims</article-title>
          , arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>06058</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Lima</surname>
          </string-name>
          ,
          <article-title>Automatic fake news detection: Are models learning to reason?</article-title>
          ,
          <source>arXiv preprint arXiv:2105.07698</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Thorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vlachos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Christodoulopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mittal</surname>
          </string-name>
          ,
          <article-title>Fever: a large-scale dataset for fact extraction and verification</article-title>
          , arXiv preprint arXiv:
          <year>1803</year>
          .
          <volume>05355</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Shu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mahudeswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          , H. Liu,
          <article-title>Fakenewsnet: A data repository with news content, social context and spatialtemporal information for studying fake news on social media</article-title>
          , arXiv preprint arXiv:
          <year>1809</year>
          .
          <volume>01286</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levy</surname>
          </string-name>
          , W. Y. Wang, r/fakeddit:
          <article-title>A new multimodal benchmark dataset for ifne-grained fake news detection</article-title>
          , arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>03854</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Reis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Melo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Garimella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Almeida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eckles</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Benevenuto</surname>
          </string-name>
          ,
          <article-title>A dataset of fact-checked images shared on whatsapp during the brazilian and indian elections</article-title>
          ,
          <source>in: Proceedings of the International AAAI Conference on Web and Social Media</source>
          , volume
          <volume>14</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>903</fpage>
          -
          <lpage>908</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pykl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Guptha</surname>
          </string-name>
          , G. Kumari,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ekbal</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <article-title>Fighting an infodemic: Covid-19 fake news dataset</article-title>
          ,
          <source>in: International Workshop on Combating Online Hostile Posts in Regional Languages during Emergency Situation</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chopra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <article-title>Factify: A multi-modal fact verification dataset</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and
          <article-title>Hate Speech Detection</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kruengkrai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A multi-level attention model for evidence-based fact checking</article-title>
          ,
          <source>arXiv preprint arXiv:2106.00950</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          , G. Ghosh,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shilon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ma</surname>
          </string-name>
          , E. Moore, G. Predovic,
          <article-title>Exploring deep multimodal fusion of text and photo for hate speech classification</article-title>
          ,
          <source>in: Proceedings of the third workshop on abusive language online</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yuan</surname>
          </string-name>
          , G. Xun,
          <string-name>
            <given-names>K.</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          , Eann:
          <article-title>Event adversarial neural networks for multi-modal fake news detection</article-title>
          ,
          <source>in: Proceedings of the 24th acm sigkdd international conference on knowledge discovery &amp; data mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>849</fpage>
          -
          <lpage>857</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          ,
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Khattar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Goud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          , Mvae:
          <article-title>Multimodal variational autoencoder for fake news detection</article-title>
          ,
          <source>in: The world wide web conference</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2915</fpage>
          -
          <lpage>2921</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          , et al.,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11929</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kim</surname>
          </string-name>
          ,
          <article-title>Vilt: Vision-and-language transformer without convolution or region supervision</article-title>
          ,
          <source>arXiv preprint arXiv:2102.03334</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          , Vlmo:
          <article-title>Unified vision-language pre-training with mixture-of-modality-experts</article-title>
          ,
          <source>arXiv preprint arXiv:2111.02358</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P.</given-names>
            <surname>Patwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suryavardan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chopra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakraborty</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ekbal</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ahuja</surname>
          </string-name>
          ,
          <article-title>Benchmarking multi-modal entailment for fact verification</article-title>
          , in: Proceedings of De-Factify: Workshop on Multimodal Fact Checking and
          <article-title>Hate Speech Detection</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>10084</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Madry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Makelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsipras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vladu</surname>
          </string-name>
          ,
          <article-title>Towards deep learning models resistant to adversarial attacks</article-title>
          ,
          <source>arXiv preprint arXiv:1706.06083</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T.-Y. Liu, R-drop:
          <article-title>Regularized dropout for neural networks</article-title>
          ,
          <source>arXiv preprint arXiv:2106.14448</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T.</given-names>
            <surname>Miyato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Goodfellow</surname>
          </string-name>
          ,
          <article-title>Adversarial training methods for semi-supervised text classification</article-title>
          ,
          <source>arXiv preprint arXiv:1605.07725</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shafahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Najibi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Ghiasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dickerson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Studer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Davis</surname>
          </string-name>
          , G. Taylor, T. Goldstein,
          <article-title>Adversarial training for free!</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>T.-Y. Lin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dollár</surname>
          </string-name>
          ,
          <article-title>Focal loss for dense object detection</article-title>
          ,
          <source>in: Proceedings of the IEEE international conference on computer vision</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>2980</fpage>
          -
          <lpage>2988</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>