<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Chuck Schumer was one of Hedil
Fleiss' top clients. Look it up.
Doug Masters (@protestertrophy)
January</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>UNIPI-NLE at CheckThat! 2020: Approaching Fact Checking from a Sentence Similarity Perspective Through the Lens of Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lucia C. Passaro</string-name>
          <email>lucia.passaro@fileli.unipi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Bondielli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Lenci</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Marcelloni</string-name>
          <email>francesco.marcelloni@unipi.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Florence</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Pisa</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>23</volume>
      <issue>2019</issue>
      <abstract>
        <p>This paper describes a Fact Checking system based on a combination of Information Extraction and Deep Learning strategies to approach the task named \Veri ed Claim Retrieval" (Task 2) for the CheckThat! 2020 evaluation campaign. The system is based on two main assumptions: a claim that veri es a tweet is expected i) to mention the same entities and keyphrases, and ii) to have a similar meaning. The former assumption has been addressed by exploiting an Information Extraction module capable of determining the pairs in which the tweet and the claim share at least a named entity or a relevant keyword. To address the latter, we exploited Deep Learning to re ne the computation of the text similarity between a tweet and a claim, and to actually classify the pairs as correct matches or not. In particular, the system has been built starting from a pre-trained Sentence-BERT model, on which two cascade ne-tuning steps have been applied in order to i) assign a higher cosine similarity to gold pairs, and ii) classify a pair as correct or not. The nal ranking produced by the system is the probability of the pair labelled as correct. Overall, the system reached a 0.91 MAP@5 on the test set.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The great proliferation of online misinformation and fake news in the last few
years encouraged the development of several fact-checking initiatives by various
actors including journalists, governments, organizations, and companies. In the
past, fact-checking was typically performed manually, resulting in the collection
of large amounts of annotated resources for this speci c task. More recently,
researchers have started to use such resources with the aim of training automatic
fact-checking systems [
        <xref ref-type="bibr" rid="ref19 ref25 ref33">19, 25, 33</xref>
        ]. A common phenomenon in social media is that
viral claims often come back after a while [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], increasing the probability that
a particular claim has been previously fact-checked by a trusted organization.
Therefore, systems able to decide whether a claim has been already fact-checked
have become particularly relevant, because they contribute to breaking down the
costs of verifying both old and new viral claims. In this scenario, the
CLEF2020CheckThat! task 2 [
        <xref ref-type="bibr" rid="ref1 ref2 ref26 ref8">1, 2, 8, 26</xref>
        ] has been organized with the goal of supporting
journalists and fact-checkers when trying to determine whether a claim has been
already fact-checked.
      </p>
      <p>The goal of the task is speci ed as follows: \Given a check-worthy claim and
a dataset of veri ed claims, rank the veri ed claims, so that those that verify
the input claim (or a sub-claim in it) are ranked on top".3</p>
      <p>
        This paper describes a system that approaches such task by exploiting a
combination of Information Extraction (IE) and Deep Learning (DL) strategies
to associate a tweet with the most probable claim that veri es it. The task is
indeed strongly related to the concepts of information extraction and text
similarity. Intuitively, to guess if two claims are related to each other, it is important
to establish whether i.) they share some linguistic properties (e.g., mentioned
entities) and ii.) they are in general semantically similar. In order to deal with
i.), traditional IE methods are very useful and accurate [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] when extracting
information such as Named Entities (e.g., persons, locations and organizations)
and content words (e.g., nouns, verbs). On the other hand, ii.) requires a deeper
representation of text meaning, which can be obtained with Neural Language
Models (NLMs) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. State-of-the-art NLMs [
        <xref ref-type="bibr" rid="ref12 ref22">12, 22</xref>
        ] based on Transformer
architectures and attention mechanisms [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] have become increasingly popular in the
last couple of years, thanks to their ability to model whole text sequences and
generate pre-trained representations that can be ne-tuned for di erent tasks.
An important feature of the word representations produced by such models is
that they are contextualized (i.e., they di er depending on the word context),
thereby improving model performance in tasks based on word [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and sentence
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] similarity.
      </p>
      <p>The rest of the paper is organized as follows: Section 2 presents an overview of
both fact-checking frameworks and NLP resources relevant for the task. Section
3 describes the proposed approach to solve the fact-checking task, consisting in
the creation of two ne-tuned models to handle the claim semantic relatedness.
Sections 4 and 5 focus on results and discussion, respectively. Finally, Section 6
draws some conclusions and describes future research directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        A key aspect of the process of building trustworthy data sets of fake and reliable
news is actually how the Fact-Checking process is performed [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In the last years,
several approaches have been proposed for di erent purposes. For example, Fact
Check Explorer, developed by Google, 4 browses and searches for fact checks
      </p>
      <sec id="sec-2-1">
        <title>3 https://github.com/sshaar/clef2020-factchecking-task2</title>
      </sec>
      <sec id="sec-2-2">
        <title>4 http://toolbox.google.com/factcheck/explorer</title>
        <p>
          by exploiting mentions and topics and by o ering several lters to re ne the
queries. Similarly, ClaimsKG [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] o ers a Knowledge Graph to search the claims
containing particular named entities or keyphrases.
        </p>
        <p>
          Over many years, fact-checking has been performed manually by
journalists, by exploiting available tools online [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. Earliest work on automated
factchecking de ne the task as the assignment of a truth value to a claim made in
a particular context [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. Most of the approaches on automated fact-checking
exploit the reliability of a source and the stance of its claims with respect to
other claims and already veri ed information. The assignment of the truth value
is often based on the way in which particular claims (or rumors) are spread on
social media [
          <xref ref-type="bibr" rid="ref13 ref28 ref7 ref9">7, 9, 13, 28</xref>
          ] or on the Web [
          <xref ref-type="bibr" rid="ref17 ref20">17, 20</xref>
          ]. Other approaches use Wikipedia
[
          <xref ref-type="bibr" rid="ref18 ref30">18, 30</xref>
          ] or other knowledge graphs [
          <xref ref-type="bibr" rid="ref11 ref27">11, 27</xref>
          ] to fact-check claims. More recently, a
novel approach has been proposed that exploits Sentence-BERT [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] to re-rank
claims [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] in order to predict whether a claim has been fact-checked before.
        </p>
        <p>
          Indeed, DL models have proven to be among the most e ective techniques
for Language Modelling. In addition, the availability of new DL architectures
such as Transformers [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] has led to a signi cant performance improvement in a
wide range of NLP tasks. Transformers have two main advantages over previous
Language Modelling architectures. First, thanks to the attention mechanism
each element of a text sequence can access information of all the other elements.
Thus, the meaning of words (and sentences) in context can be modelled more
e ectively. Second, the Transformer architecture is geared towards exploiting the
full potentialities of transfer learning for NLP tasks. The idea behind transfer
learning is that the knowledge learnt on a more general task can be exploited to
specialize a model on new problems for which the amount of data is much more
limited. One particular instance of transfer learning consists of two di erent
training paradigms, namely pre-training and ne-tuning. During pre-training,
language models are typically trained with an unsupervised learning tasks on
vast collections of textual data. For example, models can be trained to predict
speci c words in a sequence based on their surrounding context, and to predict
whether two sentences are sequential or not [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. During ne-tuning, the
pretrained model is further trained, this time for a limited number of epochs, on
supervised learning tasks such as for example sequence labeling or sequence pair
classi cation. The main idea is that, the initial weights (or a subset thereof)
of the pre-trained model, are further adjusted to model the ne-tuning task.
Typically, the pre-training step is very time-consuming and computationally
expensive, but the same resulting model can be used as a starting point to solve
a wide range of tasks. On the other hand, the ne-tuning step is less
resourcedemanding and requires less labelled data.
        </p>
        <p>
          Transformer-based architectures such as BERT [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and XLNet [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] have
obtained state-of-the-art results in most NLP tasks they have been applied to. One
advantage of such architectures is that they learn contextualized representations
that allow models to capture word polysemy. Conversely, more traditional
language models such as Skip-Gram and Continuous-Bag-of-Words algorithms [
          <xref ref-type="bibr" rid="ref15 ref16 ref4">4,
15, 16</xref>
          ] learn non-contextualized embeddings and store a single vector for each
word type belonging to the training set, independently of its context. Moreover,
Transformers have been also exploited to obtain context-aware sentence
representations that have been proven to enable semantic comparison of sentences
with promising performances [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The UNIPI-NLE approach</title>
      <p>Given a tweet and a set of already veri ed claims (vclaims), the goal of the
task is to predict, for every target tweet-vclaim pair, the likelihood of the vclaim
verifying the tweet. Indeed, among the target tweet-vclaim pairs, there exists
only a gold pair whose vclaim veri es the tweet, which therefore is a correct
match. The goal is achieved by ranking, for each tweet, the claims that are more
likely to verify it. The dataset is composed of three elements:
1. the veri ed claims used for fact checking, each of them provided with an
identi er, a title, and the actual claim;
2. the training tweets, associated with an identi er and a textual content;
3. the correct pairing between tweets and veri ed claims.</p>
      <p>The training set consists of 1; 003 tweets (803 for training, 200 for development)
and 10; 373 already veri ed claims. The test set consists of 200 additional tweets.</p>
      <p>
        The UNIPI-NLE system is based on two main assumptions: the claims that
verify a tweet are expected to mention the same entities and keyphrases and
should have a similar meaning. To address the rst point, among target
tweetvclaim pairs, we identify the subset of candidate pairs (also referred as potential
pairs) in which the tweet and the vclaim share at least a named entity or a
content word. We refer to the step of identifying candidate pairs among target
ones as the IE step. The DL modules described below have been fed only with the
portion of the dataset consisting of such candidate pairs. In order to estimate the
text similarity between a tweet and a vclaim, we exploit Siamese BERT networks
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] to create a language model that is able to better deal with sentence-level
textual similarity. This model is then used to learn if a claim can be used to
verify a tweet. In particular, we perform two cascade ne-tuning steps aimed
at i.) assigning a higher cosine similarity to gold tweet-vclaim pairs and ii.)
actually classifying a target tweet-vclaim pair, and more speci cally a candidate
tweet-vclaim pair, as a correct match (gold) or not.
      </p>
      <p>
        Figure 1 shows the neural components of the system architecture. The
rst white stack (bert-base-uncased + bert-base-nli-mean-tokens)
consists of the pre-trained model released by [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and trained on SNLI [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and MultiNLI dataset [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] to create universal sentence embeddings. The
black boxes show our two ne-tuned models: Our Sentence-BERT model
(bert-base-nli-factcheck-cos) follows a training paradigm similar to the one
described in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], but it is speci cally geared to assigning a higher cosine
similarity to gold tweet-vclaim pairs. The last level of the architecture represents
the nal Transformer-based classi er trained to decide, given a candidate
tweetvclaim pair, whether the tweet is actually veri ed by that claim or not. The
classi er ne-tunes bert-base-nli-factcheck-cos on the fact-checking task,
by labelling candidate tweet-vclaim pairs as correct matches (gold) or not.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Information Extraction (IE) step</title>
        <p>
          Starting from the assumption that similar claims tend to mention the same
entities and keyphrases, we developed an IE module to nd potential
tweetvclaim pairs. Such a module is based on Stanza [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], a state of the art natural
language analysis package. We processed each text fragment (i.e., a tweet, a
vclaim or a vclaim title) with Sentence Splitting, PoS-tagging, Lemmatization,
and Named Entity Recognition. Thus, each text is associated with its keywords,
consisting of its content words (nouns, verbs, and adjectives) and named entities.
        </p>
        <p>Given a tweet, in order to retrieve potential claims that verify it, we used
two di erent functions based on the keywords:</p>
        <p>IE function { the overlapping score is simply computed by counting the
number of elements (cf. the keywords eld in Table 1 and Table 2) shared by
the tweet and the claim. Candidate tweet-vclaim pairs are required to share
at least one lowercased element (named entity or content word).
IEElastic function { it exploits Elasticsearch5 to nd the potential
candidate pairs. Speci cally, for each tweet, candidate claims consist of the top
1; 000 matches ranked by relevance, using the scoring function provided by</p>
        <sec id="sec-3-1-1">
          <title>5 https://www.elastic.co/</title>
          <p>vclaim and title
keywords
[`chuck schumer', `hollywood',
title: Was Sen. Chuck Schumer a `heidi eiss's', `chuck schumer',
Client of `Hollywood Madam' `hollywood', `heidi eiss', `sen.',
Heidi Fleiss? `chuck', `schumer', `phone',
`numvclaim: Sen. Chuck Schumer's ber', ` nd', `hollywood', `madam',
name and/or phone number were `heidi', ` eiss', `black', `book',
found in "Hollywood Madam" `client', `sen.', `chuck', `schumer',
Heidi Fleiss's black book of clients. `client', `hollywood', `madam',
`heidi', ` eiss']
keywords
[`chuck schumer', `hedil eiss',
`doug masters', `january 23, 2019',
`chuck', `schumer', `hedil', ` eiss',
`client', `look', `doug', `masters',
`@protestertrophy', `january']
the task organizers for the baseline. Such scoring function is an Elasticsearch
multi-match query based on both the vclaim and its title and the tweet itself.
Candidate tweet-claim pairs obtained with the IE overlapping function were
used to train the model bert-base-nli-factcheck-cos. Candidate tweet-claim
pairs obtained with the IE and IEElastic functions have been used at inference
time to obtain the nal predictions submitted for evaluation, namely
respectively T2-EN-UNIPI-NLE-BERT2IE (contrastive run) and
T2-EN-UNIPI-NLEBERT2IEElastic (primary run). Moreover, we used both the IE and IEElastic
functions to simply rank the potential claims associated with each tweet
according to the overlapping score. The score provided by IEElastic corresponds to the
task o cial baseline. The results of these rankings are reported in Section 4.
3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Fine-tuning of Transformer models</title>
        <p>
          Siamese BERT networks [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] can be used to create language models specialized
on tasks related to Semantic Textual Similarity (STS) [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. In order to train our
system to recognize gold tweet-vclaim pairs, we rst model the textual similarity
between tweets and claims belonging to the same candidate tweet-vclaim pairs.
To this purpose, we started from one of the available ne-tuned Sentence-BERT
models,6 namely bert-base-nli-mean-tokens [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. Such model was originally
        </p>
        <sec id="sec-3-2-1">
          <title>6 https://github.com/UKPLab/sentence-transformers</title>
          <p>
            trained on SNLI [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] and MultiNLI dataset [
            <xref ref-type="bibr" rid="ref34">34</xref>
            ] and tested on the STSbenchmark
[
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. The training phase was performed by classifying a pair of sentences with
the labels entail, contradict, and neutral while evaluation was performed on the
STSbenchmark [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] dataset, which contains sentence pairs and their similarity
score. The trained model was exploited to infer sentence pair similarity via cosine.
The bert-base-nli-mean-tokens achieved 77:12 Pearson correlation with gold
scores on the STSbenchmark test set.
          </p>
          <p>We added two levels of ne-tuning to this model, in order to adapt the
sentence pair similarity task to the fact-checking one (i.e., gold tweet-vclaim
pairs are associated with the maximum cosine similarity), as well as to fact-check
a pair with a classi cation layer (i.e., gold tweet-vclaim pairs are associated with
the positive label). To this end, we exploited both the vclaim text content and
its title, namely the vclaim title.</p>
          <p>
            In fact, each vclaim is provided with a title that can be considered as a
summary of the vclaim itself and therefore, very similar to it. The usage of both
the vclaim and vclaim title for training has two main advantages. First, it allows
to increase the size of the dataset so that the model is shown more positive
examples, that are under-represented. Second, it helps to add variability to the
training examples, both positive and negative. For example, a title may contain
an acronym such as \KKK", whereas the claim may contain its extended form, in
this case \Ku Klux Klan". In our experiments we noticed that such a variability
was very helpful to improve the overall performances of our models.
Sentence pair similarity We modeled the ne-tuning step like Reimers and
Gurevych [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ] to estimate the semantic similarity between two sentences. The
authors used the STSbenchmark [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] dataset, containing pairs of sentences with
a similarity score ranging from 0 (no similarity) to 5 (maximum similarity). The
Sentence-BERT model was ne-tuned using the regression objective function on
the training set [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ]. Therefore, for each epoch, loss was computed by considering
the correlation between the gold similarity judgments and the predicted cosine
similarity between sentence embeddings.
          </p>
          <p>Our goal was to train the model to identify the gold tweet-vclaim pair among
the set of potential ones (i.e., those ltered with the IE step). To this aim,
we tried to separate the gold tweet-vclaim pairs from the other candidates. In
particular, given the assumption that a claim that veri es a tweet is semantically
similar to it, we built our training set as follows:
1. we created two positive examples from a gold tweet-vclaim pair, the rst
one composed by the tweet and the vclaim itself (tweet-vclaim pair), and
the second one composed by the tweet and the title of the claim
(tweetvclaim title). Both the positive pairs were assigned with a cosine similarity
value of 1:0. This forces the model to boost the similarity between the texts
belonging to gold pairs;
2. for each gold tweet-vclaim pair, 20 other tweet-vclaim pairs were randomly
selected as negative examples from the list of candidate pairs obtained with
the IE overlapping function (cf. Section 3.1). The similarity of the negative
examples was computed as the cosine similarity between vectors predicted by
bert-base-nli-mean-tokens, modi ed by the tanh function. This has the
e ect of decreasing the cosine similarity, thus e ectively penalising negative
examples.</p>
          <p>The bert-base-nli-factcheck-cos model was trained for 4 epochs with a
batch size of 8, and 10% of data was used for the model warm up.
Classi cation The bert-base-nli-factcheck-cos model is used to initialize
the weights for the classi er. In this case, the model is trained on a simple binary
classi cation task to distinguish between matching (gold) pairs, labelled as 1, and
non-matching ones, labelled as 0. Similarly to the previous ne-tuning step, we
selected negative examples among candidate tweet-vclaim pairs returned by the
IE module. Like for the sentence pair similarity model, for each tweet, the
tweetvclaim and the tweet-vclaim title pairs were used as positive examples. However,
in this case 2 negative examples were selected among the tweet-vclaim candidate
pairs, in order to better balance the training data for the classi cation.</p>
          <p>
            Our model bert-base-nli-factcheck-clas is therefore a Transformer with
a classi cation head on top of it, implemented with the Huggingface library.7
The model was trained for 3 epochs with a batch size of 8. We used the AdamW
optimizer with a learning rate of 2e 5 [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Inference step</title>
        <p>The inference step was performed by classifying the candidate tweet-vclaim pairs.
To retrieve the potential candidates, in fact, we applied the functions IE and
IEElastic described in section 3.1 to obtain, respectively, the
T2-EN-UNIPI-NLEBERT2IE and the T2-EN-UNIPI-NLE-BERT2IEElastic predictions. Moreover,
we also tested a run in which we classi ed all the target tweet-vclaim pairs
with no-preselection. The results of this additional experiment are shown in
Table 4. In all cases, we used the probability of the class 1 predicted by the
bert-base-nli-factcheck-clas model to rank the vclaims for each tweet.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The evaluation metric used in the competition is the Mean Average Precision
@5 (MAP@5) calculated over the gold ranking. The overall performance of our
models is reported in Table 3. For the sake of comparison, we also show the
performances obtained by the top models and by the o cial baseline as well.</p>
      <p>Our systems, namely the T2-EN-UNIPI-NLE-BERT2IE and the
T2-ENUNIPI-NLE-BERT2IEElastic, which di er for the overlapping function used at
inference time, obtained respectively 0:9160 and 0:9120 (cf. tables 4 and 5).</p>
      <p>Moreover, in order to explain the e ectiveness of each module for the nal
predictions, we computed their performances on the test set. Table 6 shows the</p>
      <sec id="sec-4-1">
        <title>7 https://huggingface.co</title>
        <p>Buster.ai
Buster.ai
UNIPI-NLE
UNIPI-NLE
type
primary
contr.-2
primary
contr.-1
Task Organizers baseline
0.907
0.877
performances of each module obtained on the task by ranking the claim for
a tweet according to several measures. More speci cally, for each module, we
report the model name, the type of the ne-tuning we applied, the function used
at inference time for selecting candidates and the MAP@5 obtained with the
o cial scorer.</p>
        <p>As for the IE step, given a tweet, we ranked the claims according to the
overlapping function for both the IE and the IEElastic methods. The IEElastic
method coincides actually with the baseline provided by the task organizers.
To assess the performances of the Sentence-BERT model ne-tuned on cosine
similarity, namely the bert-base-nli-factcheck-cos, we ranked the claim
according to the adjusted cosine similarity. Finally, we show the nal submitted
results. At inference time, the model bert-base-nli-factcheck-clas was fed
with the candidate tweet-vclaim pairs calculated with both the IE and the
IEElastic methods. In addition, we also report the results obtained by making the
predictions for all the target tweet-vclaim pairs.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>Several remarks can be made to comment our results. By looking at the model
scores, we see that our approach is able to outperform the baseline by a wide
margin, despite the fact that the Elasticsearch based approach proposed by the
task organizer was shown to be very e ective nonetheless. In addition, our system
ranked second among the participants of the task, obtaining performances that
are only slightly worse than the winning system, which obtained a MAP@5 of
0:938.</p>
      <p>Moreover, we can draw some interesting insights by considering the various
steps and data selection strategies. We notice that the IE baseline appears to
be less e ective as a standalone tool for selecting the best candidates among
claims for each tweet, with results well below the IEElastic one. However, the IE
method performs optimally when used as a selection criterion at inference time.
In fact, we experimented three methods for selecting candidate pairs at inference
time. In addition to the IE method and the IEElastic one, we assessed the
MAP
MAP
MAP
MAP
MAP
MAP
Precision
Precision
Precision
Precision
Precision
Precision
Rec Rank
Rec Rank
Rec Rank
Rec Rank
Rec Rank
Rec Rank
1
3
5
10
20
all
10
20
all
1
3
5
1
3
5
10
20
all
score
metric
performance with no pre-selection as well. In this case, the classi er was shown
with all possible tweet-vclaim pairs. We see that the IE method performs best,
but only slightly better than IEElastic. However, both pre-selection methods
outperform the model for which no pre-selection is made. We can argue that
this is because, during training, our objective was to enable the classi er to
distinguish between the gold tweet-vclaim pair and other pairs that share similar
features but are in fact incorrect. Therefore, the classi er may be more prone to
errors when tweet-vclaim pairs, which di er greatly from each other, are shown
as it never saw such examples during training. Experimental results seem to
con rm such hypothesis. This could be seen as a potential shortcoming for the
classi er itself. However, we can argue that considering the system as a whole,
it can bring two potential advantages. First, the classi er needs less negative
examples for an e ective training. We can argue that it is more di cult to
decide between two similar claims for a tweet, rather than between two very
di erent ones. Therefore, we chose to train the classi er to solve the \harder"
IEElastic baseline
IE baseline
classi cation
classi cation
classi cation
bert-base-nli-factcheck-cos</p>
      <p>cosine similarity
bert-base-nli-factcheck-cos
cosine similarity</p>
      <p>IEElastic
Fine-tuining</p>
      <p>Inference
IEElastic
IEElastic</p>
      <p>IE
IE
IE
problem, and addressed the \simpler" one with a less sophisticated, yet e ective,
approach. Second, the classi cation of each tweet-vclaim pair is time consuming.
On our machine, equipped with a Nvidia TitanXp graphic card, the inference
step considering all pairs took around ten hours. When performing the
preselection of pairs with our IE method, the inference step took less than four
hours. The IE method is e cient because it only needs to extract content words
and named entities for each pair, a task that is almost trivial in terms of time
complexity with modern NLP toolkits and current hardware.</p>
      <p>
        Finally, it is interesting to point out the contribution of the cosine similarity
adaptation performed with Sentence-BERT. Clearly, the model itself does not
perform well on the present task. However, two observations can be made. First,
during development we noticed that, by using a standard BERT model such as
BERT-base-uncased for representing sentences (i.e., by averaging word-level
representations obtained from the model), ranking claims based on cosine similarity
was completely ine ective, obtaining a very low MAP@5. Instead, by exploiting a
pre-trained Sentence-BERT model, we obtained much more encouraging results,
that were subsequently improved thanks to our cascade ne-tuning strategy. This
serves as additional evidence for the fact that standard BERT models are not
able to represent sentences in a semantically proper way, as already claimed in
the literature [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Second, by exploiting the ne-tuned Sentence-BERT model
(bert-base-nli-factcheck-cos) for obtaining the initial weights for the
classier (bert-base-nli-factcheck-clas), we clearly outperformed a model based
on BERT-base-uncased and trained in the same way. More speci cally, on the
development set we obtained a 0.72 MAP@5 for a bert-base-uncased ne-tuned
model and a MAP@5 of 0.78 for the bert-base-nli-factcheck-cos.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and future directions</title>
      <p>The approach to Fact Checking performed by the UNIPI-NLE team is based on a
combination of IE and DL strategies. The choice has been led by the assumptions
that supporting claims tend to mention the same entities and keywords of the
target tweet, and are semantically similar to it. On the one hand, a standard IE
module extracts relevant words and entities from texts and is used to construct
the training set for the following DL modules and to constrain the inference
process. In fact, the IE step is also crucial to turn down the processing time by
ltering the candidate pairs to be classi ed. Transformers, on the other hand,
are very useful to carry out e ective transfer learning, by ne-tuning large
pretrained models for speci c tasks such as the fact-checking one. In this paper,
ne-tuning is exploited both for modeling textual similarity and for classifying
text pairs to decide if a member of the pair veri es the other.</p>
      <p>The UNIPI-NLE approach strongly outperforms the baseline and ranked
second among the primary submissions of the task. In the future, we plan to
perform some additional hyperparameter tuning on the models. Moreover, we
would like to test this approach in similar tasks such as Fake News identi cation.
We are con dent that by exploiting the dynamic selection of training data in
addition to an e ective and e cient information extraction strategy, we will
obtain strong performances also to solve this harder task.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>We gratefully acknowledge the support of NVIDIA Corporation with the
donation of the Titan Xp GPU used for this research. This work was partially
supported by the University of Pisa in the context of the project \Event
Extraction for Fake News Detection" in the framework of the MIT-Unipi program, and
by the Italian Ministry of Education and Research (MIUR) in the framework of
the CrossLab project (Departments of Excellence).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arampatzis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.):
          <article-title>Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <source>Interaction Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ).
          <source>LNCS (12260)</source>
          , Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Barron-Ceden~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Elsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , Da San Martino, G.,
            <surname>Hasanain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Haouari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Babulkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Hamdan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Nikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Shaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Sheikh Ali</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          : Overview of CheckThat! 2020:
          <article-title>Automatic identi cation and veri cation of claims in social media</article-title>
          .
          <source>In: Arampatzis et al. [1]</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ducharme</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vincent</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jauvin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A neural probabilistic language model</article-title>
          .
          <source>Journal of machine learning research 3(Feb)</source>
          ,
          <volume>1137</volume>
          {
          <fpage>1155</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>5</volume>
          ,
          <issue>135</issue>
          {
          <fpage>146</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bondielli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcelloni</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A survey on fake news and rumour detection techniques</article-title>
          .
          <source>Information Sciences</source>
          <volume>497</volume>
          ,
          <volume>38</volume>
          {
          <fpage>55</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angeli</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potts</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.:</given-names>
          </string-name>
          <article-title>A large annotated corpus for learning natural language inference</article-title>
          .
          <source>In: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>632</volume>
          {
          <fpage>642</fpage>
          . Association for Computational Linguistics, Lisbon, Portugal (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Canini</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suh</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pirolli</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          :
          <article-title>Finding credible information sources in social networks based on content and social structure</article-title>
          .
          <source>In: 2011 IEEE Third International Conference on Privacy, Security, Risk and Trust and 2011 IEEE Third International Conference on Social Computing</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cappellato</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . (eds.): Working Notes of CLEF 2020|
          <article-title>Conference and Labs of the Evaluation Forum (</article-title>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poblete</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Information credibility on twitter</article-title>
          .
          <source>In: Proceedings of the 20th international conference on World wide web</source>
          . pp.
          <volume>675</volume>
          {
          <issue>684</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Cer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agirre</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez-Gazpio</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specia</surname>
          </string-name>
          , L.:
          <article-title>SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation</article-title>
          .
          <source>In: Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval2017)</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>14</fpage>
          . Association for Computational Linguistics, Vancouver, Canada (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ciampaglia</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shiralkar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bollen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Computational fact checking from knowledge networks</article-title>
          .
          <source>PloS one 10(6)</source>
          ,
          <year>e0128193</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          : BERT:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers). pp.
          <volume>4171</volume>
          {
          <fpage>4186</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Gorrell</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochkina</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liakata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zubiaga</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derczynski</surname>
          </string-name>
          , L.:
          <article-title>SemEval-2019 task 7: RumourEval, determining rumour veracity and support for rumours</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>845</volume>
          {
          <fpage>854</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Loshchilov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hutter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Decoupled weight decay regularization</article-title>
          .
          <source>In: In Proceedings of the 2019 International Conference on Learning Representations</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>CoRR abs/1301</source>
          .3781 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Proceedings of the 26th International Conference on Neural Information Processing Systems - Volume</source>
          <volume>2</volume>
          . p.
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          . NIPS'
          <volume>13</volume>
          , Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Leveraging joint interactions for credibility analysis in news communities</article-title>
          .
          <source>In: Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          . pp.
          <volume>353</volume>
          {
          <issue>362</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Nie</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bansal</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Combining fact extraction and veri cation with neural semantic matching networks</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on Arti cial Intelligence</source>
          . vol.
          <volume>33</volume>
          , pp.
          <volume>6859</volume>
          {
          <issue>6866</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Popat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Strotgen, J.,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Where the truth lies: Explaining the credibility of emerging claims on the web and social media</article-title>
          .
          <source>In: Proceedings of the 26th International Conference on World Wide Web Companion</source>
          . pp.
          <volume>1003</volume>
          {
          <issue>1012</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Popat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Strotgen, J.,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Where the truth lies: Explaining the credibility of emerging claims on the web and social media</article-title>
          .
          <source>In: Proceedings of the 26th International Conference on World Wide Web Companion</source>
          . pp.
          <volume>1003</volume>
          {
          <issue>1012</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Stanza: A Python natural language processing toolkit for many human languages</article-title>
          .
          <source>In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Improving language understanding by generative pre-training (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amodei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Language models are unsupervised multitask learners</article-title>
          .
          <source>OpenAI Blog</source>
          <volume>1</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Reimers</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Sentence-BERT: Sentence embeddings using siamese BERT-networks</article-title>
          .
          <source>In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>3982</volume>
          {
          <fpage>3992</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Shaar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babulkov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , Da San Martino, G.,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>That is a known lie: Detecting previously fact-checked claims</article-title>
          .
          <source>In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          . pp.
          <volume>3607</volume>
          {
          <fpage>3618</fpage>
          . Association for Computational Linguistics,
          <string-name>
            <surname>Online</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Shaar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Babulkov</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alam</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barron-Ceden</surname>
            ~o,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elsayed</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasanain</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suwaileh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haouari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , Da San Martino, G.,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Overview of CheckThat! 2020 English:
          <article-title>Automatic identi cation and veri cation of claims in social media</article-title>
          .
          <source>In: Cappellato et al. [8]</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Shiralkar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciampaglia</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          :
          <article-title>Finding streams in knowledge graphs to support fact checking</article-title>
          .
          <source>In: Proceedings of the 2017 IEEE International Conference on Data Mining (ICDM)</source>
          . pp.
          <volume>859</volume>
          {
          <fpage>864</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sliva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu, H.:
          <article-title>Fake news detection on social media: A data mining perspective</article-title>
          .
          <source>ACM SIGKDD explorations newsletter 19(1)</source>
          ,
          <volume>22</volume>
          {
          <fpage>36</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Tchechmedjiev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fafalios</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boland</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasquet</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zloch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zapilko</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dietze</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Todorov</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Claimskg: a knowledge graph of fact-checked claims</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          . pp.
          <volume>309</volume>
          {
          <fpage>324</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Thorne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vlachos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christodoulopoulos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mittal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>FEVER: a large-scale dataset for fact extraction and VERi cation</article-title>
          .
          <source>In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long Papers). pp.
          <volume>809</volume>
          {
          <fpage>819</fpage>
          . Association for Computational Linguistics, New Orleans,
          <string-name>
            <surname>Louisiana</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>5998</volume>
          {
          <issue>6008</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Vlachos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riedel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Fact checking: Task de nition and dataset construction</article-title>
          .
          <source>In: Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science</source>
          . pp.
          <volume>18</volume>
          {
          <fpage>22</fpage>
          . Association for Computational Linguistics, Baltimore,
          <string-name>
            <surname>MD</surname>
          </string-name>
          , USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , W.Y.:
          <article-title>"liar, liar pants on re": A new benchmark dataset for fake news detection</article-title>
          . In: Barzilay,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Kan</surname>
          </string-name>
          , M.Y. (eds.)
          <source>ACL (2)</source>
          . pp.
          <volume>422</volume>
          {
          <fpage>426</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nangia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A broad-coverage challenge corpus for sentence understanding through inference</article-title>
          .
          <source>In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long Papers). pp.
          <volume>1112</volume>
          {
          <fpage>1122</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Carbonell, J.,
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Xlnet: Generalized autoregressive pretraining for language understanding</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>5753</volume>
          {
          <issue>5763</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>