<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>bigIR at CLEF 2019: Automatic Veri cation of Arabic Claims over the Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fatima Haouari</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zien Sheikh Ali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tamer Elsayed</string-name>
          <email>telsayedg@qu.edu.qa</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Qatar University</institution>
          ,
          <addr-line>Doha</addr-line>
          ,
          <country country="QA">Qatar</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the proliferation of fake news and its prevalent impact on democracy, journalism, and public opinions, manual fact-checkers become unscalable to the volume and speed of fake news propagation. Automatic fact-checkers are therefore needed to prevent the negative impact of fake news in a fast and e ective way. In this paper, we present our participation in Task 2 of CLEF-2019 CheckThat! Lab, which addresses the problem of nding evidence over the Web for verifying Arabic claims. We participated in all of the four subtasks and adopted a machine learning approach in each with di erent set of features that are extracted from both the claim and the corresponding retrieved Web search result pages. Our models, trained solely over the provided training data, for the di erent subtasks exhibited relatively-good performance. Our o cial results, on the testing data, show that our best performing runs achieved the best overall performance in subtasks A and B among 7 and 8 participating runs respectively. As for subtasks C and D, our best performing runs achieved the median overall performance among 6 and 9 participating runs respectively.</p>
      </abstract>
      <kwd-group>
        <kwd>Fact Checking Arabic Retrieval Learning to Rank Classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Fake news is witnessing an explosion recently, and it is considered as one of
the biggest threats to democracy, journalism, and public trust in governments.
In combating fake news, the number of manual fact-checking organizations
increased by 239% in a period of four years, where it reached 149 fact-checkers in
2018 as apposed to only 44 in 20141.</p>
      <p>
        One of the main challenges is that manual fact-checking does not scale with
the volume of daily fake news. This mismatch can be attributed to the gap
between the time the claim is made and the time the claim is checked and
published, as it is very time-consuming for journalists to nd check-worthy claims
and verify them. Another challenge is that fact-checking requires advanced
writing skills in order to convince the readers whether the claim is true or false [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In
fact, it is estimated that check-worthiness of a claim and writing an article about
it can take up to one day [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Moreover, manual fact-checkers are outdated [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Most of the fact-checking frameworks adopt the old content management systems
specialized for traditional blogs and newspapers, but not built for the current
modern journalism. A new approach is therefore needed for automated fake news
detection and veri cation.
      </p>
      <p>
        The industry and academia have shown an overwhelming interest in fake
news to address the challenges of its detection and veri cation. Many pioneering
ideas were proposed to address many aspects of fact-checking systems with their
focus varies between detecting check-worthy claims [
        <xref ref-type="bibr" rid="ref10 ref7 ref8">8, 10, 7</xref>
        ], checking claims
factuality [
        <xref ref-type="bibr" rid="ref11">11, 15, 16, 20</xref>
        ], checking news media factuality [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and proposing full
automatic fact-checking systems [
        <xref ref-type="bibr" rid="ref14 ref9">9, 14, 22</xref>
        ]. There are also some shared tasks
proposed and open to the research community interested in the problem such as
FEVER-2018 task for fact extraction and veri cation [18] and CheckThat! 2018
lab on automatic identi cation and veri cation in political debates at CLEF [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        This year, CLEF-2019 CheckThat! Lab [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] introduced two tasks to tackle
two main problems of automated fact-checking systems. The main objective of
the rst is to detect check-worthy-claims to be prioritized for fact-checking [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
while the second focuses on evidence extraction to support fact-checking a claim
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In this paper, we present the approach adopted by our bigIR group at Qatar
University to address the second task.
      </p>
      <p>Task 2 (Evidence and Factuality) addresses the problem of nding evidence
over the Web for verifying Arabic claims. It assumes the system is given an
Arabic claim (as a short sentence) and a corresponding ranked list of Web pages
that were retrieved by a Web search engine for that claim. The system then
needs to address four sub-problems, each is de ned as a subtask as follows:
1. Subtask A: Rank the retrieved pages based on how useful they are for
verifying the claim.
2. Subtask B: Classify the Web pages as \very useful" for veri cation,
\useful", \not useful", or \not relevant".
3. Subtask C: Within each useful page, identify which passages are useful for
claim veri cation.
4. Subtask D: Determine the true factuality of the claim, i.e., whether it is
"True" or "False".</p>
      <p>We have participated in all of the four subtasks. Since it is the rst year of
the task (and thus our rst attempt), we generally adopted a simple machine
learning approach, where learning models were trained only on the given training
data over hand-crafted features. We applied feature ablation to assess the impact
of each feature on the performance of our models.</p>
      <p>For subtask A, to re-rank the pages based on their usefulness, we adopted
a pairwise learning-to-rank approach with features extracted either from the
page as a whole (such as source popularity, URL links, and number of quotes),
from the relevant segments in the page (such as the similarity score of the most
relevant sentence), or from the search results (such as the original rank of the
page). Additionally, we extracted claim-dependent features such as the similarity
between the claim and the title and the snippet of the page.</p>
      <p>For subtask B, we adopted a multi-class classi cation approach to classify the
Web pages. We considered several features including word embeddings, named
entities, similarity scores, number of relevant sentences in the page, and
URLbased features (such as URL length, URL scheme, and URL domain).</p>
      <p>For subtask C, we adopted a binary classi cation approach to classify the
passages within a useful page. Features included Bag-Of-Words (BOW), named
entities, number of quotes, score of most relevant sentence from each passage,
and the similarity score between the claim and the passage.</p>
      <p>For subtask D, we also adopted a binary classi cation approach to discover
the claim's factuality given the retrieved Web pages. To classify the claim, we
rst identify the most similar pages to the claim for feature extraction. For the
selected pages, we consider their similarity scores, source popularity, and the
sentiment of the page.</p>
      <p>Our contribution in this work is two-fold:
1. We participated in all of the four subtasks adopting a machine learning
approach with relatively-di erent set of features in each. The features are
extracted from both the claims and the retrieved Web pages.
2. Our best performing runs exhibited the best performance in both subtasks</p>
      <p>A and B among the submitted runs.</p>
      <p>The remainder of this paper is organized as follows. Section 2 describes how
we processed and extracted features from the claims and retrieved pages.
Sections 3, 4, 5, and 6 outline our approach and discuss our experimental evaluation
in detail for subtask A, B, C, and D respectively. Finally, Section 7 concludes
and discusses possible future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Preprocessing &amp; Feature Extraction</title>
      <p>In our work, we apply common main preprocessing for all subtasks to parse
documents, identify relevant segments, and extract features. However, we include
or exclude some features in each subtask. In this section, we describe in detail
the preprocessing steps and introduce and motivate the features we extracted at
all levels. For each page, we extract two types of features: features that depend
on the claim/page relationship (claim-dependent) and features that depend
solely on the page (page-dependent).</p>
      <p>In what follows, a text segment in a page is centered by one sentence, but
also includes both the sentence that precedes and the sentence that follows it,
as de ned by Yasser et al. [21], to consider the context of the sentence.
2.1</p>
      <sec id="sec-2-1">
        <title>HTML Parsing</title>
        <p>As the Web pages are in raw HTML format, we parse each page by extracting
only the clean version of the textual body discarding images, videos, and scripts
using newspaper2 and BeautifulSoup3 Python libraries. We removed stopwords
using Python NLTK4 Arabic stopwords. We also discard the sentences containing
less than 3 words, motivated by the empirical study done by Zhi et al. [22].
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Text Vector Representations</title>
        <p>
          In extracting our features, we consider two text vector representations:
{ Bag-of-Words (BOW): We consider BOW representation to represent full
passages (mainly for subtask C). We considered only the terms that appeared
at least 7 times in the training data, based on some preliminary experiments.
{ Distributed Representation (W2V): We consider word2vec embeddings [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
to represent the claim and the segments of a page; each is represented as the
average vector of the embeddings of terms in the claim/segment. We used
the pre-trained AraVec embeddings model proposed by Soliman et al. [17].
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Relevant Segments Identi cation</title>
        <p>To identify relevant segments in a page for a given claim, we represent the claim
and each sentence in the page by their average of term W2V vectors. We then
compute the cosine similarity between the vectors of the claim and each segment.
Segments are considered relevant if the similarity score is higher than a threshold.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Page-Dependent Features</title>
        <p>
          We extracted two types of page-dependent features: credibility and content.
Credibility Features To indicate the credibility of the page, we consider the
following features:
{ Source Popularity (SrcPop): This feature may indicate trustworthiness,
as it captures how popular a particular website is. We used Amazon Alexa
rank5 motivated by Baly et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] that used this feature to estimate the
reliability of media sources. We consider this feature as a categorical feature
by binning the ranking values into 10 categories, then we convert it to a one
hot encoding vector of 10 binary features.
2 https://pypi.org/project/newspaper3k/
3 https://pypi.org/project/bs4/
4 https://pypi.org/project/nltk/
5 https://www.alexa.com/
{ URL Features: these features were used by Baly et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to detect the
reliability of web sources. We used Python URL handling library urlib6 to
parse the URL and extract the following orthographic features:
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Length (URLLen) and Number of Sections (URLSecs): The</title>
        <p>length of the URL path and the number of sections separated by `/'
help indicate whether the website is legitimate, irregular, or a phishing
website.</p>
        <p>Scheme (URLScheme): The URL protocol (https or http) indicates
the trustworthiness of the website. We extracted the URL scheme then
we used scikit-learn label encoder7 to encode string values of schemes to
integers.</p>
        <p>Domain Su x(URLSfx): The su x of a URL domain determines
the source and credibility of the website. For example, a website with
domain su x .gov is a federal government site and is more credible than
a commercial website with a su x of .com. We used label encoder to
encode their string values into integers.</p>
        <p>Content Features From page body, we extract the following linguistic and
similarity features:
{ Number of Quotes (NQts): For each page, we count the number of quotes
in all relevant segments. This feature may be very useful to rank web pages
and decide how useful they are for claim veri cation as it may indicate the
credibility of the page by quoting sources. In our work, we considered only
quotes with ve words or more.
{ Number of URL links (NLinks): This feature represents the number of
URL links in the retrieved page. It may indicate the credibility of the source
by giving references.
{ Named Entities (NEs): Pages mentioning named entities may indicate the
truthfulness of the page. We used Python polyglot NLP tool8 to recognize
location, organizations, and persons entities in the most relevant segment of
the page. We form a vector of 3 integer values representing the number of
occurrences of every entity type in the segment.
2.5</p>
      </sec>
      <sec id="sec-2-6">
        <title>Claim-Dependent Features</title>
        <p>We extracted the following features based on the claim-page interaction:
{ Original Rank (Rank): This feature is available from the search results
and it represents how the page is potentially-relevant to the claim according
to the search engine.
{ Similarity: This includes cosine similarity between claim and title (ClmTtlSim),
claim and snippet(ClmSnptSim), and claim and a passage (ClmPsgSim).
6 https://pypi.org/project/urllib3/
7 scikit-learn.org/stable/modules/generated/sklearn.preprocessing.LabelEncoder.html
8 https://github.com/aboSamoor/polyglot
{ Number of Relevant Sentences (NRelSent): For every page, we
compute the similarity between the claim and each sentence. We count the
number of relevant sentences in each page as it might indicate the relevance of
the page.
{ Number of relevant webpages (NRelPages): For every claim, we count
the number of webpages with a similarity score between claim and most
relevant sentence higher than a certain threshold.
{ Score of the most Relevant Segment (MostRelSeg): This feature
indicates how similar the most relevant segment is to the claim.
{ Sentiment (SntCnt): Sentiment analysis can help identify if the stance
of the page is positive, negative, or neutral. This may help in identifying
whether the page agrees with the claim or not. We use polyGlots Sentiment
model9 to extract sentiments. From the most relevant segment, we get two
values, the number of words with positive polarity and the number of words
with negative polarity.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Subtask A: Reranking Retrieved Pages</title>
      <p>
        In this subtask [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the goal is to rerank the retrieved pages based on their
usefulness for verifying a speci c claim. In this section, we present our proposed
approach, experimental setup and results, our selected runs for CLEF
submissions, and nally we will present the CLEF results.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Approach</title>
        <p>Our approach is based on learning-to-rank (L2R). We propose a pairwise L2R
model considering three di erent L2R classi ers, namely, SVM C-Support Vector
Classi cation (SVC), which is implemented based on libsvm10, Gaussian Nave
Bayes (Gaussian NB), and the ensemble classi er Random Forest (RF), using
Scikit-learn Python library.11 We consider the following features (discussed in
Section 2):
{ Basic features: Rank, SrcPop, and MostRelSeg.
{ Similarity features: ClmTtlSim and ClmSnptSim.
{ NLinks.</p>
        <p>{ NQts.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experimental Setup</title>
        <p>Parameters We experimented with the three di erent classi ers mentioned
in 3.1. We set the kernel for SVC to linear, and set the number of estimators for
the RF models to 100 (based on preliminary experiments). For the NB models,
we did not tune any hyper-parameters and used the default settings.
9 https://polyglot.readthedocs.io/en/latest/Sentiment.html
10 https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVC.html
11 https://scikit-learn.org/stable/index.html
Baselines We compare our models against a baseline that returns the pages
ranked in their original ranks (i.e., based on relevance scores of the search engine,
not on usefulness for fact-checking).
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Evaluation on Training</title>
        <p>As we were constrained by the size of the training data, containing only 10
claims, we adopted leave-one-claim-out (LOO) cross validation to evaluate the
trained models. We optimized our models using the graded relevance measure
NDCG@20.</p>
        <p>We rst experimented with di erent values of the cosine similarity threshold
(0.4, 0.5, 0.6, and 0.7) when extracting relevant segments. In our unreported
preliminary experiments, we observed that the best performing models were the
ones trained with features extracted using a similarity threshold of 0.4 and 0.7,
presented in Fig. 1 and Fig. 2 respectively. We also tried di erent combinations
of features as shown in both gures.</p>
        <p>The results show that our models could not beat the baseline with only
the basic features. However, NB models outperformed the baseline when other
features were introduced. We also notice that introducing the ClmTtlSim and
ClmSnptSim to the basic features improved the performance of our models, while
excluding the SrcPop feature improved the performance. Moreover, our proposed
NLinks and NQts features did not have a noticeable impact on the performance
of the models.
Runs As shown in Fig. 1, NB models outperform other L2R models over the
training data, therefore we picked the 3 best NB models to submit to CLEF:.
1. NB trained with Basic, ClmTtlSim, and ClmSnptSim features, and excluding</p>
        <p>SrcPop.
2. NB trained with Basic, ClmTtlSim, ClmSnptSim, and NQts features, and
excluding SrcPop.
3. NB trained with Basic, ClmTtlSim, ClmSnptSim, NQts, and NLinks
features, and excluding SrcPop.</p>
        <p>Moreover, when the cosine similarity threshold was set to 0.7, RF outperformed
other models, as shown in Fig. 2, so we also picked its best performing model:
4. RF trained with Basic, ClmTtlSim, and ClmSnptSim features, and excluding</p>
        <p>SrcPop.</p>
        <p>Results As shown in Table 1, the o cial CLEF evaluation shows that our best
performing model on the test data was the NB model trained with basic,
ClmTtlSim, and ClmSnptSim features (excluding SrcPop) which achieved NDCG@20
value of 0.55. This was the maximum score achieved among 7 runs submitted
for this subtask. We observed that the performance of our models on training
data was better than on testing data; this can be attributed to the small size of
the training dataset, containing only 395 pages from 10 claims, which could be
insu cient and not a good representative to train the models.
fBasic+Simg</p>
        <p>-SrcPop
fBasic+Sim+NQtsg</p>
        <p>-SrcPop
fBasic+Simg</p>
        <p>-SrcPop
fBasic+Sim
+NQts+NLinksg
-SrcPop</p>
        <p>RF
NB
NB
NB</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Subtask B: Classifying Retrieved Pages</title>
      <p>
        The main goal of this subtask [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is to classify all retrieved Web pages based
on how useful they are in detecting the claim's veracity. A webpage is useful if
it has enough evidence to verify the claim and if its source is trustworthy. In
this section, we present our approach, experimental setup, training results, and
CLEF results for our submitted runs.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Approach</title>
        <p>In our approach for this subtask, we use di erent machine learning algorithms
to perform multi-class classi cation. We consider SVC as it shows to learn well
from small datasets. We also include Gradient Boosting (GB) and RF as an
ensemble model. As mentioned in 3.1, we use Scikit-learn Python library for our
implementation. We consider the following features:
{ Basic features: Rank, SrcPop and MostRelSeg.
{ NEs in the relevant segment.
{ NQts.
{ URL features.</p>
        <p>{ W2V representation of both the claim and the relevant segment.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Experimental Setup</title>
        <p>Parameters For SVC, we used an RBF kernel with regularization parameter
C = 15 and L2 penalty, and we set to 0.01 to avoid over- tting. For GB and
RF models, we set the number of estimators to 100 and 150 respectively (based
on preliminary experiments).</p>
        <p>Baselines As a baseline we adopted Wang et al. [19] method for feature
extraction and classi cation. Their dataset consists of short passages where passages
are classi ed into ve di erent categories. This baseline was selected because
the feature extraction methods are implemented on short passages similar to the
size of our extracted relevant segments. Moreover, they are working on ne-grain
classi cation.</p>
        <p>Since our training data is highly imbalanced, we also used the Zero Rule
algorithm as a baseline for this subtask. Zero Rule algorithm predicts the majority
class in the dataset. In our training data, class -1 (non-relevant) is the majority
class with 65% of the labels.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation on Training</title>
        <p>We conducted multiple experiments in attempt to nd which features
combination will result in the best F1 score. We split our dataset into 70% for training
and 30% for testing. From our experiments, we noticed that varying the
similarity threshold when extracting relevant segments had a signi cant impact on
the overall score. We concluded that our best performing models were the ones
trained with features extracted with similarity thresholds of 0.4 and 0.7. Fig.
3 and Fig. 4 show the results obtained from our experiments using similarity
thresholds 0.4 and 0.7 respectively.</p>
        <p>We observed that when training the classi ers with basic features and NEs
the performance improved. On the other hand, incorporating some content
features like URL features and W2V vectors had a negative impact on the
performance of the classi ers. We also note that ensemble classi ers (GB and RF)
outperformed the baselines and other classi ers all the time.
Runs As concluded in section 4.3, ensemble classi ers have outperformed SVC
classi ers. So, for our runs we picked the GB and RF models. We selected the
following models with cosine similarity threshold of 0.7:</p>
        <sec id="sec-4-3-1">
          <title>1. GB Classi er trained with basic features. 2. GB Classi er trained with basic features and NEs. We also picked the following models when cosine similarity threshold is set to 0.4:</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>3. GB Classi er trained with basic features and NQts.</title>
          <p>4. RF Classi er trained with basic features and NQts.</p>
          <p>Results Table 2 shows our training results compared to the o cial CLEF testing
results. We notice that our best validation model with F1 score of 0.52 that
combines basic features with NEs has achieved lower testing score. Meanwhile,
our model that combines basic features with NQts has scored a testing F1 score
of 0.31. The inconsistency between train and test F1 scores can be justi ed due
to the small training dataset of only 395 webpages. Also, the imbalance in the
classes of the dataset could have caused the models to over t. Our best model
that achieved F1 score value of 0.31 is the highest among all submitted runs for
this subtask.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Subtask C: Classifying Passages</title>
      <p>
        In this subtask [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the goal is to extract useful passages for claim veri
cation within the useful retrieved pages. In this section, we present our proposed
methodology, experimental evaluation, selected runs for this subtask, and CLEF
results.
Deciding whether a passage within a useful page is useful or not is a classi
cation problem. Therefore, our methodology is based on using di erent machine
learning classi ers namely SVC, NB, and RF. We consider the following features
for this subtask:
{ BOW of the passage.
{ MostRelSeg in the passage.
{ ClmPsgSim.
{ NQts in the passage.
      </p>
      <p>{ NEs in the passage.
5.2</p>
      <sec id="sec-5-1">
        <title>Experimental Setup</title>
        <p>Parameters The three di erent classi ers mentioned in section. 5.1 were used
in our experiments. We set the kernel for SVC to linear, and the number of
estimators for the RF models to 100 in all the experiments. For the Gaussian
NB, we did not tune any hyperparameters and we based our experiments on the
default settings.</p>
        <p>Baselines We compare our models against the majority baseline.
5.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>Evaluation on Training</title>
        <p>Since we have only 6 claims in the dataset provided for subtask C, which contains
only 167 passages from 31 di erent pages, we considered LOO cross validation
in our experiments. We used F1 score as our evaluation metric. As shown in
Fig. 5, SVC outperformed all other models with all groups of features. However,
when the BOW features were excluded, the Gaussian NB achieved the best
among all. We also observed that the two best performing models are the SVC
model when the NEs features were excluded, and the SVC model when the NQts
feature was excluded achieving an F1 score of 0.444 and 0.43 respectively. We
also noticed that the performance of the SVC model trained with all features
improved compared to when trained with BOW features only, achieving an F1
score of 0.427 as apposed to 0.387.
Runs As shown in Fig. 5, SVC models outperformed other classi ers except
when the BOW features were excluded, in which case the NB model achieved
the best F1 score. Therefore, we picked the 3 best SVC models and the best NB
model to submit:</p>
        <sec id="sec-5-2-1">
          <title>1. SVC trained with all features. 2. SVC trained with all features excluding the NQts feature. 3. SVC trained with all features excluding NEs features. 4. NB trained with all features excluding BOW features.</title>
          <p>Results As shown in Table 3, in the o cial CLEF evaluation, our best
performing model in the test phase was the SVC model trained with all features
excluding the NQts features, which achieved F1 score value of 0.4. The low F1 of
our models can be attributed to the big di erence in training and testing data
including passages from 6 claims and 59 claims respectively. Our highest scoring
model is ranked 3rd out of the six runs submitted to the lab, and the maximum
score achieved among all runs submitted for this subtask was 0.56.
6</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Subtask D: Verifying Claims</title>
      <p>The goal of this subtask is to identify whether the claim is "True" or "False". For
a claim to be true, it should have supporting evidence that veri es its factuality.
In this section, we present our approach, experimental setup, and training results
for verifying the claims. Then, we discuss CLEF results for our submitted runs.
Deciding the factuality of a claim is a binary classi cation problem. Therefore,
we propose a supervised learning approach using di erent classi ers: GB, RF
and Linear Discriminant Analysis (LDA).</p>
      <p>For this subtask, we select the most signi cant features from webpages to
classify the claim. Unlike previous tasks, we consider SntCnt features to nd the
polarity of the webpage. In addition, we consider the usefulness of the article by
using the most relevant segment extracted as explained in Section 2 to represent
the webpage. In our experiments, we consider the following features for our
binary classi ers:
{ Similarity Scores: out of all webpages associated with a claim, we only
consider three di erent scores: maximum ClmTtlSim, ClmSnptSim, and MostRelSeg.
{ NRelPages.
{ For every claim, we select the webpage with maximum MostRelSeg value
and extract the following features from it: SrcPop and SntCnt.
6.2</p>
      <sec id="sec-6-1">
        <title>Experimental Setup</title>
        <p>Parameters For GB and RF classi ers, we found that the default parameters
are the best (based on preliminary experiments). For LDA classi er, we found
that using 5 components for linear discrimination is most e ective in terms of
accuracy.</p>
        <p>
          Baseline As a baseline for this subtask, we implemented Karadzhov et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]
method. They classify claims as "True" or "False" based on the top returned
search results from several engines. They used an SVC classi er with RBF kernel
in their experiments. The inputs to the classi er are word embeddings of the most
relevant segment in the webpage, webpage snippet, and the claim. In addition
to the word embeddings, the average and maximum similarity scores of the
segments and snippets are included as features. We also adopt their method of
segment extraction to compare with our approach.
We conducted multiple experiments to nd which features combination will
result in the best factuality classi cation. Due to the limitation in the size of
training dataset, we used 8-fold cross validation on all our models for this
subtask.
        </p>
        <p>We rst experimented with di erent values of the cosine similarity threshold
(0.4, 0.5, 0.6, and 0.7) when extracting relevant segments. In our unreported
preliminary experiments, we observed that the best performing models were the
ones trained with features extracted using a similarity threshold of 0.6 presented
in Fig. 6. We noticed that the GB model trained with all features outperformed
all other models. We also observed that our models outperformed the baseline
score most of the time except when the NRelPages were excluded from the
features. Furthermore, we conclude that NRelPages and SntCnt features are
useful in classi cation of a claim.
Runs Based on our training results presented in section 6.3, we decided to use
the models trained on all features to classify the claims factuality on testing data.
We selected the best ensemble classi ers with two di erent similarity thresholds.</p>
        <sec id="sec-6-1-1">
          <title>1. GB classi er, with similarity threshold 0.7. 2. GB classi er, with similarity threshold 0.4. 3. RF classi er, with similarity threshold 0.4. 4. RF classi er, with similarity threshold 0.6.</title>
          <p>Results Table 4 shows our training results compared to the o cial CLEF testing
results. Runs for subtask D were submitted over two cycles. In the rst cycle, we
classify the claims factuality using all webpages provided. In the second cycle, we
classify the claims factuality using only useful webpages. We present the results
for the second cycle in this section.</p>
          <p>As presented in Table 4, we notice that all models achieved very similar F1
test scores. However, our GB model trained with all features has the highest
training and testing scores, achieving F1 score of 0.91 and 0.53 for training and
testing respectively. Our highest scoring model is ranked 4th out of the nine runs
submitted to the lab, and the maximum score achieved among all runs submitted
for this subtask was 0.62.
In this paper, we present our approach for task 2 of CLEF-2019 CheckThat!
Lab. For subtask A, we proposed pairwise learning-to-rank approach using
different learning models to rank the retrieved pages based on their usefulness.
Our best performing model trained using the basic and similarity features
(excluding source popularity) achieved an NDCG@20 of 0.55, which is the highest
score among 7 runs submitted for this subtask. For subtask B, we proposed a
classi cation model incorporating source popularity feature along with named
entities. Our best performing model achieved an F1 score of 0.31, which is the
highest score achieved among the 8 runs submitted for this subtask. For subtask
C, we proposed a classi cation model considering BOW, named entities, and the
number of quotes features extracted from passages. Our best performing model
trained with all features (excluding the number of quotes) achieved an F1 score
of 0.4 and got 3rd place. For subtask D, we proposed a classi cation model using
sentiment features to nd the polarity of the page, in addition to the number of
potentially-relevant pages. Our best model trained with all features achieved an
F1 score of 0.53 and got 4th place.</p>
          <p>That was our rst attempt using a very small training data that was provided
by the track organizers. With larger datasets, we plan to improve our classi
cation models with more features including word embeddings, that are trained
speci cally for this task, and probably with deep learning models as well.
15. Rashkin, H., Choi, E., Jang, J.Y., Volkova, S., Choi, Y.: Truth of varying shades:
Analyzing language in fake news and political fact-checking. In: Proceedings of
the 2017 Conference on Empirical Methods in Natural Language Processing. pp.
2931{2937 (2017)
16. Ruchansky, N., Seo, S., Liu, Y.: Csi: A hybrid deep model for fake news detection.</p>
          <p>In: Proceedings of the 2017 ACM on Conference on Information and Knowledge
Management. pp. 797{806. ACM (2017)
17. Soliman, A.B., Eissa, K., El-Beltagy, S.R.: Aravec: A set of arabic word embedding
models for use in arabic nlp. Procedia Computer Science 117, 256{265 (2017)
18. Thorne, J., Vlachos, A., Christodoulopoulos, C., Mittal, A.: Fever: a large-scale
dataset for fact extraction and veri cation. arXiv preprint arXiv:1803.05355 (2018)
19. Wang, L., Wang, Y., de Melo, G., Weikum, G.: Five shades of untruth:
Finergrained classi cation of fake news. In: 2018 IEEE/ACM International Conference
on Advances in Social Networks Analysis and Mining (ASONAM). pp. 593{594.</p>
          <p>IEEE (2018)
20. Wang, W.Y.: \Liar, liar pants on re": A new benchmark dataset for fake news
detection. arXiv preprint arXiv:1705.00648 (2017)
21. Yasser, K., Kutlu, M., Elsayed, T.: Re-ranking Web Search Results for Better
Fact-Checking: A Preliminary Study. In: Proceedings of 27th ACM International
Conference on Information and Knowledge Management (CIKM). pp. 1783{1786.</p>
          <p>ACM, Turin, Italy (2018)
22. Zhi, S., Sun, Y., Liu, J., Zhang, C., Han, J.: Claimverif: a real-time claim veri
cation system using the web and fact databases. In: Proceedings of the 2017 ACM
on Conference on Information and Knowledge Management. pp. 2555{2558. ACM
(2017)</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Atanasova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohtarami</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Da San Martino, G.:
          <article-title>Overview of the CLEF-</article-title>
          2019
          <source>CheckThat! Lab on Automatic Identi cation and Veri cation of Claims. Task</source>
          <volume>1</volume>
          : Check-Worthiness
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Baly</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexandrov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glass</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Predicting Factuality of Reporting and Bias of News Media Sources</article-title>
          . CoRR abs/
          <year>1810</year>
          .01765 (
          <year>2018</year>
          ), http://arxiv.org/abs/
          <year>1810</year>
          .01765
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Elsayed</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Barron-Ceden~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Hasanain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            , Da San Martino, G.,
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          : Checkthat! at clef 2019:
          <article-title>Automatic identi cation and veri cation of claims</article-title>
          .
          <source>In: European Conference on Information Retrieval</source>
          . pp.
          <volume>309</volume>
          {
          <fpage>315</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Elsayed</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Barron-Ceden~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Hasanain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            , Da San Martino, G.,
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Overview of the CLEF-2019 CheckThat!: Automatic Identi cation and Veri cation of Claims. In: Experimental IR Meets Multilinguality, Multimodality, and</article-title>
          <string-name>
            <surname>Interaction. LNCS</surname>
          </string-name>
          , Lugano, Switzerland (
          <year>September 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hasanain</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suwaileh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elsayed</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>Barron-Ceden~o,</article-title>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Overview of the CLEF-</article-title>
          2019
          <source>CheckThat! Lab on Automatic Identi cation and Veri cation of Claims. Task</source>
          <volume>2</volume>
          : Evidence and Factuality
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adair</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>J.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tremayne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The quest to automate fact-checking</article-title>
          .
          <source>In: Proceedings of the 2015 Computation+ Journalism Symposium</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arslan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tremayne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Toward automated fact-checking: Detecting check-worthy factual claims by claimbuster</article-title>
          .
          <source>In: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <year>1803</year>
          {
          <year>1812</year>
          . ACM (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tremayne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Detecting check-worthy factual claims in presidential debates</article-title>
          .
          <source>In: Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          . pp.
          <year>1835</year>
          {
          <year>1838</year>
          . ACM (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hassan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>Arslan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caraballo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gawsane</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joseph</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulkarni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nayak</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          , et al.:
          <article-title>Claimbuster: The rst-ever endto-end fact-checking system</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          <volume>10</volume>
          (
          <issue>12</issue>
          ),
          <year>1945</year>
          {
          <year>1948</year>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jaradat</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gencheva</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barron-Cedeno</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Claimrank:
          <article-title>Detecting check-worthy claims in arabic and english</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>07587</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Karadzhov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <article-title>Barron-Ceden~o,</article-title>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Koychev</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          :
          <article-title>Fully automated fact checking using external sources</article-title>
          .
          <source>arXiv preprint arXiv:1710.00341</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Barron-Ceden~o,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Elsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Marquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Zaghouani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Kyuchukov</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , Da San Martino, G.:
          <article-title>Overview of the CLEF2018 CheckThat! lab on automatic identi cation and veri cation of political claims</article-title>
          .
          <source>In: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          . pp.
          <volume>372</volume>
          {
          <fpage>387</fpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Popat</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Strotgen, J.,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Credeye: A credibility lens for analyzing and explaining misinformation</article-title>
          .
          <source>In: Companion of the The Web Conference 2018 on The Web Conference</source>
          <year>2018</year>
          . pp.
          <volume>155</volume>
          {
          <fpage>158</fpage>
          .
          <string-name>
            <surname>International World Wide Web Conferences Steering Committee</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>