<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TOBB-ETU at CLEF 2019: Prioritizing Claims Based on Check-Worthiness</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bahadir Altun</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mucahid Kutlu</string-name>
          <email>m.kutlug@etu.edu.tr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TOBB University of Economics and Technology</institution>
          ,
          <addr-line>Ankara</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years, we witnessed an incredible amount of misinformation spread over the Internet. However, it is extremely time consuming to analyze the veracity of every claim made on the Internet. Thus, we urgently need automated systems that can prioritize claims based on their check-worthiness, helping fact-checkers to focus on important claims. In this paper, we present our hybrid approach which combines rule-based and supervised methods for CLEF-2019 Check That! Lab's Check-Worthiness task. Our primary model ranked 9th based on MAP, and 6th based on R-P, P@5, and P@20 metrics in the o cial evaluation of primary submissions.</p>
      </abstract>
      <kwd-group>
        <kwd>Fact-Checking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Check-Worthiness Learning-to-Rank</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        With the fast developing information retrieval (IR) technologies and social media
platforms such as Twitter and Facebook, it is very easy to reach any kind of
information we are looking for and share it with other people. Therefore, any
information can be popular worldwide by the incredible sharing behavior of the
Internet users. As a worrisome nding, a recent study [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] reports that false news
spread eight times faster than true news.
      </p>
      <p>
        Obviously, false news can have many undesired consequences and be very
harmful for our daily life. Many researchers and journalists are trying to
minimize the spread of misinformation and its negative impacts. For example,
factchecking websites, such as Snopes1, investigate the veracity of claims spread
on the Internet and publish their ndings. However, these valuable e orts are
not enough to e ectively combat the false news because fact-checking is a very
time-consuming task, taking around one day for a single claim [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>Considering the vast amount of claims spread on the Internet on a daily basis
and high cost of fact-checking, there is an urgent need to develop systems that
Algorithm 1 Claim Ranking Algorithm
1: Input: Training Data TrD
2: Input: Test Data TD
3: FT rD = extract f eatures(T rD)
4: model = train M ART (FT rD)
5: FT D = extract f eatures(T D)
6: ranked claims = test(model; FT D)
7: ranked claims = apply rules(ranked claims; RU LES)
8: return ranked claims
assist fact-checkers to focus on the important claims instead of spending their
precious time for less important claims.</p>
      <p>
        In this paper, we present our method for CLEF-2019 Check That Lab!'s[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
Check Worthiness task[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We use a hybrid approach in which claims are ranked
using a supervised method and then reranked based on hand-crafted rules. In
particular, we use MART [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] learning-to-rank algorithm to rank the claims.
Features we investigate include topical category of claims, named entities,
partof-speech tags, bigrams, and speakers of the claims. We also develop rules to
detect the statements that are not likely to be a claim, and put those statements
at the very end of our ranked lists. Our primary model ranked 9th based on
MAP, and 6th based on R-P, P@5, and P@20 metrics in the o cial evaluation
of primary submissions.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Proposed Approach</title>
      <p>We propose a hybrid approach which uses both rule-based and supervised
methods to rank claims based on their check-worthiness. Now we explain our approach
in detail.
2.1</p>
      <sec id="sec-3-1">
        <title>Hybrid Approach</title>
        <p>
          We rst rank claims using a supervised method and then rerank the claims
using a rule-based method. Algorithm 1 shows the steps in our proposed hybrid
approach. For the supervised ranking phase, we investigated logistic regression
(LR) and various learning-to-rank (L2R) algorithms including MART [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
RankBoost [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], and RankNet[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. In our not-reported initial experiments, we observed
that MART outperforms other L2R methods. Thus, we focus on only MART
and LR in developing our primary and contrastive systems.
        </p>
        <p>
          In the training dataset, we observed that many sentences contain phrases that
are not likely to be a part of a check-worthy claim such as thanks. In addition,
speaker of many statements are de ned as system and these statements contain
non-claim phrases such as applause indicating the actions of the audience during
the debate. Thus, in rule-based reranking phase, inspired from [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], we set score
of a sentence to the minimum score if it contains any of the thanks, thank you,
welcome, and goodbye phrases, or 2) its speaker is system.
2.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Features</title>
        <p>The features we use can be categorized into ve categories: 1) Named entities,
2) part-of-speech tags, 3) topical category, 4) bigrams, and 5) speakers. Now we
explain our features used in our supervised methods.</p>
        <p>Named Entities: Claims about people and institutions are likely to be
check-worthy because people and institutions can be negatively a ected by those
claims. In addition, a check-worthy claim should be also veri able. The location
information, numerical values in claims are easy to verify because they provide
unambiguous information. Therefore, detecting whether a statement is about a
person or an institution, and existence of numerical values, location, and date
information might be good indicators for check-worthy claims. Thus, we identify
the named entities using Stanford Named Entity Tagger2 which tags entities
such as person, location, organization, money, date, time and percentage. We
use a binary vector of size 7 representing the presence of each named entity type
separately.</p>
        <p>Part-of-speech (POS) Tags: Sentences without informative words are less
likely to be check-worthy. POS tags of words might help us to detect the amount
of information in a sentence. For instance, a sentence without any noun such as
"That is great!" does not contain any information that can be fact-checked.
Thus, we detect POS tags in the statements using Stanford POS toolkit3. For
this set of features, we again use a binary feature vector of size 36 in which
each feature represents existence or absence of a particular POS tag in the
corresponding statement.</p>
        <p>
          Topical Category: Topic of a claim might be an e ective indicator for
check-worthy claims [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. For instance, a claim about celebrities might be less
check-worthy than a claim about economics or wars. Therefore, we detect
categories of statements using IBM-Watson's Natural Language Understanding
Tool4. This categorization classi es every statement into topics such as nance,
law, government, and politics, where each topic could be branched up to two
levels of subtopics. We only use the main topic and their immediate subtopics,
yielding 298 features in total. The topics are represented as binary features as
we do for named entities and POS tags.
        </p>
        <p>Bigram: Instead of using a large list of bigrams, we use a limited but
powerful set of bigrams. In this set of features, we use bigrams that appear at least
N times only in check-worthy or only in not check-worthy claims in the training
set. We set N to 50 empirically based on our initial (not-reported) experiments,
yielding 47 features. Some of these bigrams are "All right", "I know", and "go
2 https://nlp.stanford.edu/software/CRF-NER.html
3 https://nlp.stanford.edu/software/tagger.html
4 https://www.ibm.com/watson/services/natural-language-understanding/</p>
        <p>Speaker List: A claim's check worthiness may depend on the person who
makes it. Claims of in uential people are more important than claims of
random people. In addition, some people might make check-worthy claims more
frequently than others. For instance, in the training dataset, we found that
3.8% of Donald Trump's statements and 2.6% of Hilary Clinton's statements
are labeled as check-worthy, showing that Donald Trump is more likely to make
check-worthy claims than Hilary Clinton. Therefore, we use the 33 speakers
appearing in the training data as binary features.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results</title>
      <p>In this section, we present experimental results on both training and test data
using di erent sets of features and machine learning algorithms.</p>
      <p>The training and test datasets contain 19 and 7 les at varying lengths,
respectively, where each le corresponds to transcription of a debate for the US
2016 presidential election. We perform 19-fold cross validation on the training
data where each fold is a separate debate, and calculate mean average precision
(MAP). We use Ranklib5 library in our L2R approach and Scikit toolkit6 for LR.
We set the number of trees and the number of leaves parameters of MART to
50, and 2, respectively based on our initial experiments with di erent parameter
con gurations. We use default parameters for LR.</p>
      <p>We rst investigate the impact of each feature group using LR and MART. In
particular, for each feature group we use all other features to build the model and
observe how the performance would change without the corresponding feature
group. Table 1 shows MAP scores for various feature sets.
5 https://sourceforge.net/p/lemur/wiki/RankLib/
6 https://scikit-learn.org/
category features cause around 2.96%, and 2.28% decrease in the performance
of MART, respectively. However, for LR, POS tags seem the most e ective
features while named entities and topical categories have very small impact on the
performance. Nevertheless, we achieve the best results when we use LR with all
features (0.2198) while LR with All-fNamed Entitiesg, MART with All-fSpeaker
List g, and MART with all features have also similar results.</p>
      <p>For our primary submission, we selected MART algorithm with all features
except speaker list. This is because its performance is very close to our best
performing model and it does not use any speaker list which makes it more
debate-independent model. For our contrastive-1 and contrastive-2 submissions,
we elected MART and LR using all features.</p>
      <p>For the evaluation on test data, we train our selected models with all training
data. Table 2 presents o cial performance scores of all our submissions based
on various evaluation metrics. 12 groups submitted 25 runs in total. Our primary
submission ranked 9th among primary models based on MAP which is the o cial
evaluation metric of the task. Our primary submission is also ranked 6th based
on R-P, P@5, and P@20 metrics, among primary models.
In this section, we take a deeper look into our results and conduct a qualitative
analysis. We manually investigate the top 10 claims in our primary system's
output in each test le. Table 3 shows a sample of claims with their labels and
ranks in our system's output.</p>
      <p>Comparing Row 1 and 2, while both of the statements are about the change
in unemployment, one of them is labeled as check-worthy while the other one
is not. We also observe a similar issue in statements shown in Row 3 and 4
where both statements are about the amount of trade and use similar words
with similar sentence structure but their labels are di erent. This might be due
to many reasons. First, check-worthy claims can be subjective, making the task
even more challenging. Second, we might need background information about the
claims because a claim which has been already discussed among people is more
check-worthy than the one which is not discussed by anyone. The requirement
of background information suggests that context the claim made and external
resources such as web pages, social media posts about the claim can be useful
to detect check-worthiness of claims.</p>
      <p>Regarding the claims in Row 5 and 6, the statements contain veri able claims
and their topic is about economics which can be considered as an important
topic. In addition, the statement in Row 5 has also an explicit accusation
regarding to an administration. However, they are both labeled as not check-worthy.
These samples provide further evidence that detecting check-worthiness of a
claim requires utilizing external resources and its context instead of just using
the claim itself.</p>
      <p>Regarding the claims in Row 7, 8, and 9, we can see that all three claims are
about trade but only claim in Row 9 is labeled as check-worthy. This might be
because the amount of money mentioned in Row 7 and 8 are not exact numbers,
making them hard to verify. On the other hand, the claim in Row 9 mentions an
exact number about the trade de cit. While our features take account whether
a claim has a numerical value or not, this suggest that we have to handle special
cases where numbers are not exact.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we described our methods for CLEF-2019 Check That! Lab's
Check-Worthiness task and discuss results conducting a qualitative analysis. We
proposed a hybrid approach which combines rule-based and supervised
methods together. We rst rank claims using MART algorithm, and then change
the scores of some statements based on our hand-crafted rules. We investigated
various features including topical category, POS tags, named entities, bigrams,
and speakers of claims. Our primary model ranked 9th based on MAP, and 6th
based on R-P, P@5, and P@20 metrics, among primary models. In our
qualitative analysis we discussed that detecting check-worthiness is a subjective task
and linguistic analysis of claims are not enough, requiring utilization of other
resources to capture the background information about claims.</p>
      <p>In the future, we plan to investigate linguistic features to capture the
contextual information about the claims, and develop more e ective hand-crafted
rules.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          , G. Karadzhov,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mohtarami</surname>
          </string-name>
          , and G. Da San Martino.
          <article-title>Overview of the clef-2019 checkthat! lab on automatic identi cation and veri cation of claims. task 1: Check-worthiness</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          , Lugano, Switzerland,
          <year>2019</year>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shaked</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Renshaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lazier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Deeds</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. N.</given-names>
            <surname>Hullender</surname>
          </string-name>
          .
          <article-title>Learning to rank using gradient descent</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on Machine learning (ICML-05)</source>
          , pages
          <fpage>89</fpage>
          {
          <fpage>96</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>T.</given-names>
            <surname>Elsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barron-Ceden</surname>
          </string-name>
          ~o,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hasanain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Suwaileh</surname>
          </string-name>
          , G. Da San Martino, and
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          .
          <article-title>Overview of the clef-2019 checkthat!: Automatic identi cation and veri cation of claims</article-title>
          .
          <source>In Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction,
          <string-name>
            <surname>LNCS</surname>
          </string-name>
          , Lugano, Switzerland,
          <year>September 2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Freund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Iyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <article-title>An e cient boosting algorithm for combining preferences</article-title>
          .
          <source>Journal of machine learning research</source>
          ,
          <volume>4</volume>
          (Nov):
          <volume>933</volume>
          {
          <fpage>969</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Friedman</surname>
          </string-name>
          .
          <article-title>Greedy function approximation: a gradient boosting machine</article-title>
          .
          <source>Annals of statistics</source>
          , pages
          <volume>1189</volume>
          {
          <fpage>1232</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>N.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tremayne</surname>
          </string-name>
          .
          <article-title>Detecting check-worthy factual claims in presidential debates</article-title>
          .
          <source>In Proceedings of the 24th acm international on conference on information and knowledge management</source>
          , pages
          <year>1835</year>
          {
          <year>1838</year>
          . ACM,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Vosoughi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Aral</surname>
          </string-name>
          .
          <article-title>The spread of true and false news online</article-title>
          .
          <source>Science</source>
          ,
          <volume>359</volume>
          (
          <issue>6380</issue>
          ):
          <volume>1146</volume>
          {
          <fpage>1151</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>K.</given-names>
            <surname>Yasser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kutlu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Elsayed</surname>
          </string-name>
          . bigir at clef 2018:
          <article-title>Detection and veri cation of check-worthy political claims</article-title>
          .
          <source>In CLEF (Working Notes)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>C.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Karakas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          .
          <article-title>A hybrid recognition system for checkworthy claims using heuristics and supervised learning</article-title>
          .
          <source>In CLEF 2018 Working Notes. Working Notes of CLEF 2018-Conference and Labs of the Evaluation Forum</source>
          , volume
          <volume>8995</volume>
          , pages
          <fpage>171</fpage>
          {
          <fpage>175</fpage>
          . CEUR-WS. org,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>