<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Fact Checking: Detecting and Verifying Facts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bilal Ghanem</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Polit`ecnica de Val`encia</institution>
        </aff>
      </contrib-group>
      <fpage>19</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>With the uncontrolled increasing of fake news, untruthful claims, and rumors over the web, recently different approaches have been proposed to address the problem. To distinguish false claims from truthful ones may positively affect the society in different aspects. In this paper we describe the motivations towards fake news research topic in the recent years, we present similar and related research topics, and we show our preliminary work and future plans.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Fake news is an important topic so the
research community become interested in its
identification necessity. Fake news as a topic
has been defined as: “An inaccurate,
sometimes sensationalistic report that is created
to gain attention, mislead, deceive or
damage a reputation”1. Unlike misinformation,
in fake news, authors have previous intention
to pose a misleading sentence. Whereas in
misinformation, authors have inaccurate or
confused information about specific topic.</p>
      <p>The spreading of this phenomenon has
increased in a massive way in online sites. The
openness of the web has led to increase the
number of online news agencies, social
media networks, and online blogs. These
platforms allow certain organizations or
individuals to pose fake news, because they guarantee
them several attractive factors, such as
privacy, free-access, availability, and large
audience. In the same time, a large number of
untrusted news agencies have appeared. These
sites can affect the public opinions about
spe1https://whatis.techtarget.com/definition/fakenews, visited in May 2018.
cific issues, where they have political, social,
or financial agendas.</p>
      <p>
        In a recent approach
        <xref ref-type="bibr" rid="ref11">(Nelson and Taneja,
2018)</xref>
        , the authors investigated the user’s
behavior when visiting these sites. They
examined online visitations data across
different Internet devices. During the US elections
in 2016, they found that the number of
devices that visited trusted news sites was 40
times larger than fake news ones. This
observation shows a hypothesis that the Internet
users are aware of fake news and they are able
to discriminate them from others. Later on,
        <xref ref-type="bibr" rid="ref2">(Bond Jr and DePaulo, 2006)</xref>
        showed that
humans can detect lies only 4% better than
random chance. They analyzed more than
200 meta-data of trusted sites and showed
that their good reputation may contribute to
have more visitors. This study revealed that
the users are not able to detect fake news
and they are affected by the good reputation
of online sites.
      </p>
      <p>
        Improving the search engines ranks of
untrusted sites might make them more popular,
and maybe, truthful. Therefore, fake news
sites’ admins employ website spoofing or
authentic news styling technique to mimic the
hight reputation of the authentic sites in an
attempt to make their sites closely similar.
The last US election has drawn a real fear in
the American nation about fake news
        <xref ref-type="bibr" rid="ref11">(Nelson and Taneja, 2018)</xref>
        . The fast spreading of
fake news motivated the owner of Wikipedia
encyclopedia to create a news site called
WikiTribune2 to promote Evidence-based
journalism.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Rumors</title>
      <p>
        A similar research topic that has attracted
the research community attention recently is
rumors detection. Despite the different
definitions of rumors in the literature, the most
accepted is: “Unverified and instrumentally
relevant information statements in
circulation”
        <xref ref-type="bibr" rid="ref4">(DiFonzo and Bordia, 2007)</xref>
        . Both,
rumors and fake news are similar research
topics, although they have been tackled
independently. Unlike fake news, rumors may turn
out to be true, false, or partly true, wherein
the time of posting the veracity is unknown.
Rumors normally remain in circulation (ex.
retweeted by Twitter users) until a trusted
destination uncover or verify the truth.
Another difference has been shown by
        <xref ref-type="bibr" rid="ref13">(Zubiaga
et al., 2018)</xref>
        : rumors cannot just be
classified by their veracity type (true, false,
halftrue), but also by the credibility degree (high
or low). Previously,
        <xref ref-type="bibr" rid="ref1">(Allport and Postman,
1947)</xref>
        studied rumors from a psychological
perspective. They were interested in
answering why people spread rumors in their
environment. In that time (1947), it was
difficult and complex to find a clear answer. But
in the recent years, a renewed interests
conducted to answer the question: people spread
rumors when there is uncertainty, when they
feel anxiety, and when the information is
important.
      </p>
      <p>
        According to
        <xref ref-type="bibr" rid="ref13">(Zubiaga et al., 2018)</xref>
        ,
rumors in literature have been studied from
different perspectives: rumors detection
(rumor or not), tracking (in social media,
detecting posts dealt with these rumors), stance
classification (how each user or post is
oriented towards a rumor’s veracity) and
veracity prediction (true, false, or unverified).
In rumourEval shared task at SemEval-2017
        <xref ref-type="bibr" rid="ref3">(Derczynski et al., 2017)</xref>
        , the organizers
proposed two different subtasks: stance
classifi2https://www.wikitribune.com, visited in May
2018.
cation, and veracity prediction.
        <xref ref-type="bibr" rid="ref5">(Enayet and
El-Beltagy, 2017)</xref>
        have achieved the highest
result in veracity prediction. They used the
percentage of replying tweets, the existence of
hashtags, and the existence of URLs as
features for a classification model. According
to the previous tasks, the veracity prediction
was the one more related to fake news
detection.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Fact Checking in Presidential</title>
    </sec>
    <sec id="sec-4">
      <title>Debates</title>
      <p>Fake news gained attention in the 2016 US
presidential elections, where both Democrats
and Republicans blamed each other for
spreading false information. According to a
public poll, many people after the election
believed that fake news had affected the election
results. Even during the presidential debate,
journalists found that there were many false
claims between candidates. Their goal was
to weaken or to build bad reputation of the
contenders. These false claims could draw
real effect on the elections results. Especially,
presidential debates are long by their nature
and to detect false claims may need more
time than the needed to spread these false
claims among the public. Also, these debates
contain large number of sentences accused by
each candidate and most of these claims are
un-factual claims (opinions). These opinions
composed another challenge to the
journalists to filter them from the factual claims.
This real issue has taken attention by
CLEFLab 2018. In this lab, two shared tasks have
been proposed: check worthiness and
factuality. In the following, we will give brief review
about these two tasks, and show our
preliminary approach
3.1</p>
      <p>
        CLEF-2018 Check That Lab
As mentioned above, two different tasks have
been proposed at CLEF-2018 Check That lab
        <xref ref-type="bibr" rid="ref10">(Nakov et al., 2018)</xref>
        . The first task was
similar to rumor detection. The idea is to
detect claims that are worthy for checking. The
other task is complementary, where the
factuality of these factual claims is needed to be
checked.
      </p>
      <p>
        Task 1 - Check-Worthiness: A set of
presidential debates from the US
presidential election is presented for the task, where
each claim in the debate text has been tagged
manually as worth to be checked or not. The
full text of the debates is used in the task to
allow participants to exploit contextual
features in the debates. The task goal is to
detect claims that are worthy of checking and
to rank them from the most worthy one for
checking to the lowest. Our preliminary
approach for this task
        <xref ref-type="bibr" rid="ref6 ref7">(Ghanem et al., 2018b)</xref>
        was inspired from previous works proposed
by
        <xref ref-type="bibr" rid="ref8">(Granados et al., 2011)</xref>
        and
        <xref ref-type="bibr" rid="ref12">Stamatatos
(2017)</xref>
        . The authors in the former work have
used a text distortion technique to enhance
thematic text clustering by maintaining the
words that have a low frequency in
documents. Similarly, in the latter work, the
same text distortion technique was used for
authorship attribution. The authors
maintained the words that had the highest
frequency in the documents to detect the
author from his/her writing style. In this work,
we used the same text distortion technique to
detect worthy claims. We believed that this
type of tasks is more thematic than
stylistic, where the writing style is not as
important as the thematic words. We have
maintained the thematic words (that have the
lowest frequency) using a ratio of C; the higher
value of C is, the more thematic words are
maintained. Also, we maintained a set of
linguistic cue words (LC) that were used
previously by
        <xref ref-type="bibr" rid="ref9">(Mukherjee and Weikum, 2015)</xref>
        to infer the news credibility. Additionally,
we maintained also named entities from
being distorted, such as: Iraq, Trump,
America. Through the manually checking of the
claims, we found the checking worthy claims
tended to list different types of named
entities. After applying the distortion process,
the new version of the text was used by
Bagof-Chars using the Tf-Idf weighting scheme.
The new distorted text became less biased by
the high frequency words, such as stopwords.
For the ranking purpose, we used a method
inspired from the K-Nearest Neighbor (KNN)
classifier to rank these worthy claims based
on the distance to the nearest neighbor. For
the classification process, we used KNN
classifier. During the training phase of the task,
the Average Precision @N was used. In the
testing phase, a set of testing files were
provided by the organizers and the Mean
Average Precision (MAP) measure was employed.
In Table 1 the results for the task are
presented. It is worth to mention that the task
has been organized also in Arabic, where the
English claims were translated manually. In
the English part of the task, our approach
      </p>
      <p>Team
Prise de Fer
Copenhagen
UPV-INAOE</p>
      <p>bigIR
Fragarach</p>
      <p>blue
RNCC</p>
      <p>English
0.1332
0.1152
0.1130
0.1120
0.0812
0.0801
0.0632
(UPV-INAOE) has achieved the third
position among seven teams. In the Arabic part,
only two teams have submitted their results.
Similarly to English, the results are close and
there is not a big difference among them. We
believed that the low result of our approach
in the Arabic part is because we translated
the LC lexicons automatically and a manual
translation might have been more reliable.
Task 2 - Factuality: As we mentioned
above, this task concerns with detecting the
factuality of the claims from the US
presidential election. The claims that are
unworthy for checking have not been annotated
and kept in the debates to maintain the
context. Factual claims have been tagged as
True, False, and Half-True. These debates
are provided in two languages, English and
Arabic, similarly to the previous task. The
macro F1 score was used as the performance
measure.</p>
      <p>
        Our approach for this task was based on
the hypothesis that factual claims have been
discussed and mentioned in online news
agencies. In our approach
        <xref ref-type="bibr" rid="ref10 ref13 ref6 ref7">(Ghanem et al., 2018a)</xref>
        ,
we used the distribution of these claims in
the search engines results3. Furthermore, we
supposed that truthful claims have been
mentioned more by trusted web news agencies
than untruthful ones. Thus, our approach
depended on modeling the returned results
from search engines using similarity measures
with the reliabilities of the sources. Our
feature set consists of two types: dependent
and independent features. For the
dependent features, we used cosine over
embedding between a claim query and each of the
first N results from the search engines. We
used the main sentence components to built
the sentences embeddings, discarding
stop3We used in our experiments both Google and
Bing search engines.
words. Also, for each result, we extracted
the AlexaRank value for its site, to capture
the reliability of the result source. Finally,
another text similarity feature was used to
measure the similarity using the full
sentences, without using embeddings and
discarding stopwords. From these set of
features, we built other dependent features that
capture the distribution of some of these
previous ones. We used Standard Deviation and
Average of the previous cosine similar values,
and in a similar manner, for AlexaRank
values.
      </p>
      <p>The official results of task 2 are shown in
Table 2. In general, the low results of both
complex tasks can give an intuition of how
much this research topic is. Further work is
needed.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Future Work</title>
      <p>Fact checking became recently an even more
interesting research topic. We think that
detecting worthy claims should be the first step
in fake news identification. We will address
these issues both in political debates and
social media. To the best of our knowledge,
the previous works on fact checking only
concentrated on validating the facts using
external resources. We will work on
investigating facts from different aspects:
linguistically, structurally, and semantically. In this
vein, SemEval-2019 lab4 proposes two tasks
for the next year: first to determine rumor
veracity and support for rumors, which is
similar to the one that was proposed
previously in SemEval-2017. Secondly, fact
checking in community question answering forums,
which is a new environment for
investigating facts veracity. This shows the high
interest of the research community in these two
research topics. Therefore, participating in
these tasks is one of our future plans.</p>
      <p>4http://alt.qcri.org/semeval2019/index.php?id=tasks,
visited on May 2018</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Allport</surname>
            ,
            <given-names>G. W.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Postman</surname>
          </string-name>
          .
          <year>1947</year>
          .
          <article-title>The psychology of rumor</article-title>
          . Oxford, England: Henry Holt.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Bond</given-names>
            <surname>Jr</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. F.</surname>
          </string-name>
          and
          <string-name>
            <surname>B. M.</surname>
          </string-name>
          <year>DePaulo</year>
          .
          <year>2006</year>
          .
          <article-title>Accuracy of deception judgments</article-title>
          .
          <source>Personality and social psychology Review</source>
          ,
          <volume>10</volume>
          (
          <issue>3</issue>
          ):
          <fpage>214</fpage>
          -
          <lpage>234</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Derczynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Procter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. W. S.</given-names>
            <surname>Hoi</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Semeval-2017 task 8: Rumoureval: Determining rumour veracity and support for rumours</article-title>
          .
          <source>arXiv preprint arXiv:1704</source>
          .
          <fpage>05972</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>DiFonzo</surname>
            , N. and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Bordia</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Rumor, gossip and urban legends</article-title>
          .
          <source>Diogenes</source>
          ,
          <volume>54</volume>
          (
          <issue>1</issue>
          ):
          <fpage>19</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Enayet</surname>
            ,
            <given-names>O. and S. R.</given-names>
          </string-name>
          <string-name>
            <surname>El-Beltagy</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Niletmrg at semeval-2017 task 8: Determining rumour and veracity support for rumours on twitter</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017)</source>
          , pages
          <fpage>470</fpage>
          -
          <lpage>474</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Ghanem</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
            Montes-y G`omez,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Rangel</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          . 2018a.
          <article-title>Upv-inaoe - check that: An approach based on external sources to detect claims credibility</article-title>
          .
          <source>In Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum, CLEF '18</source>
          ,
          <string-name>
            <surname>Avignon</surname>
          </string-name>
          , France, September.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Ghanem</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
            Montes-y G`omez,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Rangel</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
          </string-name>
          . 2018b.
          <article-title>Upv-inaoe - check that: Preliminary approach for checking worthiness of claims</article-title>
          .
          <source>In Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum, CLEF '18</source>
          ,
          <string-name>
            <surname>Avignon</surname>
          </string-name>
          , France, September.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Granados</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cebrian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Camacho</surname>
          </string-name>
          , and F. de Borja Rodriguez.
          <year>2011</year>
          .
          <article-title>Reducing the loss of information through annealing text distortion</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          ,
          <volume>23</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1090</fpage>
          -
          <lpage>1102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Weikum</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Leveraging joint interactions for credibility analysis in news communities</article-title>
          .
          <source>In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management</source>
          , pages
          <fpage>353</fpage>
          -
          <lpage>362</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , A. Barr´on-Ceden˜o,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Suwaileh</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>M`arquez</article-title>
          , W. Zaghouani,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atanasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kyuchukov</surname>
          </string-name>
          , and G. Da San Martino.
          <year>2018</year>
          .
          <article-title>Overview of the CLEF-2018 CheckThat! Lab on automatic identification and verification of political claims</article-title>
          . In P. Bellot,
          <string-name>
            <given-names>C.</given-names>
            <surname>Trabelsi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mothe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Murtagh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Soulier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sanjuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappellato</surname>
          </string-name>
          , and N. Ferro, editors,
          <source>Proceedings of the Ninth International Conference of the CLEF Association: Experimental IR Meets Multilinguality</source>
          , Multimodality, and Interaction, Lecture Notes in Computer Science, Avignon, France, September. Springer.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Nelson</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Taneja</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The small, disloyal fake news audience: The role of audience availability in fake news consumption. new media &amp; society</article-title>
          , page
          <volume>1461444818758715</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Authorship attribution using text distortion</article-title>
          .
          <source>In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>1</volume>
          ,
          <string-name>
            <surname>Long</surname>
            <given-names>Papers</given-names>
          </string-name>
          , volume
          <volume>1</volume>
          , pages
          <fpage>1138</fpage>
          -
          <lpage>1149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Zubiaga</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liakata</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Procter</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Detection and resolution of rumours in social media: A survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>51</volume>
          (
          <issue>2</issue>
          ):
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>