<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CLEF</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Sadness and Fear: Classification of Fake News Spreaders' Content on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>ILC CNR</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy irene.russo@ilc.cnr.it</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>22</volume>
      <fpage>22</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>The vast amount of accurate and inaccurate information circulating on the internet requires computational methodologies to detect low-quality content. This kind of content often constitutes fake news, as in the PAN @ CLEF 2020 competition Profiling Fake News Spreaders on Twitter. This competition asks for systems that identify possible fake news spreaders on social media as a first step to prevent fake news from being propagated among online users. In this paper, the methodology used for this classification task is reported. Preprocessing of the data and the features extracted to classify fake news spreaders is explained. A regression-as-classification approach that enables the representation of being a fake news spreader as a gradable one is proposed. The performance (accuracy) on the training and the test set with the different sets of features is reported.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Nowadays, news production is not exclusive to official media outlets: everybody can
report about events. This tendency has positive consequences on the freedom of speech
- especially in countries where this fundamental human right is menaced - but it also
presents several risks. The vast amount of accurate and inaccurate information
circulating on the internet requires computational methodologies to detect low-quality content.
The fake information that spreads on social media can be dangerous for public debates
on societal issues, increasing the general level of anxiety and affecting the behavior of
the population in case of emergency [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. User-generated content such as pictures and
short videos are a potential source of rumors that should be carefully verified [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Similarly, reporting news with link sharing can be harmful, especially if the news source is
not reliable [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        On social media information, disinformation (intentionally false content, created to
cause harm) and misinformation (false content shared without the user realized it)
coexist [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], making it hard to detect reliable channels of information.
      </p>
      <p>
        We can dedicate a limited amount of time and attention to the verification of a source of
information; moreover, repetitions of rumors make them more plausible for everyone
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Harmful content is often labeled as fake news, as in the PAN @ CLEF 2020
competition Profiling Fake News Spreaders on Twitter[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This competition asks for systems
that identify possible fake news spreaders on social media as a first step to prevent fake
news from being propagated among online users.
      </p>
      <p>Fake news is a standard label used in the NLP community, a trendy term denoting, in
reality, many different phenomena that require various features/approaches to be
detected and classified: rumors, propaganda, satire, hoaxes, etc. The ubiquity of the term
hides the fact that the NLP community lacks an informed typology of fake news types.
This typology would need insights from political science and cognitive psychology to
discover the most harmful kind of fake news: someone spreading inaccurate
information about a pandemic is not dangerous as someone tweeting about last gossip involving
Jennifer Lopez.</p>
      <p>
        Fake news are not directly the focus of PAN @ CLEF 2020 competition Profiling Fake
News Spreaders on Twitter [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The organizers are instead concerned with the
identification of Twitter profiles that frequently share articles with inaccurate information
(intentionally or not), contributing to the creation and propagation of fake news online.
According to the organizers, the identifying possible fake news spreaders on social
media is the first step towards preventing fake news from being propagated among online
users.
      </p>
      <p>In this paper, the experiments aiming at this classification task are reported. Pre-processing
of the data and the features extracted to classify fake news spreaders is explained. A
regression-as-classification approach that enables the representation of being a fake
news spreader as a gradable one is proposed. The performance (accuracy) on the
training and the test set with the different sets of features is reported.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>
        The starting hypothesis of the methodology reported in this paper is that emotions have
a crucial role in identifying fake news spreaders because fake content contains
emotionally charged words. The role of emotions and emotions’ intensity for the detection
of fake news has been investigated by [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that propose EmoCred, a system
incorporating emotional signals into a LSTM neural network for the classification of credible
and non-credible claims contained in fact-checking datasets. They experimented with
lexicon-based emotional analysis, the emotional intensity of words, and a neural
network for generating the intensity level of emotional reactions, achieving an accuracy
ranging from 0.608 to 0.628 (depending on the dataset and the methodology tested).
The results improve the baseline - a LSTM for classification of texts - showing the
relevance of emotional signals for fake news classification.
      </p>
      <p>With a focus on the personality traits of fake news spreaders and fake news checkers
on Twitter, the methodology proposed by [?] addresses the problem of fake news at the
users’ level using also linguistics patterns found in users’ posts to decide if a user is a
potential spreader or checker. Their system - CheckerOrSpreader - is a model based on
a CNN network and handcrafted features that refer to the linguistic patterns and
personality traits and can classify a user as a potential fake news checker or spreader (0.59
as F1 score).</p>
      <p>
        Linguistic features that characterize fact-checkers on Twitter have been analyzed by [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
to create a deep learning framework that generates responses with fact-checking
intention. Fact-checkers prefer a formal language, avoiding swear words and Internet slang;
the text generation framework proposed outperforms other text generation approaches
quantitatively and qualitatively.
      </p>
      <p>
        Apart from textual features, user social engagements can be used to distinguish users
that share real news from users who share fake news on Twitter [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. For example, a
comparative analysis of explicit and implicit profile features reveals that users sharing
fake news tend to express more “favor” actions. Their predicted age is slightly bigger
when compared with users that share real news.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        With the assumption that the property of being a fake news spreader could be a gradable
one dependent on several characteristics of user-generated content, the classification
task proposed by PAN @ CLEF 2020 competition Profiling Fake News Spreaders on
Twitter was addressed as a regression one. A random forest regressor [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] was
implemented because it outperforms other regressions algorithms on this dataset. Since the
random forest regressor’s output is a decimal number, it has been rounded to get the
reliability class for each processed instance and then compute the accuracy.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Data Pre-processing</title>
        <p>
          The set of features used in the regression-as-classification experiments concern stylistic
aspects (e.g., use of emphatic punctuation marks), intended communicative functions of
tweets (e.g., mentioning other users) and the emotional profiles of the feed. Concerning
this latter aspect, the occurrences of emotion words in aggregated tweets can help to
detect the tendency to be a fake news spreader. Thanks to the NRC Affect Intensity
Lexicon [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], a manually annotated dataset of 6,000 English words collected with a
technique called best–worst scaling (BWS), an intensity value for eight emotions can
be derived.
        </p>
        <p>– RTTR: root type-token ratio is a measure commonly used in NLP to assess the
complexity of a text;
– mentions: number of mentioned users in the Twitter feed. Since the training set has
been anonymized, it is impossible to have an idea of the variability and the type of
mentioned users;
– replies: number of replies in the Twitter feed;
– urls: number of URLs in the Twitter feed;
– hashtags: number of hashtags in the Twitter feed;
– emoticon: number of emoticons in the Twitter feed;
– emphatic?: number of question marks in the Twitter feed;
– emphatic!: number of exclamation marks in the Twitter feed
– rich_people: the sum of occurrences of rich people’s names in the Twitter feed. The
list is composed by the world’s highest-paid celebrities according to Forbes;
– all_emotion: the sum of values for all lemmas associated with all the emotions in
the Twitter feed;
– fear: the sum of values for all lemmas associated with this emotion;
– trust: the sum of values for all lemmas associated with this emotion;
– anger: the sum of values for all lemmas associated with this emotion;
– sadness: the sum of values for all lemmas associated with this emotion;
– joy: the sum of values for all lemmas associated with this emotion;
– disgust: the sum of values for all lemmas associated with this emotion;
– anticipation: the sum of values for all lemmas associated with this emotion;
– surprise: the sum of values for all lemmas associated with this emotion.</p>
        <p>To understand which features could be more discriminative for the two classes, a
correlation analysis between each feature and the class value is proposed in Table 1.
In Table 2, the accuracy for different combinations of features on the training set is
reported, applying random forest regressor and evaluating with 10-cross fold validation.
Random forest regressors are not deterministic; for this reason, the mean accuracy for
ten runs is reported.</p>
        <p>– all features: ’rttr’,’mentions’,’urls’,’hashtags’, ’replies’,’rich_people’, ’emoticons’,
’emphatic?’, ’emphatic!’, ’emotions_words’, ’emotions_fear’, ’emotions_trust’,
’emotions_anger’,’emotions_sadness’, ’emotions_joy’, ’emotions_disgust’,’emotions_anticipation’,
’emotions_surprise’
– communicative features: ’rttr’,’mentions’,’urls’,’hashtags’, ’replies’,’rich_people’
– stylistic features: ’emoticons’, ’emphatic?’, ’emphatic!’
– emotions words: ’emotions_fear’, ’emotions_trust’, ’emotions_anger’,’emotions_sadness’,
’emotions_joy’, ’emotions_disgust’,’emotions_anticipation’, ’emotions_surprise’
– best features: ’hashtags’, ’emotions_fear’, ’emotions_sadness’
features FNS_En FNS_Es
best features 0.702 0.698
emotion words 0.635 0.527
stylistic features 0.517 0.575
communicative features 0.629 0.687
all features 0.687 0.713</p>
        <p>Table 2. Accuracy results for the training set.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results on the test set</title>
        <p>The test set of PAN @ CLEF 2020 competition Profiling Fake News Spreaders on
Twitter is composed by 400 Twitter feeds. The best model, including number of hashtags
and resulting from 10 cross-fold validation on the training set has been used for the
regression-as-classification on the test set. Results on the test set are reported in
Table 3.</p>
        <p>dataset Accuracy Accuracy baseline
FNS_En 0.58 0.74
FNS_Es 0.5150 0.79</p>
        <p>Table 3. Accuracy results for the test set.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>In this paper, the methodology used to identify fake news spreaders for the PAN @
CLEF 2020 competition Profiling Fake News Spreaders on Twitter is described. After
explaining data’s pre-processing and the features extracted to classify fake news
spreaders, a regression-as-classification approach is proposed; it represents being a fake news
spreader as a gradable property. Performance (accuracy) on the training and the test set
with the different sets of features is reported. The accuracy is below the baseline
provided for this task and not in line with the results obtained on the trained set with the
same methodology. As a consequence, further investigations are needed.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alexander</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          :
          <article-title>Social media in disaster risk reduction and crisis management</article-title>
          .
          <source>Science and Engineering Ethics</source>
          ,
          <volume>20</volume>
          pp.
          <fpage>717</fpage>
          -
          <lpage>733</lpage>
          (
          <year>2014</year>
          ). https://doi.org/10.1007/s11948-013-9502-z
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Random forests</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>45</volume>
          pp.
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Leveraging emotional signals for credibility detection</article-title>
          .
          <source>In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . p.
          <fpage>877</fpage>
          -
          <lpage>880</lpage>
          . SIGIR'
          <volume>19</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA (
          <year>2019</year>
          ). https://doi.org/10.1145/3331184.3331285, https://doi.org/10.1145/3331184.3331285
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Gorrell</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochkina</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liakata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zubiaga</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Derczynski</surname>
          </string-name>
          , L.:
          <article-title>SemEval-2019 task 7: RumourEval, determining rumour veracity and support for rumours</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <fpage>845</fpage>
          -
          <lpage>854</lpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (Jun
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>S19</fpage>
          -2147, https://www.aclweb.org/anthology/S19-2147
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>McCreadie</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Richard,
          <string-name>
            <given-names>M.C.</given-names>
            ,
            <surname>Iadh</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          :
          <article-title>Crowdsourced rumour identification during emergencies</article-title>
          .
          <source>In: Proceedings of the 24th International Conference on World Wide Web</source>
          . pp.
          <fpage>965</fpage>
          -
          <lpage>970</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Word affect intensities</article-title>
          .
          <source>In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ).
          <article-title>European Language Resources Association (ELRA), Miyazaki</article-title>
          , Japan (May
          <year>2018</year>
          ), https://www.aclweb.org/anthology/L18-1027
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghanem</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 8th Author Profiling Task at PAN 2020: Profiling Fake News Spreaders on Twitter</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Eickhoff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Névéol</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.)
          <article-title>CLEF 2020 Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR Workshop Proceedings (Sep</source>
          <year>2020</year>
          ),
          <article-title>CEUR-WS</article-title>
          .org
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Scheufele</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>N.M.:</given-names>
          </string-name>
          <article-title>Science audiences, misinformation, and fake news</article-title>
          .
          <source>PNAS 16</source>
          ,
          <fpage>7662</fpage>
          -
          <lpage>7669</lpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1073/pnas.1805871115
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Liu, H.:
          <article-title>Understanding user profiles on social media for fake news detection</article-title>
          .
          <source>In: 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR)</source>
          . pp.
          <fpage>430</fpage>
          -
          <lpage>435</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Vo</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Learning from fact-checkers: Analysis and generation of fact-checking language</article-title>
          .
          <source>In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . p.
          <fpage>335</fpage>
          -
          <lpage>344</lpage>
          . SIGIR'
          <volume>19</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA (
          <year>2019</year>
          ). https://doi.org/10.1145/3331184.3331248, https://doi.org/10.1145/3331184.3331248
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wardle</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Understanding Information Disorder</article-title>
          . https://firstdraftnews.org/wpcontent/uploads/2019/10/Information_Disorder_Digital_AW.pdf?x76701 (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>