<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Approaches to the Pro ling Fake News Spreaders on Twitter Task in English and Spanish</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jacobo Lopez Fernandez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Antonio Lopez Ram rez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politecnica de Valencia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper discusses the decisions made approaching PANs Pro ling Fake News Spreaders on Twitter Task at CLEF 2020. We brie y describe how we combined author tweets to create samples that do or do not represent a Fake News Spreader. We decided to handle both languages proposed for this task: Spanish and English; and the methodologies that we suggested were Linear Support Vector Machines (SVMs) and Gradient Boosting, respectively. Other approaches such as Long ShortTerm Memory (LSTM) were taken into account in the process of nding a model with the best accuracy results and these were also reported in this paper. We made use of the cross-validation scenario to obtain accuracy results due to the reduced amount of data. We have managed to achieve average accuracy scores of 0.735 for the Spanish language identi cation task and 0.685 for the English language identi cation task.</p>
      </abstract>
      <kwd-group>
        <kwd>author pro ling</kwd>
        <kwd>fake news</kwd>
        <kwd>multilingual</kwd>
        <kwd>social media</kwd>
        <kwd>spreaders</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The trust in information read through social media has steadily increased in the
last few years. However, allow their accounts to publish and propagate
misinformation with severe consequences for our society. First of all, we should make it
clear that there are di erent types of misinformation and disinformation, such
as fake news, satire or rumours that go viral in online social networks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In
addition, psycho-linguistic information as emotion, sentiment or informal
language should be previously analised. Exploiting information extracted from user
pro les and user interactions, we should be able to classify them depending on
the information obtained.
      </p>
      <p>A great amount of fake news and rumors are propagated in online social
networks with the aim, usually, to deceive users and formulate speci c opinions
[15]. Users play a critical role in the creation and propagation of fake news
online by consuming and sharing articles with inaccurate information either
intentionally or unintentionally.</p>
      <p>To prevent dissemination of misinformation and disinformation we may
prole fake news spreaders automatically. Author pro ling is based on the detection
of certain characteristics in pro les making use of linguistic pattern recognition
techniques.</p>
      <p>
        This paper discusses the decisions made approaching PANs' Pro ling Fake
News Spreaders on Twitter Task at CLEF 2020 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The task consists in, given
a Twitter feed, determine whether its author is keen to be a fake news spreader.
So, it focuses on identifying possible fake news spreaders on social media as
a rst step towards preventing fake news from being propagated among online
users. The task has a multilingual perspective, so that includes tweets in English
and Spanish. It is de ned as a binary classi cation task.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Fake news detection has attracted a lot of research attention in the last years.
Guess et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] made an approach to this eld by doing research during the
elections, in particular the 2016 US election procedure. In that paper, they
propose a system that obtained features from polls published on Facebook. Popat
et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] suggested an end-to-end model to evaluate trust on random texts,
without human supervision. Subsequently, they presented a biLSTM neural network
model which aggregates signals from external evidence articles, the language of
these articles and the trustworthiness of their sources.
      </p>
      <p>
        Shu et al. [14] pointed that fake news spreaders cannot be pro led precisely
based only on text content, but we should understand the correlation between
user pro les on social media and fake news. They state that social engagements
should be used as auxiliary information to improve fake news detection systems.
In addition, Sliva et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], distinguished that, approaching content from a data
mining perspective, we could identify patterns that could mark a text as fake.
These patterns, such as clearness which make the text more readable and could
convince the receiver even when it is fake. Collecting this kind of information
produces a huge, unestructured, incomplete and noisy data; di cult and
expensive to manage. Giachanou et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] proposed EmoCred that incorporates
emotions that are expressed in the claims into an LSTM network to di erentiate
between fake and real claims.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Fake News Spreaders Detection Systems</title>
      <p>First, we apply the same type of preprocessing for both, the English and Spanish
tasks data, following the next steps:
{ Load tweets from XML les.
{ Concatenate tweets forming a chain for every author. All tweets in this chain
are separated by a blank space. We apply this technique on the English
dataset and the Spanish dataset.</p>
      <p>
        With the concatenated data, we vectorized our samples. The vectorizers used
to perform this task were CountVectorizer and T dfVectorizer [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
CountVectorizer creates valuable data from counting words in samples while T
dfVectorizer takes into consideration more common words in detriment of those which
are less common.
      </p>
      <p>Concurrently, the tokenizer selected was casual tokenize, an implementation
of TweetTokenizer from NLTK, due to its suitability to manage characters and
expressions commonly used on the Twitter social network.</p>
      <p>After all these transformations, we ended up getting a feature matrix for each
of the languages proposed, English and Spanish.
3.1</p>
      <sec id="sec-3-1">
        <title>First approaches</title>
        <p>The rst method applied to classify our samples employed Recurrent Neural
Networks (RNN), making use of pretrained word embeddings in order to help
representing words as real-valued vectors and lead to better performance of our
neural network system.</p>
        <p>
          For the Spanish language task, the embeddings loaded by our system were the
Spanish Billion Words Corpus and Embeddings [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] which had been trained using
word2vec and consists of near 1 million words where each of them is represented
as a vector with a size of 300.
        </p>
        <p>
          For the English language task, the embeddings loaded by our system had
been trained using GloVe from Stanford [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and collected by Laurence Moroney,
composed of near 6000 billion words where each of them is represented as a
vector with a size of 100.
        </p>
        <p>Once the embeddings were loaded, we trained our RNN model with LSTM
and the following topology:
Despite the fact that the results obtained making use of RNN were close to
those reported in the state of the art for this kind of tasks, we did not reach
promising results as we will explain later in this paper. At that point, we made
use of classi ers provided by the framework scikit-learn and chose the Gradient
Boosting algorithm and the linear SVM algorithm for the English and Spanish
tasks, respectively.</p>
        <p>
          The main core of Gradient Boosting [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] consists of a predictive model based
on decision trees, built step-by-step allowing the optimization of a di erentiable
loss function. For this function we made use of linear regression or 'sigmoid',
called 'deviance' in the scikit-learn framework, whose mathematical expression
is represented as the following:
(1)
(2)
(3)
(4)
where fm stands for the mth weak classi er and m is the corresponding
weight. It is exactly the weighted combination of M weak classi ers. The function
which gives the weight for the mth weak classi er is the following:
where m is the lowest weighted classi cation error.
        </p>
        <p>
          From another point of view, the main core of SVM [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] is based on the concept
of separating a group of points (samples) into two di erent categories. As a
consequence, our model had to be able to classify the sample correctly into
its category. SVM looks for a hyperplane which optimally separates the points
belonging the two classes. Subsequently, we look for the hyperplane with the
longest distance (margin) to the closest points to it.
        </p>
        <p>The equation of the hyperplane in the 'M' dimension can be given as:
(z) =</p>
        <p>1
1 + e z</p>
        <p>
          We also experimented using the AdaBoost [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] algorithm along with the loss
function, called 'exponential' in the scikit-learn framework. This algorithm
focuses on classi cation problems and aims to convert a set of weak classi ers into
a strong one. The nal equation for classi cation can be represented as:
M
F (x) = sign( X
m=1
        </p>
        <p>mfm(x))
m =
where wi are vectors, b is biased term and xi are input variables.</p>
        <p>Furthermore, given a group of points S = (x1; c1); :::; (xN ; cN ) and a constant
C&gt;0, we should obtain weights 2 &lt;d, 0 2 &lt; and the tolerance parameter
&amp; 2 &lt;N to minimize the following expression:
(5)
(6)
(7)
(8)
Dependant of the following two expressions:</p>
        <p>N
X &amp;n
n=1
cn( txn + 0)
1
&amp;n; 1
n</p>
        <p>N
&amp;n
In this section we discuss about the dataset provided and the experimental
adjustments we made.
The given dataset is composed of two folders, a folder for the Spanish language
and a folder for the English language, which contain:
{ A XML le by author or Twitter user pro le. There are 100 tweets in each</p>
        <p>XML le.</p>
        <p>{ A text le with the list of authors and the ground truth.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experimental Settings</title>
        <p>
          The submission of our system was made from the TIRA platform [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. As
the participants were only provided with the training data, we applied
crossvalidation into 10 folds in order to test our system. Di erent classi ers were
tested such as Support Vector Machine, Gaussian Naive-Bayes, Gradient
Boosting, Stochastic Gradient Descent, K-nearest Neigbours and two Neural Network
approaches (Multilayer Perceptron and Convolutional Recurrent Neural
Networks with LSTM). Regarding the English task we obtained the best results
with a Gradient Boosting model and regarding the Spanish task we obtained
the best results with a linear SVM model. Concerning SVM, we set the penalty
parameter 'C' to 100 and the tolerance parameter to 0.01, with the maximum
number of stages set to 100. In our SVM implementation we operated with the
hinge loss function whereas in our Gradient Boosting implementation we
operated with the deviance function. Learning rate was set to 0.01 for Gradient
Boosting and the number of boosting stages to perform was 250. To implement
our system we made use of the scikit-learn framework 1.
        </p>
        <p>For the Spanish language task we used a language model based on bigrams
and trigrams where punctuation marks are processed and a vocabulary is built
considering the top 1000 features, ordered by frequency from the whole corpus.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Table 1 shows the results achieved by our system in the Spanish and English
tasks in terms of precision over the training data. As we mentioned earlier in
this article, we are using a 10 fold cross-validation where the average precision
of the 10 folds gives the nal precision result for each classi er. The best results
in terms of precision were those given by the linear SVM model for the Spanish
language task and the Gradient Boosting model for the English language task.
We should mention that the results shown in this table were obtained without
modifying any hyperparameter of the two classi ers.
1 https://scikit-learn.org
In this paper we described our system for PANs Pro ling Fake News Spreaders
on Twitter Task at CLEF 2020. Regarding the English language task we
proposed a system trained with Gradient Boosting algorithm while for the Spanish
language task we proposed a system based on linear SVM. The state of the
art tells us that Neural Networks are currently the best solution for this kind
of classi cation tasks, but the results that we achieved do not match with this
statement. We can consider that some tasks are not suitable to be addressed
with this kind of systems so far, as we saw with our implementation of
Convolutional Recurrent Neural Network with LSTM. Eventually, a traditional machine
learning algorithm performed better.</p>
      <p>Our results showed that the input data processing considerably conditions
performance in our system. In addition, from our results we can observe that
our Spanish language solution performs better compared to our English language
solution, as we managed to achieve 0.735 and 0.685 accuracy respectively.
14. Kai Shu, Suhang Wang, and Huan Liu. Understanding user pro les on social media
for fake news detection. In Proceedings - IEEE 1st Conference on Multimedia
Information Processing and Retrieval, MIPR 2018, pages 430{435. Institute of
Electrical and Electronics Engineers Inc., June 2018. 1st IEEE Conference on
Multimedia Information Processing and Retrieval, MIPR 2018 ; Conference date:
10-04-2018 Through 12-04-2018.
15. Xinyi Zhou and Reza Zafarani. A survey of fake news: Fundamental theories,
detection methods, and opportunities. 12 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Cristian</given-names>
            <surname>Cardellino</surname>
          </string-name>
          .
          <article-title>Spanish billion words corpus and embeddings</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Theodoros</given-names>
            <surname>Evgeniou</surname>
          </string-name>
          and
          <string-name>
            <given-names>Massimiliano</given-names>
            <surname>Pontil</surname>
          </string-name>
          .
          <article-title>Support vector machines: Theory and applications</article-title>
          . volume
          <volume>2049</volume>
          , pages
          <fpage>249</fpage>
          {
          <fpage>257</fpage>
          , 01
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Freund</surname>
          </string-name>
          and
          <string-name>
            <given-names>Robert E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          .
          <article-title>A short introduction to boosting</article-title>
          .
          <source>Journal of Japanese Society for Arti cial Intelligence</source>
          ,
          <volume>14</volume>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jerome</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Friedman</surname>
          </string-name>
          .
          <article-title>Greedy function approximation: A gradient boosting machine</article-title>
          .
          <source>Annals of Statistics</source>
          ,
          <volume>29</volume>
          :
          <fpage>1189</fpage>
          {
          <fpage>1232</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Bilal</given-names>
            <surname>Ghanem</surname>
          </string-name>
          , Paolo Rosso, and Francisco Rangel.
          <article-title>An Emotional Analysis of False Information in Social Media and News Articles</article-title>
          .
          <source>ACM Transactions on Internet Technology (TOIT)</source>
          ,
          <volume>20</volume>
          (
          <issue>2</issue>
          ):1{
          <fpage>18</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Anastasia</given-names>
            <surname>Giachanou</surname>
          </string-name>
          , Paolo Rosso, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Crestani</surname>
          </string-name>
          .
          <article-title>Leveraging emotional signals for credibility detection</article-title>
          .
          <source>In Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR'
          <volume>19</volume>
          , page
          <volume>877</volume>
          {
          <fpage>880</fpage>
          , New York, NY, USA,
          <year>2019</year>
          .
          <article-title>Association for Computing Machinery</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Guess</surname>
          </string-name>
          , Jonathan Nagler, and
          <string-name>
            <given-names>Joshua</given-names>
            <surname>Tucker</surname>
          </string-name>
          .
          <article-title>Less than you think: Prevalence and predictors of fake news dissemination on facebook</article-title>
          .
          <source>Science Advances</source>
          ,
          <volume>5</volume>
          :
          <fpage>eaau4586</fpage>
          ,
          <year>01 2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Je rey Pennington, Richard Socher, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>1532</fpage>
          {
          <fpage>1543</fpage>
          ,
          <string-name>
            <surname>Doha</surname>
          </string-name>
          , Qatar,
          <year>October 2014</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Kashyap</given-names>
            <surname>Popat</surname>
          </string-name>
          , Subhabrata Mukherjee,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Yates</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Weikum</surname>
          </string-name>
          . Declare:
          <article-title>Debunking fake news and false claims using evidence-aware deep learning</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1809</year>
          .06416,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Martin</surname>
            <given-names>Potthast</given-names>
          </string-name>
          , Tim Gollub, Matti Wiegmann, and
          <article-title>Benno Stein. TIRA Integrated Research Architecture</article-title>
          . In Nicola Ferro and Carol Peters, editors,
          <source>Information Retrieval Evaluation in a Changing World</source>
          . Springer,
          <year>September 2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. Francisco Rangel, Anastasia Giachanou, Bilal Ghanem, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          .
          <source>Overview of the 8th Author Pro ling Task at PAN</source>
          <year>2020</year>
          :
          <article-title>Pro ling Fake News Spreaders on Twitter</article-title>
          . In Linda Cappellato, Carsten Eickho , Nicola Ferro, and Aurelie Neveol, editors,
          <source>CLEF 2020 Labs and Workshops</source>
          , Notebook Papers. CEURWS.org,
          <year>September 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Omid</surname>
            <given-names>Shahmirzadi</given-names>
          </string-name>
          , Adam Lugowski, and
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Younge</surname>
          </string-name>
          .
          <article-title>Text similarity in vector space models: A comparative study</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1810</year>
          .00664,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kai</surname>
            <given-names>Shu</given-names>
          </string-name>
          , Amy Sliva, Suhang Wang,
          <string-name>
            <given-names>Jiliang</given-names>
            <surname>Tang</surname>
          </string-name>
          , and Huan Liu.
          <article-title>Fake news detection on social media: A data mining perspective</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          ,
          <volume>19</volume>
          , 08
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>