<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CLEF</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Analyzing User Profiles for Detection of Fake News Spreaders on Twitter</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>María S. Espinosa</institution>
          ,
          <addr-line>Roberto Centeno, and Álvaro Rodrigo</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Natural Language Processing and Information Retrieval Group Universidad Nacional de Educación a Distancia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>22</volume>
      <fpage>22</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>The massive spread of digital information to which our society is subjected nowadays has led to a great amount of false or extremely biased information being shared and consumed by Internet users every day. Disinformation, including misleading and even false information, is a major issue for our current society. The impact of fake news on politics, economy and even public health is yet to be specified. Internet users must face a high amount of false information in digital media such as rumours, fake news, and extremely biased news. Given the crucial role that the spread of fake news plays in our current society, it is becoming essential to design tools to automatically verify the veracity of online information. In order to address this issue, the PAN@CLEF 2020 competition has proposed a task focused on the detection of fake news spreaders on Twitter. In this paper, we offer a detailed description of the system developed for this competition. Our system relies on psychological features for modelling the behaviour of users.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The rise of social media in the past years has changed the way people consume
information, especially news. The amount of time spent online, immediacy and lower to
nonexistent price are decisive factors for this global change in the ways of news
consumption.</p>
      <p>According to a study conducted by the Pew Research Center in 2016 in the United
States, the percentage of American adults getting their news through social media
increased from 49% in 2012 to 62% in 20161. The same study in 2018 reported a value of
68%, confirming that this number is still increasing2. The shift in how people consume
news during the past few years is undeniable. Nowadays, people are more likely than
ever to use social networks for news instead of more traditional sources such as printed
newspapers and television [23].</p>
      <p>Our society is subjected to a massive exposition of information nowadays, and this
has led to a great amount of false or extremely biased information being shared and
consumed by Internet users every day. In 2013, the need to combat fake news was
already appointed by the World Economic Forum’s Global Risks Report that warned
that “digital wildfires" could spread false information rapidly3.</p>
      <p>
        Fake news has become a global issue, especially since recent social and political
events such as the 2016 U.S. Presidential Election. Online political discussion was
strongly influenced by social media users and bots spreading misinformation, which
potentially altered public opinion and endangered the integrity of the elections.
Furthermore, a paper published in 2019 conducting a study over a dataset with 171 million
tweets in the five months preceding the election day founded that, from the 30
million tweets which contained a link to news outlets, 25% of them spread either fake or
extremely biased news [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The impact of fake news in global economy, public health and even in the creation
of panic in society has been extensively documented in the past few years with
countless examples, such as [15], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [22] and [17]. These are examples of the high cost
associated to the spread of fake news: the absence of control and verification of the
information, which makes social media a fertile ground for the spread of unverified or
false information.
      </p>
      <p>
        With this in mind, we can affirm that the magnitude, diversity and substantial
dangers of fake news and, in more general terms, the disinformation circulating on social
media is becoming a reason of concern due to the potential social cost it may have in
the near future [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. As a consequence, the research community has launched several
evaluation competitions to foster the development of systems able to detect false
information. The Author Profiling task of Profiling Fake News Spreaders on Twitter at the
PAN@CLEF 2020 competition is a good example of one of such competitions [14].
In this paper, we describe the proposal send to the competition, analysing the main
contributions and errors detected.
      </p>
      <p>Our proposal focuses on the content created by a user instead of the content created
by other users as for example retweets. Then, we use a combination of psychological
and linguistic features aimed at modelling the behaviour of a user to detect if they are
spreader or non-spreader.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Over the past few years, several definitions have been given for the term fake news. One
of the most frequent definitions is one that overlaps with the contents of misinformation
and disinformation as well: Fake news are fabricated information that mimics news
media content intentionally created to deceive, mislead or misinform readers [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The
intention behind the creation and dissemination of fake news often has a political or
economic component. Given the crucial role that the spread of fake news plays in our
current society, research on this topic is developing significantly. In fact, the number
of published papers indexed in the the Scopus database concerning the topic of fake
      </p>
      <sec id="sec-2-1">
        <title>3 http://www3.weforum.org/docs/WEF_GlobalRisks_Report_2013.pdf</title>
        <p>news has increased considerably from less than 20 in 2006 to more than 200 in 2018
[25]. These works concentrate on understanding how false information spreads through
social media, and how can it be efficiently detected in order to reduce its negative impact
on society. This task has been approached from different perspectives, such as Natural
Language Processing (NLP), Data Mining (DM), and Social Media Analysis (SMA).</p>
        <p>Recent research proposes an approach which combines text generation and
factchecking in order to mitigate the effects of fake news spreading [24]. In many cases the
task is treated as a binary classification problem where a news piece is classified as fake
or real. However, there are cases in which this classification may not be adequate since
the news could be partially true and partially false. For this reason, systems capable of
multi-class classification have also been proposed [16].</p>
        <p>In the field of Natural Language Processing, research has been focusing on the
detection and intervention of fake news using techniques such as Machine Learning
and Deep Learning [18], and taking into account:
– Content-based features contain information that can be extracted from the text, such
as linguistic features.
– Context-based features contain surrounding information such as user
characteristics, social network propagation features, or users’ reactions to the information.</p>
        <p>Detecting fake news in the context of social media presents characteristics and
challenges that result in content-based methods not being effective on their own. Fake news
are intentionally created to deceive, making it difficult identify them only from their
textual content. For this reason, it is common to use surrounding information such as
the way in which they are disseminated and the behavior of the users involved in this
dissemination, as well as information related to the author of the news [19].</p>
        <p>
          Recent research has demonstrated that studying the correlation of the user profile
and the spread of fake news works for the identification of those users mere likely to
believe fake news and for the differentiation of those more likely to believe real news
[20]. Approaches considering context-based features in combination with content-based
features have been gaining popularity in the past years, due to the promising results
obtained in recent studies, such as [21] and [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Mainly three aspects of this type of
information can be studied:
– User information, such as location, age, number of followers, etc.
– The responses generated by fake news, which can stand as an important source of
detection not only because users use responses to express their opinions but also
because they can help in the construction of a credibility index for users [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
– The social networks through which the news disseminate. The study of the
networks through which the information is propagated has special relevance since the
rapid diffusion of these networks is used to reach the maximum number of users in
the shortest possible time.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Dataset Description and Preprocessing</title>
      <p>The dataset provided by the organizers of the task was divided into two collections of
tweets: one in Spanish and one in English. Each of them contained 300 XML documents
containing 100 tweets written or shared by one user. There was, therefore, information
regarding 300 users. In addition, each directory contained a truth file with the user
identifier and a 0 or a 1 determining the class label 4.</p>
      <p>The user identification numbers were not the real Twitter IDs since these were
obfuscated for privacy reasons. Their purpose was to be able to identify the users in the
truth file as well as in the user XML documents.</p>
      <p>In the same way, the contents of the tweets in what regards mentions, hashtags,
URLs, and usernames were also obfuscated indicating only that one of them had
occurred by the use of a keyword in capital letters and between dollar symbols. Therefore,
a tweet containing the following text:
“RT The new president of the USA is @bartsimpson! #usa https://thenews.com/
new-president-bart-simpson”
would have the following content in the provided dataset:
“RT The new president of the USA is $USER$ ! $HASHTAG$ $URL$”
Our participation in this task was only for the English language, therefore, we only
used the data in the English directory in our model.</p>
      <p>
        With regards to the preprocessing applied to the data, there are three important steps
that we took before the feature extraction:
1. Tokenization. In this step, the words conforming the tweets were separated into
tokens using the RegexpTokenizer from the Natural Language Toolkit (NLKT)
in Python [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which splits a string into substrings using a regular expression that
matches the tokens. In this case, we selected words of 3 or more alphabetic
characters.
2. Stop words removal. After the words were tokenized, the stop words were removed
from the text. For this task, we used the stopword set form the NLTK corpus
combined with a small list of custom words added manually to the set. These words
were commonly used words such as prepositions, coordinating conjunctions, and
determiners that were not included in the NLTK corpus stop word set.
3. Tweet aggregation. After the two previous steps were executed, the resulting
tokens were aggregated conforming a single document per user. The reason for this
aggregation was to have a single piece of text per user before the processing phase.
      </p>
      <p>It is important to notice that, as we will see in the following sections, some of this
preprocessing had to be done later on the processing phase because some of the features
of our model take into account metrics such as the number of stop words, the number
of determiners, or the number of coordinating conjunctions.
4 The information regarding which class label corresponds to fake news spreaders was not
available due to GDPR reasons</p>
    </sec>
    <sec id="sec-4">
      <title>Feature Engineering</title>
      <p>The main contents that users share in social media can be divided in: (1) content
created by the user, and (2) content created by others. Our model for the task of profiling
fake news spreaders on Twitter is based on the following hypothesis: Establishing the
difference between user-created and user-shared content will reveal more accurate
features of the user’s online behavior. This hypothesis states that the individual
analysis of these groups of contents will reveal more precise features of the user profiles.</p>
      <p>For this reason, the process of feature extraction was applied differently to the
content that was originally written by the user (i.e his/her tweets) and to the content shared
by the user but originally written by other users (i.e. retweets). In this section, we will
describe all the feature engineering applied to the data detailing how the distinction
between tweets and retweets was made in each case.</p>
      <p>The complete set of features extracted from the data is depicted in Table 1. The set
of features used in the model can be divided in the following four main categories:
– Psychological features, which will help defining the user profile in order to
differentiate between fake news spreaders and real news spreaders.
– Linguistic features that will help identifying the linguistic traits that identify each
category.
– Twitter actions features. Exploring how the users behave in the social network could
offer some insights on the online behaviour of fake news spreaders.
– Headline analysis data, which can help us in the identification of news pieces that
a user shares that are actually fake.
4.1</p>
      <sec id="sec-4-1">
        <title>Psychological Features</title>
        <p>The psychological features were extracted using a third-party API developed by Symanto5.
The documents containing the aggregated tweets for each user were sent to the API in
order to retrieve the values of their (1) personality traits, (2) their communication styles,
and (3) the sentiment analysis of their text.</p>
        <p>The personality traits value would be either “emotional” or “rational” depending on
the analysis of the user’s text. The value returned by the API when the communication
styles are requested is a collection of traits, such as self-revealing, which means
sharing one’s own experience and opinion; fact-oriented, which implies focusing on factual
information, objective observations or statements; information-seeking, that is, posing
questions; and action-seeking or aiming to trigger someone’s action by giving
recommendation, requests or advice. Finally, the values returned by the sentiment analysis of
the text returns either “positive” o “negative” depending on the sentiment found in the
user’s text.</p>
        <p>All the values returned by Symanto’s API included a percentage for the predicted
value and had to be converted to a binary representation (0,1).</p>
        <sec id="sec-4-1-1">
          <title>5 https://symanto-research.github.io/symanto-docs/</title>
          <p>Category</p>
          <p>
            Feature name Description
Set of values
For the extraction of linguistic features, a natural language pipeline called Polyglot was
used [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. This library is built using distributed word representations (word embeddings)
in conjunction with traditional NLP features for over 100 different languages in order
to solve NLP tasks, such as Part-of-Speech (POS) tagging, Named Entity Recognition
(NER), sentiment analysis, etc.
          </p>
          <p>
            For our model’s set of features we choose 12 POS tagging metrics, 3 named entity
recognition metrics and total word count. Details regarding the specific metrics can be
found in Table 1.
The analysis of the activities of the users in Twitter was restricted by the data
obfuscation described in Section 3. Therefore, only 4 metrics were recorded from the actions
of the user within Twitter: the number of mentions, the URL number, the number of
retweets and the number of hashtags. The values of these metrics were counted from
the total aggregation of tweets of each user.
In this category of the model, we try to study if there are specific message characteristics
that accompany fake news articles being produced and widely shared. Recent studies
suggest that not only these characteristics exist, but also that some of them can be found
in the headline of the news article [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. Therefore, we took the 3 most significative
characteristics differentiating fake news headlines from real news headline and applied them
to the text in our dataset.
          </p>
          <p>Based on the assumption of our main hypothesis being true, we separated tweets
from retweets and applied these measurements only to the retweet subset.
5
5.1</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments and Results</title>
      <sec id="sec-5-1">
        <title>Experiments</title>
        <p>For the creation of our model we first did some experiments in order to select the most
important features as well as the best performing algorithms6. We tested the model
taking into account the set of features available in each category separately, and we also
tested the possible combinations of the features to evaluate their performance on the
data.</p>
        <p>With regards to the classification models, we used an open source machine learning
library, scikit-learn [12]. We performed a comparative analysis in which we tested the
model with some of the most pupular classification algorithms, such as Logistic
Regression, K-Neighbors, Random Forest, Decision Tree, and Support Vector Machines.
6 All the experiments and results can be found in a notebook uploaded to
https://github.com/mariaesp/PANCLEF_PAPER/blob/master/notebooks/Spreaders.ipynb.
Features Measures LogisticRegression KNeighbors RandomForest DecisionTree SVM</p>
        <p>In order to use all the available data in for the tests, we used cross-validation with 5
iterations. The results obtained for each classifier can be seen in Table 2. Results are
given in terms of accuracy, the official measure, as well as precision and recall.</p>
        <p>Features Measures LogisticRegression KNeighbors RandomForest DecisionTree SVM</p>
        <p>After comparing the results obtained with the different categories and classifiers,
we trained the model using a combination of all categories. With regards to the
classification algorithm, the Random Forest Classifier outperformed the others in 3 of the
4 categories, and also in an additional test in which the set of all the features in the
four categories were combined. The results of this experiment can be found in table 3.
Therefore, the Random Forest Classifier was chosen as the algorithm to train our model.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Final Model Definition and Results</title>
        <p>Once all the experiments allowed us to choose the classifier and the set of features for
our model, we trained the model with the data provided for the task and exported our
trained model in order to make the submission and evaluation in the TIRA7 environment
[13].</p>
        <p>There were two evaluations for our model. On the one hand, there was an early-bird
submission evaluation for the task and, on the other hand, there was the final submission
evaluation. We participated in both evaluations, first with and early model and then with
a final model. The evaluation results can be found in table 4.</p>
        <p>Data</p>
        <p>Model Phase
development efianralyl
experimentation
experimentation
test
early
final
early-bird submission
final submission</p>
        <p>Accuracy
0.67
0.68
0.67
0.64</p>
        <p>With regards to the general results of the competition, our team was in position 61
from 66 in the classification. This classification considers results in both Spanish and
English languages calculating the average from both accuracies. However, since our
team only participated in the English part of the task, we have taken into account only
the results in English language in order to see our position. In this case, our team was
positioned 45th form 66 participants. Furthermore, if we aggregate the results, that is,</p>
        <sec id="sec-5-2-1">
          <title>7 https://www.tira.io/</title>
          <p>if we count all the participants with the same results as just one participant, our result
would be 16th from 33 participants.</p>
          <p>It is important to notice that our model evolved from the early-bird submission to
the final sumbmission. There were 3 main changes performed in the model:
– The psychological features could not be included in the first evaluation due to
technical issues with the platform. The organizers helped us to solve those issues and
we could add the psychological features in the final submission.
– The separation of the user original tweets from the user retweets was done after the
early-bird sumbission was completed. As we explained in detail in section 4, this
decision was made based on the assumption of our main hypothesis. Therefore, the
way in which several features of our model were calculated changed for the final
submission as well.
– The fourth category of features, namely the headline analysis data, was added to
the model for the final evaluation submission.</p>
          <p>As it can be observed in the evaluation results, our model performed slightly better
in the evaluation of the early bird submission than in the final submission. The reasons
for this performance drop are unknown to the author, since the final model did perform
better in our experimentation. One possible explanation is that the evaluation dataset
has slightly different characteristics than the training dataset and, therefore, the results
vary accordingly. Nevertheless, the difference in the results is too small to certainly
know the causes.
6</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Future Work</title>
      <p>In this paper, we have described our proposal for the Author Profiling task of
Profiling Fake News Spreaders on Twitter at the PAN@CLEF 2020 competition. Our model
aimed to differentiate fake news spreaders from real news spreaders using a
combination of psychological and linguistic traits extracted from the user’s data, together with
characteristics extracted from both the user behaviour in the social network and the
news headline analysis.</p>
      <p>On the one hand, one of the next experiments will be to use deep learning in the
training phase of the model. The advances made in the development of Recurrent
Neural Networks (RNNs) and Convolutional Neural Networks (CNNs) in the past years
demonstrate promising results in the field of natural language processing.</p>
      <p>
        On the other hand, more work needs to be done with regards to the psychological
and psycholinguistic dimension of the model. The are several psychological models
that we want to explore in the following months, such as the Big5 personality model [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
and the Myers–Briggs Type Indicator (MBTI) model [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Due to the lack of existing
tools for the automatic labelling of these indicators, we will work in the retrieval and
labelling of a larger dataset in order to learn to automatically predict such personality
traits.
      </p>
      <p>As it can be seen in the results exposed in the previous section, our results in
development and test are consistent, which means that, despite the work that needs to be
done in order to improve it, it is a robust model with an expectable performance when
datasets vary.</p>
      <p>This work is at very early stages of development and will continue evolving towards
a more efficient and better performing system. This is why we show preliminary results
of our experiments and there is still room for improvement.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This research project has been supported by the European Social Fund through the
Youth Employment Initiative (YEI 2019) and the Spanish Ministry of Science,
Innovation and Universities (DeepReading RTI2018-096846-B-C21, MCIU/AEI/FEDER,
UE).
12. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M.,
Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D.,
Brucher, M., Perrot, M., Duchesnay, E.: Scikit-learn: Machine learning in Python. Journal
of Machine Learning Research 12, 2825–2830 (2011)
13. Potthast, M., Gollub, T., Wiegmann, M., Stein, B.: TIRA Integrated Research Architecture.</p>
      <p>In: Ferro, N., Peters, C. (eds.) Information Retrieval Evaluation in a Changing World.</p>
      <p>Springer (Sep 2019)
14. Rangel, F., Giachanou, A., Ghanem, B., Rosso, P.: Overview of the 8th Author Profiling
Task at PAN 2020: Profiling Fake News Spreaders on Twitter. In: Cappellato, L., Eickhoff,
C., Ferro, N., Névéol, A. (eds.) CLEF 2020 Labs and Workshops, Notebook Papers.</p>
      <p>CEUR-WS.org (Sep 2020)
15. Rapoza, K.: Can ‘fake news’ impact the stock market? by Forbes (2017)
16. Rashkin, H., Choi, E., Jang, J.Y., Volkova, S., Choi, Y.: Truth of varying shades: Analyzing
language in fake news and political fact-checking. In: Proceedings of the 2017 conference
on empirical methods in natural language processing. pp. 2931–2937 (2017)
17. Rich, M.: As coronavirus spreads, so does anti-chinese sentiment. The New York Times.</p>
      <p>Available from: URL: https://www. nytimes.</p>
      <p>com/2020/01/30/world/asia/coronavirus-chinese-racism. html (2020)
18. Ruchansky, N., Seo, S., Liu, Y.: Csi: A hybrid deep model for fake news detection. In:
Proceedings of the 2017 ACM on Conference on Information and Knowledge Management.
pp. 797–806 (2017)
19. Shu, K., Sliva, A., Wang, S., Tang, J., Liu, H.: Fake news detection on social media: A data
mining perspective. ACM SIGKDD explorations newsletter 19(1), 22–36 (2017)
20. Shu, K., Wang, S., Liu, H.: Understanding user profiles on social media for fake news
detection. In: 2018 IEEE Conference on Multimedia Information Processing and Retrieval
(MIPR). pp. 430–435. IEEE (2018)
21. Shu, K., Zhou, X., Wang, S., Zafarani, R., Liu, H.: The role of user profiles for fake news
detection. In: Proceedings of the 2019 IEEE/ACM International Conference on Advances in
Social Networks Analysis and Mining. pp. 436–439 (2019)
22. Takahashi, R.: Amid virus outbreak, japan stores scramble to meet demand for face masks.</p>
      <p>Japan Times. Consultado el 1 (2020)
23. Tolmie, P., Procter, R., Randall, D.W., Rouncefield, M., Burger, C., Wong Sak Hoi, G.,
Zubiaga, A., Liakata, M.: Supporting the use of user generated content in journalistic
practice. In: Proceedings of the 2017 chi conference on human factors in computing
systems. pp. 3632–3644 (2017)
24. Vo, N., Lee, K.: Learning from fact-checkers: Analysis and generation of fact-checking
language. In: Proceedings of the 42nd International ACM SIGIR Conference on Research
and Development in Information Retrieval. pp. 335–344 (2019)
25. Zhou, X., Zafarani, R.: Fake news: A survey of research, detection methods, and
opportunities. arXiv preprint arXiv:1812.00315 (2018)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Al-Rfou</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perozzi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skiena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Polyglot:
          <article-title>Distributed word representations for multilingual nlp</article-title>
          .
          <source>In: Proceedings of the Seventeenth Conference on Computational Natural Language Learning</source>
          . pp.
          <fpage>183</fpage>
          -
          <lpage>192</lpage>
          . Association for Computational Linguistics, Sofia, Bulgaria (
          <year>August 2013</year>
          ), http://www.aclweb.org/anthology/W13-3520
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Allcott</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gentzkow</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Social media and fake news in the 2016 election</article-title>
          .
          <source>Journal of economic perspectives 31(2)</source>
          ,
          <fpage>211</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bovet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Makse</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <article-title>Influence of fake news in twitter during the 2016 us presidential election</article-title>
          .
          <source>Nature communications 10(1)</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The role of personality and linguistic patterns in discriminating between fake news spreaders and fact checkers</article-title>
          .
          <source>In: Natural Language Processing and Information Systems: 25th International Conference on Applications of Natural Language to Information Systems, NLDB</source>
          <year>2020</year>
          , Saarbrücken, Germany, June 24-26,
          <year>2020</year>
          , Proceedings. p.
          <fpage>181</fpage>
          . Springer Nature
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>L.R.:</given-names>
          </string-name>
          <article-title>An alternative" description of personality": the big-five factor structure</article-title>
          .
          <source>Journal of personality and social psychology 59(6)</source>
          ,
          <volume>1216</volume>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Horne</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adali</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>This just in: fake news packs a lot in title, uses simpler, repetitive content in text body, more similar to satire than real news</article-title>
          .
          <source>In: Eleventh International AAAI Conference on Web and Social Media</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Luo</surname>
          </string-name>
          , J.:
          <article-title>News verification by exploiting conflicting social viewpoints in microblogs</article-title>
          .
          <source>In: Thirtieth AAAI conference on artificial intelligence</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>In washington pizzeria attack, fake news brought real guns</article-title>
          .
          <source>The New York Times</source>
          (
          <year>2016</year>
          ), https://www.nytimes.com/
          <year>2016</year>
          /12/05/business/media/cometping-pong
          <article-title>-pizza-shooting-fake-news-consequences</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lazer</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baum</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benkler</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berinsky</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenhill</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Metzger</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyhan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennycook</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rothschild</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>The science of fake news</article-title>
          .
          <source>Science</source>
          <volume>359</volume>
          (
          <issue>6380</issue>
          ),
          <fpage>1094</fpage>
          -
          <lpage>1096</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Nltk: The natural language toolkit</article-title>
          .
          <source>In: In Proceedings of the ACL Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics</source>
          . Philadelphia: Association for Computational Linguistics (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>I.B.</given-names>
          </string-name>
          :
          <article-title>The myers-briggs type indicator: Manual (</article-title>
          <year>1962</year>
          ).
          <article-title>(</article-title>
          <year>1962</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>