<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Don't Just Drop Them: Function Words as Features in COVID-19 Related Fake News Classification on Twiter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pascal Schröder</string-name>
          <email>pascal.schroeder@ru.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Radboud University</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This research shows that function words can be useful as features for machine learning models tasked with detecting conspiratorial content in COVID-19 related Twitter posts. A significance test exposes that the distribution of function words between fake and legitimate content varies greatly. Further, a support vector machine classifier is demonstrated to perform above chance when using function word-only features, achieving a Matthews correlation coeficient of 0.139 on unseen test data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Previous research into detecting conspiratorial online content
indicates that it can be distinguished by the author’s writing style
[
        <xref ref-type="bibr" rid="ref10 ref11 ref2">2, 10, 11</xref>
        ]. Simultaneously, function words provide a meaningful
proxy to an author’s writing style in authorship attribution [
        <xref ref-type="bibr" rid="ref1 ref12">1, 12</xref>
        ].
Therefore, function words could be valuable in aiding machine
learning models tasked with detecting conspiratorial content.
However, many approaches in fake news classification still rely on a
purely content-based approach, in which function words are
excluded as part of the preprocessing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. While these approaches
oftentimes ofer impressive performances, it is paramount that
potentially relevant features are not excluded in the process.
      </p>
      <p>
        Fortunately, recent years have seen a growing body of research
on the importance of stylistic features in misinforming and
conspiratorial content. For instance, Posadas-Durán et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] showed
that better performance levels can be reached for classifiers when
function words are incorporated in the training data, versus when
they are not. However, like many approaches, they have used
online news articles as their data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Arguably, the style entertained
by authors of social media posts will be diferent, and it is to be
expected that the results do not generalise.
      </p>
      <p>
        For social media and especially Twitter data, the literature is
rather sparse. Del Tredici and Fernández [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] have classified articles
shared on Twitter as fake or real, and enhanced their data with
the user’s post history and profile description, and found more
function words in the latter. However, they have not classified
posts directly, but rather articles linked in posts. Niven et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
have used function words as a proxy for ‘thoughtfulness’, arguing
that the latter correlates with the fakeness of a post’s content, but
have not found a significant diference in distribution between fake
and legit content. But since they only had available posts from 300
diferent users, individual authors could have possibly skewed the
distribution. Im et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] have found an above chance performance
for function word-only features when predicting whether a post
      </p>
      <p>Word
my
is
the
they
used
am
during
their
he
she</p>
      <p>P-Value
# No Consp. # Consp.
stems from a Russian troll account, which while related, is still not
focused on fake news detection specifically.</p>
      <p>This research thus aims to expand on the available literature
by analysing how function word usage is distributed between
authors of conspiratorial versus legitimate content, and quantifying
whether function words alone can act as suficient features in the
classification of fake news content.
2</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        The data were provided through the MediaEval 2021 Conference,
for the task ‘FakeNews: Corona Virus and Conspiracies Multimedia
Analysis’ [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. It consists of 1554 Twitter posts related to
COVID19 and diferent conspiracy theories. 1 Three diferent class labels
are provided: tweets that do not mention conspiracies (1), tweets
that discuss conspiracies without actively supporting them (2), and
tweets that promote or support conspiracies (3). Note that the classes
are imbalanced, with 767 examples for class 1, 271 examples for
class 2, and 516 examples for class 3.
      </p>
      <p>To investigate a possible diference in distribution between
conspiratorial and non-conspiratorial content, most of this research
thus focused on classes 1 and 3. First, function words were extracted
from all data using spaCy.2 To test for significance, a  2 test was
performed on the distribution of function words between the classes
1 and 3.</p>
      <p>Next, the usefulness of the extracted function words was tested.
As a representative model, a support vector machine (SVM) with
non-linear kernels was chosen, and evaluated using classicfiation
1All posts are written in English and were collected between January 17, 2020 and
June 30, 2021.
2For the full list of function words, see https://github.com/explosion/spaCy/blob/
master/spacy/lang/en/stop_words.py
accuracy and Matthews correlation coeficient (MCC). Two
scenarios were investigated, one with data containing all three classes,
and one for data containing the classes 1 and 3 only. To account for
random efects, 100 runs were computed for both scenarios, each
with a random 10% validation split. As a final evaluation, a single
SVM model was fitted on all available data, and evaluated on an
unseen test set. Due to organisational means, this evaluation is only
available for the 3-class case using MCC. Preprocessing of function
words was done using the TfidfVectorizer of SciKit Learn.3
3</p>
    </sec>
    <sec id="sec-3">
      <title>RESULTS AND ANALYSIS</title>
      <p>The results of the  2 test can be seen in Table 1, showing the 10
function words with lowest p-value, all of which are below 1%. 5
are pronouns (50%), which is higher than the overall frequency of
pronouns in the function words (≈11.48%). Only 2 out of the 10
words occur more often in class 3 (conspiracy), while the remaining
8 occur more often in class 1 (no conspiracy).</p>
      <p>The following are all sub 5% function words, 40 in total, sorted
from lowest to highest p-value:
my, is, the, they, used, am, during, their, he, she,
by, this, him, serious, doing, might, his, if, us, but,
be, these, all, seem, about, part, her, along, could,
your, due, have, are, here, using, at, per, when,
would, now</p>
      <p>Of these words, 12 are pronouns (30%), marked in italic. 7 belong
to class 3, marked in bold, while 33 belong to class 1.</p>
      <p>Table 2 shows the results of the SVM classifiers.
4</p>
    </sec>
    <sec id="sec-4">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>The results of the  2 test show that a significant diference in the
distribution of function words can be observed between
conspiratorial and non-conspiratorial content, confirming the idea that
such words are indeed important for distinguishing fake from
legitimate content. This is further supported by the classification
results, where an above chance performance on unseen test data
was achieved. Unsurprisingly, the classification performance in the
binary case was higher than in the full case, since an overlap in
style between non-conspiratorial content as well as content which
does not actively support conspiracies, but merely discusses them,
is to be expected.</p>
      <p>
        Interestingly, the category of function words most common in the
list of sub 5% p-value words were pronouns, for which the relative
frequency was greatly increased compared to their frequency in
all function words. Further, all of these pronouns except their and
these, occurred more often in the non-conspiracy category. Most
prominent are third person singular pronouns (he, she, him, his), all
featured more in class 1, which indicates that conspiracy authors
are less likely to talk about a person at length. This could be because
giving extensive detail (e.g. ‘She said X ’) rather than implying makes
their claims falsifiable, which they might be interested to avoid [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Interestingly, this stands in contrast to Rashkin et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], who
found a higher frequency of pronouns in conspiratorial content. Of
further interest is the fact that the pronoun my displays the most
significance overall. This could again be because conspiratorial
3https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.
TfidfVectorizer.html
P. Schröder
authors want to steer the argument away from their own opinion to
a more general claim, thereby avoiding responsibility. This finding
is somewhat supported by Newman et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], who found higher
usage rates of the pronoun I in people who are lying.
      </p>
      <p>Apart from pronouns, the overall majority of function words
with p-values below 5% belong to the non-conspiratorial class. This
indicates that authors of fake content use a more simplistic style,
as the complexity of a text correlates with the number of diferent
function words used.</p>
      <p>An earlier analysis on a smaller subset of the data showed
different patterns in the function word distributions, most notably
the presence of ‘hedging’ words like quite, rather and somehow.
However, these patterns disappeared when the larger data set was
released. Therefore, it is important to note that the data set at hand,
with only 1554 total posts, is a very limited subset of all COVID-19
related data found on Twitter. Thus, it cannot be ruled out that the
patterns found in this research, although powerful in predicting on
the chosen dataset, may not generalise.</p>
      <p>This limitation is extended by the fact that the data analysed
in this report does not contain author information. As stylistic
information correlates very strongly with the author of a text, the
patterns found could, in theory, be caused by a few authors having
a disproportionately high representation in the data. This efect
unfortunately could not be accounted for due to the missing
authorship information.</p>
      <p>In conclusion, this research has shown that function words are a
strong proxy for detecting conspiratorial content in the context of
COVID-19 related fake news on Twitter. To address the limitations
of this research, future work should explore in how far these results
generalise to larger corpora of Twitter data, diferent domains of
conspiracy, as well as other social media platforms.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>I thank Martha Larson for her critical input during the development
of the research question as well as the analysis process, and Lynn
de Rijk for her input on function word usage and interesting lexical
patterns in fake news content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Levitan</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Measuring the Usefulness of Function Words for Authorship Attribution</article-title>
          .
          <article-title>Proceeding of the Joint Conference on Association for Literary and</article-title>
          Linguistic Computing/Association Computer Humanities.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Nicollas R. de Oliveira</surname>
          </string-name>
          , Pedro S. Pisa, Martin Andreoni Lopez, Dianne Scherly V. de Medeiros, and
          <string-name>
            <surname>Diogo</surname>
            <given-names>M.F.</given-names>
          </string-name>
          <string-name>
            <surname>Mattos</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Identifying fake news on social networks based on natural language processing: Trends and challenges</article-title>
          .
          <source>Information (Switzerland) 12 (Jan</source>
          .
          <year>2021</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          . Issue 1. https://doi.org/10.3390/info12010038
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Lynn de Rijk</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>You Said it? How Mis-</article-title>
          and
          <string-name>
            <surname>Disinformation Tweets</surname>
          </string-name>
          <article-title>Surrounding the Corona-5G-conspiracy Communicate Through Implying</article-title>
          .
          <source>Proceedings of the MediaEval 2020 Workshop</source>
          , Online,
          <fpage>14</fpage>
          -
          <issue>15</issue>
          <year>December 2020</year>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jane</given-names>
            <surname>Im</surname>
          </string-name>
          ,
          <string-name>
            <surname>Eshwar</surname>
            <given-names>Chandrasekharan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jackson Sargent</surname>
          </string-name>
          ,
          <string-name>
            <surname>Paige</surname>
            <given-names>Lighthammer</given-names>
          </string-name>
          , Taylor Denby, Ankit Bhargava, Libby Hemphill, David Jurgens,
          <string-name>
            <given-names>and Eric</given-names>
            <surname>Gilbert</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Still out There: Modeling and Identifying Russian Troll Accounts on Twitter</article-title>
          .
          <source>In 12th ACM Conference on Web Science (WebSci '20)</source>
          .
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . https://doi.org/10.1145/3394231.3397889
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Newman</surname>
          </string-name>
          , James Pennebaker, Diane Berry, and
          <string-name>
            <given-names>Jane</given-names>
            <surname>Richards</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Lying Words: Predicting Deception from Linguistic Styles</article-title>
          .
          <source>Personality &amp; Social Psychology Bulletin</source>
          <volume>29</volume>
          (
          <year>June 2003</year>
          ),
          <fpage>665</fpage>
          -
          <lpage>75</lpage>
          . https: //doi.org/10.1177/0146167203029005010
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Timothy</given-names>
            <surname>Niven</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hung-Yu Kao</surname>
          </string-name>
          , and
          <string-name>
            <surname>Hsin-Yang Wang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Profiling Spreaders of Disinformation on Twitter: IKMLab and Softbank Submission</article-title>
          .
          <source>In CLEF</source>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>FakeNews: Corona Virus and Conspiracies Multimedia Analysis Task at MediaEval 2021</article-title>
          .
          <source>Proceedings of the MediaEval 2021 Workshop</source>
          , Online,
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          December
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Petra Filkuková, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>WICO Text: A Labeled Dataset of Conspiracy Theory and 5G-Corona Misinformation Tweets</article-title>
          .
          <source>Proceedings of the 2021 Workshop on Open Challenges in Online Social Networks</source>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Juan</given-names>
            <surname>Pablo</surname>
          </string-name>
          Posadas-Durán, Helena Gomez-Adorno,
          <string-name>
            <given-names>Grigori</given-names>
            <surname>Sidorov</surname>
          </string-name>
          , and Jesús Jaime Moreno Escobar.
          <year>2019</year>
          .
          <article-title>Detection of fake news in a new corpus for the Spanish language</article-title>
          .
          <source>Journal of Intelligent and Fuzzy Systems</source>
          <volume>36</volume>
          (May
          <year>2019</year>
          ),
          <fpage>4868</fpage>
          -
          <lpage>4876</lpage>
          . Issue 5. https://doi.org/10.3233/ JIFS-179034
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hannah</surname>
            <given-names>Rashkin</given-names>
          </string-name>
          , Eunsol Choi, Jin Yea Jang, Svitlana Volkova, and
          <string-name>
            <given-names>Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Truth of Varying Shades: Analyzing Language in Fake News and Political Fact-Checking</article-title>
          .
          <source>In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          , Copenhagen, Denmark,
          <fpage>2931</fpage>
          -
          <lpage>2937</lpage>
          . https://doi.org/10.18653/v1/
          <fpage>D17</fpage>
          -1317
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kai</surname>
            <given-names>Shu</given-names>
          </string-name>
          , Amy Sliva, Suhang Wang,
          <string-name>
            <given-names>Jiliang</given-names>
            <surname>Tang</surname>
          </string-name>
          , and Huan Liu.
          <year>2017</year>
          .
          <article-title>Fake News Detection on Social Media: A Data Mining Perspective</article-title>
          . Special Interest Group on
          <article-title>Knowledge Discovery in Data: Explorations Newsletter 19 (Aug</article-title>
          .
          <year>2017</year>
          ). https://doi.org/10.1145/3137597.3137600
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>A Survey of Modern Authorship Attribution Methods</article-title>
          .
          <source>Journal of the Association for Information Science and Technology 60 (March</source>
          <year>2009</year>
          ),
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          . https://doi.org/10.1002/asi.21001
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Del Tredici</surname>
          </string-name>
          and
          <string-name>
            <given-names>Raquel</given-names>
            <surname>Fernández</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Words are the Window to the Soul: Language-based User Representations for Fake News Detection</article-title>
          .
          <source>In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics</source>
          , Barcelona,
          <source>Spain (Online)</source>
          ,
          <fpage>5467</fpage>
          -
          <lpage>5479</lpage>
          . https: //doi.org/10.18653/v1/
          <year>2020</year>
          .coling-main.
          <fpage>477</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>