<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Work-
shops October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Kathleen M. Carley</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Shahan Ali Memon</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>1</volume>
      <fpage>9</fpage>
      <lpage>20</lpage>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>From conspiracy theories to fake cures and
fake treatments, COVID-19 has become a
hotbed for the spread of misinformation online. It
is more important than ever to identify
methods to debunk and correct false information
online. In this paper, we present a
methodology and analyses to characterize the two
competing COVID-19 misinformation
communities online: (i) misinformed users or users
who are actively posting misinformation, and
(ii) informed users or users who are actively
spreading true information, or calling out
misinformation. The goals of this study are
twofold: (i) collecting a diverse set of annotated
COVID-19 Twitter dataset that can be used
by the research community to conduct
meaningful analysis; and (ii) characterizing the two
target communities in terms of their network
structure, linguistic patterns, and their
membership in other communities. Our
analyses show that COVID-19 misinformed
communities are denser, and more organized than
informed communities, with a possibility of
a high volume of the misinformation being
part of disinformation campaigns. Our
analyses also suggest that a large majority of
misinformed users may be anti-vaxxers.
Finally, our sociolinguistic analyses suggest that
COVID-19 informed users tend to use more
narratives than misinformed users.</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>With the emergence of COVID-19 pandemic, the
political and medical misinformation has elevated to create
what is being commonly referred to as the global
infodemic. False information has hampered proper
communication, and a ected the decision-making process
[BE+20]. This makes debunking of false information
vitally important. According to one study [TLC15],
if left undisputed, misinformation can in fact
exacerbate the spread of the epidemic itself. Process of
debunking misinformation, however, is complex and one
that is not completely understood [CJHJA17]. This
is because in order to conduct any intervention, it is
rst imperative to be able to identify the
misinformation, as well as the misinformed communities.
Because of the scarcity of data, and diversity of
misinformation themes, this is already a challenging task
in itself, but is also not enough. A second, and
arguably a more important aspect of an intervention is
to be able to correct and change the beliefs of the
misinformed communities. To be able to do this, it
is important to understand how di erent communities
interact, which sub-communities they belong to, and
what are their preferences. In this paper, we
characterize the COVID-19 misinformation communities on
Twitter in terms of their network structure,
linguistic patterns, and membership in other misinformation
and disinformation sub-communities. In the process,
we also design and collect a large annotated dataset
with a comprehensive codebook that we make
available for the community to use for further analysis and
models for misinformation detection.
2
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Background</title>
      <sec id="sec-3-1">
        <title>COVID-19 Datasets</title>
        <p>In the short amount of time, many COVID-19 datasets
have been released. Most of these datasets are generic,
and lack annotations or labels. Examples include
multilingual corpus on a wide variety of topics related
to COVID-19 [CLF20, AMEP+20, HJB+20],
longitudinal Twitter chatter dataset [BTW+20],
multilingual dataset with location information of the users
[QIO20], Twitter dataset for Arabic tweets [AAA20],
Twitter dataset for popular Arabic tweets [HHSE20],
and dataset for identi cation of stance, replies, and
quotes [VCKBC20]. Most of these datasets either have
no annotations at all, employ automated annotations
using transfer learning or semi-supervised methods, or
are not speci cally designed for misinformation.</p>
        <p>In terms of datasets collected for COVID-19
misinformation analysis and detection, examples include
CoAID [CL20] which contains automatic annotations
for tweets, replies, and claims for fake news;
ReCOVery [ZMFZ20] is a multimodal dataset annotated for
tweets sharing reliable versus unreliable news,
annotated via distant supervision; FakeCovid [SN20] is a
multilingual cross-domain fake news detection dataset
with manual annotations; and [DSW20] is a large-scale
Twitter dataset also focused on fake news. A survey
of the di erent COVID-19 datasets can be found in
[LUM+20] and [SAAA20].</p>
        <p>In terms of the diversity of the classes, and the size
of the dataset, the most relevant dataset is by Alam et
al. [ASN+20] who, like our study, present a
comprehensive codebook to annotate tweets on a ner
granularity. Their dataset, however, is limited to a few
hundred tweets, and our dataset is much more diverse in
the range of topics covered. Dharawat et al. [DLMZ20]
present a similar dataset with focus on the severity of
the misinformation. However, their dataset does not
consider the di erent \types" of misinformation.
Finally, Song et al. present a dataset in [SPJ+20] which
contains a diverse set of 10 categories, but still is not
as large, and contains fewer categories in relation to
the dataset collected within our study.
2.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Misinformation Analysis</title>
        <p>A plethora of research has already been conducted
for analysing COVID-19 misinformation online. Some
examples include categorization and identi cation
of misinformed users based on their home
countries, social identities, and political a liation [HC20,
SSM+20], characterization of di erent types
conspiracy theories propagated by Twitter bots [Fer20],
characterization of the prevalence of low-credibility
information related to COVID-19 [YTLM20], exploratory
analysis of the content of COVID-19 tweets [OPR20,
SDM20], understanding the types, sources, and claims
of COVID-19 misinformation [BSHN20], and
comparison of the credibility of COVID-19 tweets to datasets
pertaining to other health issues [BKF+20]. To the
best of our knowledge none of the studies have
characterized COVID-19 misinformation communities in
terms of their sociolinguistic patterns. In this study,
we do not characterize the misinformation content
directly. Instead, we conduct a set of analysis to
understand and characterize the competing COVID-19
communities through their content, and content-sharing
behaviors and interactions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methodology</title>
      <sec id="sec-4-1">
        <title>Data Collection</title>
        <p>To collect Twitter dataset, we use Twitter search API
using a diverse set of keywords as shown in table 1
to collect data. We collected our data on three days:
29th March 2020, 15th June 2020, and 24th June 2020.
Each of these collections extracted a set of tweets from
their corresponding week. For the annotation process,
tweets were randomly sampled from that set.</p>
        <p>Terms
bleach, vaccine, acetic acid, steroids, essential
oil, saltwater, ethanol, children, kids, garlic,
alcohol, chlorine, sesame oil, conspiracy, 5G,
Keywords cure, colloidal silver, dryer, bioweapon,
cocaine, hydroxychloroquine, chloroquine, gates,
immune, poison, fake, treat, doctor, senna
makki, senna tea
#nCoV20199, #CoronaOutbreak,
#CoronaVirus, #CoronavirusCoverup,
#CoroHashtags navirusOutbreak, #COVID19,
#Coronavirus, #WuhanCoronavirus, #coronaviris,
#Wuhan
Our annotation task aims to determine the category
to which a given tweet belongs to. After many
discussions and revisions, we identify 17 categories that a
particular tweet could classify to. These 17 categories
are de ned in table 2. These categories are de ned in
further detail along with their de nitions and
examples in our codebook which we make available for the
public to use.</p>
        <p>Based on these categories, tweets were randomly
and uniformly sampled from the data collection to
maintain diversity in terms of topics covered. In the
rst phase around 4573 tweets were annotated by a
single annotator. Table 2 shows the distribution of
the data in terms of the di erent categories as
annotated by the rst annotator. In the second phase, 651
of these annotated tweets were assigned randomly to
6 other annotators.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Data Description</title>
      <p>Our data collection strategy is di erent from others
in two main aspects: (i) we have a diverse set of
categories taking into consideration di erent types of
information and misinformation online; and (ii) our dataset
is one of the very few, if not the only one, with
emphasis on informed communities with categories such
as \True Prevention", \Calling out/correction", \True
Public Health Response", and \Sarcasm". We believe
this is necessary as building models requires not just
the annotation of false information, but as well as
complementary true information categories.</p>
      <p>At the end, we have 4573 annotated tweets,
comprising of 3629 users with an average of 1:24 tweets
per user. Our annotated data not only covers a
wide range of categories as observed in table 2, but
also covers a wide range of topics as can be seen
in gure 1. We call this dataset CMU-MisCOV19
[MC20]. In adherence to the FAIR principles, the
database and the codebook has been uploaded to
Zenodo and is accessible with the following link:
http://doi.org/10.5281/zenodo.4024154. In adherence
to the Twitter's terms and conditions, we do not
provide the full tweet JSONs, but provide the tweet IDs
so that the tweets can be rehydrated. We also
provide the annotations, and the date of creation for each
tweet for the reproduction of the results of our
analyses. The annotated tweets are included in a CSV
le with the following elds: status id (tweet id of the
tweet), status created at (timestamp of the creation of
the tweet), annotation1 (annotated class of the tweet
by the rst annotator), and annotation2 (annotated
class of the tweet by the second annotator, if exists).</p>
    </sec>
    <sec id="sec-6">
      <title>Analysis and Discussion</title>
      <sec id="sec-6-1">
        <title>Identifying Communities</title>
        <p>Conducting analyses for a competing set of
communities requires identifying those communities rst.
Because we have already annotated data across a set
of true and false information categories, we identify
the membership of the users by assigning a valence
of +1 to the categories True Treatment, True
Prevention, Correction/Calling Out, Sarcasm/Satire, and
True Public Health Response, and a valence of -1 to
the categories Conspiracy, Fake Cure, Fake Treatment,
False Fact or Prevention, and False Public Health
Response. Note that we assign the valence to the
categories (or annotations) and not the tweets themselves.
This is so that we can leverage the annotations from
multiple annotators. At the end, we compute the
valence of each user as a weighted sum of the valence of
the annotations assigned to their tweets. Then we use
the valence assigned to each user to identify their
membership i.e. if valence is greater than 0, the user is
assigned to the informed group, and if the valence is less
than 0, the user is assigned to the misinformed group.
Out of 3629 users, the community detection process
assigns 47% (1697) of the users to the informed group,
29% (1043) of the users to the misinformed group, and
24% (889) of the users to ambiguous or irrelevant
category1.
5.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Data Augmentation</title>
        <p>Because our goal is to characterize communities and
their behaviors, once we identify the two communities,
we collect the timelines of users in each community to
augment our data. Our hypothesis is that these
additional posts can be used to mitigate survivorship bias
[BGIR92] within our analyses. To conduct network
analysis, bot analysis, and sociolinguistic analysis, we
rst extract only the COVID-19 related tweets from
1Irrelevant users are users who have only posted tweets
within other categories such as \Politics" or \Emergency".
Because these categories do not have an assigned valence related
to misinformation, they are not relevant for the purposes of this
study.
the timelines of each user. We do this by ltering all
the tweets by the case-insensitive keywords \corona"
and \covid". This yields a total of 330609 tweets with
an average of 91 tweets per user.
5.3</p>
      </sec>
      <sec id="sec-6-3">
        <title>Network Analysis</title>
        <p>To conduct network analysis, we extract the retweet,
mention, and reply networks of the two target
communities, and combine those networks together. We
then compute the network density for each of the
two groups. As described in [MTMC20], network
density is de ned as the ratio of actual connections
and potential connections. In dense networks,
conformity of the ideas is highly encouraged, and di erence
of opinions is discouraged. We also use ORA-PRO
[CRC, ACR17, ACR18] to plot the network graph as
shown in gure 2</p>
        <p>We note that both the informed and misinformed
users display echo-chamberness with misinformed
subcommunities being much denser than the informed
sub-communities as shown in table 3. We do,
however, notice some two-way communication from both
sides.</p>
        <p>We also plot the retweet, mention and reply
network separately as shown in gure 3. While retweet,
and mention network show little to no two-way
communication, we can observe that the reply network,
while small in size, does in fact have much more
intergroup engagement. We hypothesize that this is likely
a consequence of the \corrective" or \calling-out"
behavior.
To understand the role of bots within the two
competing groups, we used Bot-Hunter [BC18b, BC18a,
BCB+18, BC20], which has a precision of .957 and a
recall of .704, to identify potential bot-like accounts.
We use the probability of greater than or equal to .75
as our con dence threshold to identify bots. We use
a two-sample z-test for the di erence of proportions
( = 0:05) to test the di erence in proportion of bots
between the two competing groups of users. The
results of our analyses can be found in table 4.</p>
        <p>We observe that from a total of 3629 users, 14%
(505) of the users are identi ed as bots. The
percentage of bots within identi ed misinformed users,
however, is much higher (19%) than within identi ed
informed users (11%). We nd our results to be
statistically signi cant (p &lt; 0:001; z = 6:23). This
indicates that more than 1/5th of the misinformation
related posts in our dataset are potentially a result of
disinformation campaigns related to COVID-19.
5.5</p>
      </sec>
      <sec id="sec-6-4">
        <title>Sociolinguistic Analysis</title>
        <p>To understand the linguistic di erences between the
two competing communities, we conduct a linguistic
analysis based on the tweets of the two groups by using
the Linguistic Inquiry and Word Count (LIWC)
program [PBJB15]. LIWC is a text analysis tool which
looks at the di erent lexical categories each of which
is psychologically meaningful. For a given text, LIWC
calculates the percentage of each LIWC categories. All
of these categories are based on word counts.</p>
        <p>We run the LIWC program on the timelines of all
the members for each of the two competing groups.
We only use tweets relevant to COVID-19. We also
remove users identi ed as bots. Because some users
may be more active than others, using the results of
the program as is may introduce biases in our
analyses. To account for those biases, we rst normalize
the percentages by the size of the data for each user.
We use the mean of the normalized LIWC indices of
tweets of individual users for a given lexical category
as our test statistic. We use an independent z-test for
the di erence in means to establish statistical signi
cance. For all our tests, = 0:05. Our analyses are
summarized in table 5.</p>
        <p>For this part, we focus on investigating three
linguistic dimensions, each of which, along with its
linguistic correlates, is described below.
5.5.1</p>
      </sec>
      <sec id="sec-6-5">
        <title>Narrative Discourse Structure</title>
        <p>Narratives play a central role in how individuals
process information, communicate, and reason [Ves17].
We set to test the di erences in the usage of
narratives or anecdotes between the two COVID-19
misinformation communities. The LIWC correlates for
narrative discourse structure include high usage of
function words, pronouns, analytic summary dimension,
and authenticity. High usage of function words and
pronouns happens more often when expressing
feelings and behaviors which tends to happen frequently in
narratives [Pen11]. Moreover, low analytical thinking
also suggests narrative language [PBJB15].
Furthermore, authentic individuals tend to be more personal,
humble, and vulnerable [PBJB15]. Therefore, we use
all of these as proxies to identify variation in the use
of narratives across communities.</p>
        <p>In the past [MTMC20], it has also been suggested
that misinformed communities (eg. anti-vaxxers) tend
to use many more pronouns suggesting highly
narrative discourse structure. In this analysis, however,
we nd that informed users in the COVID-19
discourse use signi cantly more pronouns, more
functional words, mention more family-related keywords,
are less analytical, and more authentic and honest in
comparison to misinformed users. All of these
suggest that informed users may use many more
narratives than misinformed users. This is an interesting
nding as it presents a dichotomy between the di
erent misinformation communities (eg. anti-vaxxers and
COVID-19 misinformed community). In hindsight,
this is also an intuitive result, as our informed group is
obtained from corrective discourse where users present
their stories of family members or friends su ering
from COVID-19 to call out conspiracies and false
information. Because the two communities still seem to
have less two-way communication, this also suggests
that just the content and framing of the message (i.e.
narratives) may not be enough, and perhaps there is
a need to connect the two groups by identifying an
e ective medium of communication.
5.5.2</p>
      </sec>
      <sec id="sec-6-6">
        <title>Tone</title>
        <p>Tone describes how positive a given text is. According
to the de nition by LIWC, the higher the tone index,
the more positive the tone. Indices less than 50
typically suggest a more negative tone. While we do not
see signi cant di erences in the emotional tone of the
competing groups, we nd both the communities to be
highly negative.
5.5.3</p>
      </sec>
      <sec id="sec-6-7">
        <title>Linguistic formality</title>
        <p>Formality of the language has often been considered
as one of the most important dimensions for stylistic
variation. In [GMC+14], authors de ne linguistic
formality as a style of writing that is meant to be precise,
coherent, articulate and convincing to an educated
audience, as opposed to informal discourse which is lled
with deictic references (eg. here, there), pronouns,
and narration. The LIWC correlates to this
dimension are swear words (swear), and informal language
(informal). Informal language in LIWC is computed
on the bases of swear words, netspeak (eg. btw, lol),
non uencies (eg. err, hmm), assents (eg. agree, OK),
and llers (eg. youknow).</p>
        <p>From table 5, it can be observed that misinformed
users tend to be more informal than informed users,
though informed users tend to use more swear words
than misinformed users. This is intuitive as many of
our informed users post corrective or sarcastic tweets
to call out misinformation. However, our results are
not signi cant, and, hence inconclusive.
5.6</p>
      </sec>
      <sec id="sec-6-8">
        <title>Vaccination Stance</title>
        <p>To understand the interplay between the di erent
kinds of misinformation themes and communities, we
identify the vaccination-related stance of the
members of the misinformed sub-community. To do that,
we rst identify the subset of misinformed community
who have posted at least one tweet related to
\vaccines" in the past. We then collect the user-to-hashtag
co-occurrence network. We use the valence of the
vaccination hashtags obtained via the label
propagationbased method mentioned in the study in [MTMC20]
to identify the stance of each member (pro vs. anti)
based on the weighted sum of the valences of the
hashtags. If the weighted sum is greater than 0, we identify
the member as pro-vaxxer, and if the weighted sum is
less than 0, we identify the member as anti-vaxxer.
The distribution of the pro- and anti-vaxxers within
the COVID-19 misinformed group is as shown in table
6.</p>
        <p>We observe that from 1027 COVID-19 misinformed
users in our dataset, 41% of the members are
identied as anti-vaxxers, whereas only 22% of the members
are identi ed as pro-vaxxers. The di erence between
the proportions of the two communities is signi cantly
high. We also identify the proportion of bots within
each of the two groups: misinformed pro-vaxxers, and
misinformed anti-vaxxers. As shown in table 6, 17%
of the misinformed pro-vaxxers are bots, which is
signi cantly lower than the proportion of bots within the
misinformed anti-vaxxers. The rst thing this suggests
is that a big chunk of COVID-19 misinformation online
may in fact be disinformation, and hence, intentional.
The existence of bots within both the informed and
misinformed communities also suggests that much of
the disinformation online may be an organized e ort
to amplify the COVID-19 debate to create discord in
the communities as seen in the past with Twitter bots
and Russian trolls [BJQ+18].
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Limitations</title>
      <p>The rst important limitation pertaining to our work
is that most of our analyses are based on the data that
has been annotated by only 1 annotator. We try to
mitigate this by having more than 1/7th of our
annotations annotated by a second annotator, and taking
into account all those annotations while computing the
membership for each user. Another limitation to our
work is that all our analyses are correlational in nature,
and do not depict causation. A limitation pertaining
to our data collection strategy is that we collect our
data across a period of three weeks, augment our data
with timelines of users, and update our list of hashtags
to account for new themes. We then sample a subset
of this data for annotation process. Because of the
way data was collected, it cannot be used for
assessing change over time. Moreover, while this ensures the
diversity of misinformation-related topics and agents,
it may limit our ability to estimate the actual extent
to which the di erent types of stories are more or less
present. Another limitation related to our bot analysis
is that we use a second-level inference from a trained
model. We try to mitigate this by using labels with
probability greater than or equal to .75 to ensure high
quality labels. Finally, unlike \vaccination" related
discourse, COVID-19 does not have a clear de nition
of the \stance" of the users. This is because there
are many sub-topics associated to COVID-19 each of
which could have its own stance. In this work, we
categorize users based on misinformation. However, the
relationship between misinformation and stance
vis-avis issues is complex, and one that needs to be
understood. In the future work, we hope to explore this
relationship to create a systematic way of
characterizing communities both in terms of misinformation, and
the di erence stances of the users.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>In this paper, we present a methodology to
characterize the competing COVID-19 misinformation
communities by comparing them in terms of their network
structure, sociolinguistic variation, and membership in
disinformation campaigns and in other health-related
misinformation communities such as anti-vaxxers. We
nd that even though COVID-19 is a recent event,
misinformation related to it has created a set of
polarized communities with high echo-chamberness.
Misinformed communities are observed to be denser than
informed communities which is in line with previous
studies such as [MTMC20]. We nd that bots
exist in both the informed and misinformed groups, but
the percentage of bots in misinformed users is
significantly higher suggesting the prevalence of
disinformation campaigns. Our sociolinguistic analysis
suggests that both the target communities depict
negative emotional tone in their posts, with signals that
informed users use many more narratives than
misinformed users. Finally, we discover that many
misinformed users may be anti-vaxxers. Our analyses
suggest that misinformation communities are much
more complex as they are highly organized, and tend
to be highly analytical. Unlike previous suggestions
[SOC19], they may not be responsive to narrative
correctives, and hence, a \one size ts all" generic
messaging intervention for debunking misinformation may
not be a feasible solution. A successful intervention
may require to identify, and ban the disinformation
campaigns. It may also be useful to identify the right
medium of communication to connect the two groups.
This can be achieved by identifying users in
misinformed communities who are not rebroadcasting, or
have high betweenness centrality to be messengers for
disseminating factual information. It may also be
useful to further understand the linguistic patterns and
preferences of these communities to create an e ective
content and framing of the messaging.
7.0.1</p>
      <sec id="sec-8-1">
        <title>Acknowledgements</title>
        <p>This work was partially supported by a fellowship
from Carnegie Mellon University's Center for Machine
Learning and Health to Shahan A. Memon. We thank
David Beskow for access to his Bot-Hunter model for
bot analysis. We also thank members of CMU's
Center for Computational Analysis of Social and
Organizational Systems (CASOS) for insightful comments
and discussions related to the data codebook and its
revisions.
[AAA20]</p>
        <sec id="sec-8-1-1">
          <title>Sarah Alqurashi, Ahmad Alhindi, and</title>
          <p>Eisa Alanazi. Large arabic twitter
dataset on covid-19. arXiv preprint
arXiv:2004.04315, 2020.
[ACR17]
[ACR18]</p>
        </sec>
        <sec id="sec-8-1-2">
          <title>Neal Altman, Kathleen M Carley, and</title>
          <p>Je rey Reminga. Ora user's guide
2017. Carnegie-Mellon Univ. Pittsburgh
PA Inst of Software Research
International, Tech. Rep., 2017.</p>
        </sec>
        <sec id="sec-8-1-3">
          <title>Neal Altman, Kathleen M Carley, and</title>
          <p>Je rey Reminga. Ora user's guide
2018. Carnegie-Mellon Univ. Pittsburgh
PA Inst of Software Research
International, Tech. Rep., 2018.
[AMEP+20] Muhammad Abdul-Mageed,
AbdelRahim Elmadany, Dinesh Pabbi, Kunal
Verma, and Rannie Lin. Mega-cov:
A billion-scale dataset of 65
languages for covid-19. arXiv preprint
arXiv:2005.06012, 2020.
[ASN+20]
[BC18a]
[BC18b]
[BC20]
[BCB+18]</p>
        </sec>
        <sec id="sec-8-1-4">
          <title>Firoj Alam, Shaden Shaar, Alex Nikolov,</title>
          <p>Hamdy Mubarak, Giovanni Da San
Martino, Ahmed Abdelali, Fahim Dalvi,
Nadir Durrani, Hassan Sajjad, Kareem
Darwish, et al. Fighting the covid-19
infodemic: Modeling the perspective of
journalists, fact-checkers, social media
platforms, policy makers, and the society.
arXiv preprint arXiv:2005.00033, 2020.</p>
        </sec>
        <sec id="sec-8-1-5">
          <title>David M Beskow and Kathleen M Carley.</title>
          <p>Bot conversations are di erent:
leveraging network metrics for bot detection
in twitter. In 2018 IEEE/ACM
International Conference on Advances in
Social Networks Analysis and Mining
(ASONAM), pages 825{832. IEEE, 2018.</p>
        </sec>
        <sec id="sec-8-1-6">
          <title>David M Beskow and Kathleen M Car</title>
          <p>ley. Bot-hunter: a tiered approach to
detecting &amp; characterizing automated
activity on twitter. In Conference
paper. SBP-BRiMS: International
Conference on Social Computing,
BehavioralCultural Modeling and Prediction and
Behavior Representation in Modeling and
Simulation, 2018.</p>
        </sec>
        <sec id="sec-8-1-7">
          <title>David Beskow and Kathleen M Carley.</title>
          <p>Social Cybersecurity. Springer, 2020.</p>
        </sec>
        <sec id="sec-8-1-8">
          <title>David Beskow, Kathleen M Carley, Halil</title>
          <p>Bisgin, Ayaz Hyder, Chris Dancy, and
Robert Thomson. Introducing
bothunter: A tiered approach to detection
and characterizing automated activity
on twitter. In International
Conference on Social Computing,
BehavioralCultural Modeling and Prediction and
[BE+20]</p>
        </sec>
        <sec id="sec-8-1-9">
          <title>Darrin Baines, RJ Elliott, et al. De ning</title>
          <p>misinformation, disinformation and
malinformation: An urgent need for clarity
during the covid-19 infodemic.
Discussion Papers, pages 20{06, 2020.</p>
        </sec>
        <sec id="sec-8-1-10">
          <title>Stephen J Brown, William Goetzmann,</title>
          <p>Roger G Ibbotson, and Stephen A Ross.
Survivorship bias in performance
studies. The Review of Financial Studies,
5(4):553{580, 1992.</p>
        </sec>
        <sec id="sec-8-1-11">
          <title>David A Broniatowski, Amelia M Jami</title>
          <p>son, SiHua Qi, Lulwah AlKulaib, Tao
Chen, Adrian Benton, Sandra C Quinn,
and Mark Dredze. Weaponized health
communication: Twitter bots and
russian trolls amplify the vaccine
debate. American journal of public health,
108(10):1378{1384, 2018.</p>
        </sec>
        <sec id="sec-8-1-12">
          <title>David A Broniatowski, Daniel Kerch</title>
          <p>ner, Fouzia Farooq, Xiaolei Huang,
Amelia M Jamison, Mark Dredze, and
Sandra Crouse Quinn. The covid-19
social media infodemic re ects uncertainty
and state-sponsored propaganda. arXiv
preprint arXiv:2007.09682, 2020.</p>
        </sec>
        <sec id="sec-8-1-13">
          <title>J Scott Brennen, Felix Simon, Philip N</title>
          <p>Howard, and Rasmus Kleis Nielsen.</p>
          <p>Types, sources, and claims of covid-19
misinformation. Reuters Institute, 7,
2020.
[BTW+20] Juan M Banda, Ramya Tekumalla,
Guanyu Wang, Jingyuan Yu, Tuo Liu,
Yuning Ding, and Gerardo Chowell.</p>
          <p>A large-scale covid-19 twitter chatter
dataset for open scienti c research{an
international collaboration. arXiv preprint
arXiv:2004.03688, 2020.</p>
        </sec>
        <sec id="sec-8-1-14">
          <title>Man-pui Sally Chan, Christopher R</title>
          <p>Jones, Kathleen Hall Jamieson, and
Dolores Albarrac n. Debunking: A
metaanalysis of the psychological e cacy
of messages countering misinformation.
Psychological science, 28(11):1531{1546,
2017.</p>
        </sec>
        <sec id="sec-8-1-15">
          <title>Xiaolei Huang, Amelia Jamison, David</title>
          <p>
            Broniatowski, Sandra Quinn, and Mark
Dredze. Coronavirus twitter data: A
[MTMC20] Shahan Ali Memon, Aman Tyagi,
David R Mortensen, and Kathleen M
Carley. Characterizing
sociolinguistic variation in the competing
vaccination communities.
            <xref ref-type="bibr" rid="ref14">arXiv preprint
arXiv:2006</xref>
            .04334, 2020.
          </p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>collection of covid-19 tweets with automated annotations</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Siddique</given-names>
            <surname>Latif</surname>
          </string-name>
          , Muhammad Usman, Sanaullah Manzoor, Waleed Iqbal, Junaid Qadir, Gareth Tyson, Ignacio Castro, Adeel Razi,
          <string-name>
            <surname>Maged N Kamel Boulos</surname>
            ,
            <given-names>Adrian</given-names>
          </string-name>
          <string-name>
            <surname>Weller</surname>
          </string-name>
          , et al.
          <article-title>Leveraging data science to combat covid-19: A comprehensive review</article-title>
          .
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Carley.</surname>
          </string-name>
          Cmu-miscov19:
          <article-title>A novel twitter dataset for characterizing covid-19 misinformation</article-title>
          ,
          <year>Sep 2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Catherine</given-names>
            <surname>Ordun</surname>
          </string-name>
          , Sanjay Purushotham, and
          <string-name>
            <given-names>Edward</given-names>
            <surname>Ra</surname>
          </string-name>
          .
          <article-title>Exploratory analysis of covid-19 tweets using topic modeling, umap, and digraphs</article-title>
          . arXiv preprint arXiv:
          <year>2005</year>
          .03082,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>James W Pennebaker</surname>
          </string-name>
          , Ryan L Boyd,
          <string-name>
            <surname>Kayla Jordan</surname>
            ,
            <given-names>and Kate</given-names>
          </string-name>
          <string-name>
            <surname>Blackburn</surname>
          </string-name>
          .
          <article-title>The development and psychometric properties of liwc2015</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>James W Pennebaker.</surname>
          </string-name>
          <article-title>The secret life of pronouns</article-title>
          . New Scientist,
          <volume>211</volume>
          (
          <issue>2828</issue>
          ):
          <volume>42</volume>
          {
          <fpage>45</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Umair</given-names>
            <surname>Qazi</surname>
          </string-name>
          , Muhammad Imran, and
          <article-title>Ferda O i. Geocov19: a dataset of hundreds of millions of multilingual covid19 tweets with location information</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>SIGSPATIAL Special</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):6{
          <fpage>15</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Junaid</given-names>
            <surname>Shuja</surname>
          </string-name>
          , Eisa Alanazi, Waleed Alasmary, and Abdulaziz Alashaikh.
          <article-title>Covid19 open source data sets: A comprehensive survey</article-title>
          . medRxiv,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Gautam</given-names>
            <surname>Kishore</surname>
          </string-name>
          <string-name>
            <surname>Shahi</surname>
          </string-name>
          , Anne Dirkson, and
          <article-title>Tim A Majchrzak. An exploratory study of covid-19 misinformation on twitter</article-title>
          .
          <source>arXiv preprint arXiv:2005.05710</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Gautam</given-names>
            <surname>Kishore</surname>
          </string-name>
          Shahi and
          <string-name>
            <given-names>Durgesh</given-names>
            <surname>Nandini</surname>
          </string-name>
          .
          <article-title>Fakecovid{a multilingual crossdomain fact check news dataset for covid19</article-title>
          .
          <source>arXiv preprint arXiv:2006.11343</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[LUM+20] [MC20] [OPR20] [PBJB15] [Pen11] [QIO20] [SAAA20] [SDM20] [SN20] [SPJ+20] [SSM+20] [TLC15] Angeline Sangalang</source>
          , Yotam Ophir, and
          <string-name>
            <surname>Joseph N Cappella.</surname>
          </string-name>
          <article-title>The potential for narrative correctives to combat misinformation</article-title>
          .
          <source>Journal of communication</source>
          ,
          <volume>69</volume>
          (
          <issue>3</issue>
          ):
          <volume>298</volume>
          {
          <fpage>319</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Xingyi</given-names>
            <surname>Song</surname>
          </string-name>
          , Johann Petrak, Ye Jiang, Iknoor Singh,
          <string-name>
            <given-names>Diana</given-names>
            <surname>Maynard</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Kalina</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <article-title>Classi cation aware neural topic model and its application on a new covid-19 disinformation corpus</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          arXiv preprint arXiv:
          <year>2006</year>
          .03354,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Karishma</given-names>
            <surname>Sharma</surname>
          </string-name>
          , Sungyong Seo, Chuizheng Meng, Sirisha Rambhatla, and Yan Liu.
          <article-title>Covid-19 on social media: Analyzing misinformation in twitter conversations</article-title>
          .
          <source>arXiv preprint arXiv:2003.12309</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>Journal of Communication</source>
          ,
          <volume>65</volume>
          (
          <issue>4</issue>
          ):
          <volume>674</volume>
          {
          <fpage>698</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [VCKBC20]
          <string-name>
            <given-names>Ramon</given-names>
            <surname>Villa-Cox</surname>
          </string-name>
          , Sumeet Kumar, Matthew Babcock, and
          <string-name>
            <surname>Kathleen</surname>
            <given-names>M Carley.</given-names>
          </string-name>
          <article-title>Stance in replies and quotes (srq): A new dataset for learning stance in twitter conversations</article-title>
          .
          <source>arXiv preprint arXiv:2006.00691</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Ves17] [YTLM20] [ZMFZ20]
          <string-name>
            <given-names>Marcela</given-names>
            <surname>Veselkova</surname>
          </string-name>
          .
          <article-title>Narrative policy framework: Narratives as heuristics in the policy process</article-title>
          .
          <source>Human A airs</source>
          ,
          <volume>27</volume>
          (
          <issue>2</issue>
          ):
          <fpage>178</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Kai-Cheng</surname>
            <given-names>Yang</given-names>
          </string-name>
          , Christopher TorresLugo, and Filippo Menczer.
          <article-title>Prevalence of low-credibility information on twitter during the covid-19 outbreak</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .14484,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Xinyi</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Apurva Mulay, Emilio Ferrara, and
          <string-name>
            <given-names>Reza</given-names>
            <surname>Zafarani</surname>
          </string-name>
          .
          <article-title>Recovery: A multimodal repository for covid-19 news credibility research</article-title>
          . arXiv preprint arXiv:
          <year>2006</year>
          .05557,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>