<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting Bot Behaviour in Social Media using Digital DNA Compression</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nivranshu Pasricha</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Conor Hayes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics Data Science Institute National University of Ireland Galway</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>A major challenge faced by online social networks such as Facebook and Twitter is the remarkable rise of fake and automated bot accounts over the last few years. Some of these accounts have been reported to engage in undesirable activities such as spamming, political campaigning and spreading falsehood on the platform. We present an approach to detect bot-like behaviour among Twitter accounts by analyzing their past tweeting activity. We build upon an existing technique of analysis of Twitter accounts called Digital DNA. Digital DNA models the behaviour of Twitter accounts by encoding the post history of a user account as a sequence of characters analogous to an actual DNA sequence. In our approach, we employ a lossless compression algorithm on these Digital DNA sequences and use the compression statistics as a measure of predictability in the behaviour of a group of Twitter accounts. We leverage the information conveyed by the compression statistics to visually represent the posting behaviour by a simple two dimensional scatter plot and categorize the user accounts as bots and genuine users by using an o -the-shelf implementation of the logistic regression classication algorithm.</p>
      </abstract>
      <kwd-group>
        <kwd>Social Media</kwd>
        <kwd>Twitter</kwd>
        <kwd>Online Social Networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Twitter is a popular online social network with 139 million average daily active
users (DAU) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. It has become one of the largest sources of news in the world
over the last few years. Twitter was launched in 2006 with the idea of using
an SMS service to send messages to other people in a group. Registered users
can post short messages of up to 280 characters (increased from 140 characters
in 2017) called tweets on the platform through di erent means including the
Twitter website, native smartphone &amp; desktop applications and other
thirdparty client applications.
      </p>
      <p>Twitter's API (Application Programming Interface) allows programmers and
developers to create services and applications which can be programmed to read
tweets from other users or automatically post tweets to a user's feed. This makes</p>
      <p>
        Twitter a breeding ground for a variety of automated software agents, also called
bots, which can interact with other users on the platform. These bots are
primarily used for harmless activities such as @year progess1, a bot account that
posts passage of time in an year as a progress bar and @MuseumBot2, an account
that posts random images from the Metropolitan Museum of Art. Some bot
accounts such as @EarthquakesSF3 spread helpful content in the time of disasters
like earthquakes and another bot account @tinycarebot4 posts reminders for
self care actions to its followers. However, there have been instances in the recent
past where bot accounts have been accused of malpractices such as spamming [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],
spreading misinformation [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], amplifying misconceptions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and manipulating
political discussions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Detecting automated bot accounts or bot-like behaviour on Twitter has
become a topic of interest among academics and researchers over the last few years.
Since this task includes analyzing thousands of accounts and millions of tweets,
it is unfeasible for humans to do this quickly and e ciently. The task of bot
detection can be modelled as a supervised learning task where Twitter accounts
can be classi ed into one of two categories of bots and genuine users. A
supervised machine learning or a deep learning classi cation algorithm can be used to
train a model on a labelled dataset with appropriate features such as user pro le
attributes and past tweeting activity in order to identify accounts as genuine
users or bots with high accuracy.</p>
      <p>
        The main contributions of this paper are as follows:
{ To classify Twitter accounts as bots or genuine users, we start with the
assumption that the long-term behaviour of a bot account on Twitter is less
random than the behaviour of a genuine user, or in other words, the
behaviour of a bot account is more predictable. We measure this randomness
and unpredictability in the posting behaviour of a Twitter account by
employing a lossless compression algorithm on the Digital DNA sequence of the
account. The technique of Digital DNA was designed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to encode the
posting history of a Twitter account as a sequence of characters, just like
the actual human DNA sequence.
{ We use this string compression approach on a group of Twitter accounts and
obtain compression statistics, namely the size of the uncompressed DNA
sequence, the size of the compressed DNA sequence and the value of the
compression ratio. These compression statistics are then used as a measure
of randomness and predictability in the behaviour of the accounts under
study.
{ Based on the compression statistics, we present a 2-dimensional scatterplot
as a visual representation of the predictability in the behaviour of a group
of Twitter accounts.
1 https://twitter.com/year_progress
2 https://twitter.com/MuseumBot
3 https://twitter.com/EarthquakesSF
4 https://twitter.com/tinycarebot
{ We then use the compression statistics as features in a supervised machine
learning algorithm and learn to classify Twitter accounts as bots and genuine
users.
{ Finally, we evaluate the machine learning model using various scoring metrics
and compare them with existing state-of-the-art techniques of bot detection.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The issue of dealing with bot and automated accounts on Twitter has become a
hot topic in recent years. Various studies and experiments have been carried out
to automatically identify bot accounts and study their impact on other users of
the online platform. All the existing bot detection techniques in the literature
look to identify a combination of features related to user pro le, tweet/retweet
content and past tweeting activity of the user in order to look for patterns that
can be used to distinguish bot accounts from genuine users on the platform.</p>
      <p>
        A framework to identify social bots on Twitter was presented in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] which
deployed more than one thousand features to train a classi er on a dataset
containing Twitter bots and human users. The features used for the task of
classi cation were categorized into one of six categories of user meta-data features
(number of friends and followers, counts of tweets, retweets and replies, pro le
description etc.), friends features, network features, temporal features (number
of tweets posted in a xed time interval and distribution of time intervals
between posts), sentiment features and content features (statistics about the length
and the entropy of the text of the tweet). The sentiment and content features
extracted from the tweets were found to be important features along with the
statistical properties of retweet networks for the tweets posted by the accounts.
      </p>
      <p>
        Another approach proposed to identify bot accounts on Twitter divided a
dataset of Twitter accounts into 4 distinct bands (1K, 100K, 1M and 10M)
depending upon the number of followers of an account [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A Random Forest
classi er was then trained using a set of 21 features to classify the accounts
either as a human or a bot. General features related to an account such as the
age of the account, the count of tweets, retweets, replies and favourites, the
tweeting frequency, friend-to-follower ratio and source of Twitter activity were
used along with a few novel features such as the count of URLs per tweet, the
type of device used for tweeting and the size of the media content posted by an
account.
      </p>
      <p>
        Instead of categorizing Twitter accounts into two groups of humans and bots,
the authors in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed a system to classify the accounts to one of three
categories - human, bot and cyborg. A cyborg account here referred to either a
bot-assisted human or a human-assisted bot. This system consisted of four major
components - an entropy component that used tweeting interval as a measure
for behavioural complexity of the account, a spam detection component that
looked for patterns in the content of the tweets posted by an account, an account
properties component to study the account properties like device information,
and a decision maker which used a combination of the features generated by
the other components to nally assign a label to an account. The authors also
created a metric called account reputation, which was de ned as the normalized
ratio for the number of accounts followed by a Twitter account and the number
of accounts befriended by that account. Human accounts were found to have
highest values of account reputation, closely followed by cyborg accounts and bot
accounts had much lower values, typically less than 0.5. In terms of frequency of
posting tweets, bots were found to exhibit burstiness, i.e., posting a number of
tweets in a small interval of time interspersed with long periods of hibernation
whereas tweets posted by human accounts had large intervals (typically hours
or days) between successive tweets. Bot accounts also exhibited regular posting
behaviour while human accounts displayed more complex posting behaviour and
higher entropy.
      </p>
      <p>
        Most of these approaches have been trained to classify accounts using a
supervised learning algorithm, however some unsupervised approaches have also
been studied for this task. A methodology to identify spam campaigns on Twitter
and Facebook modelled a set of user pro les using a weighted social graph where
users were inter-linked based on values of a prede ned set of 14 features [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
social graph was then clustered using Markov clustering into groups of users that
exhibit similar behaviour and activities.
      </p>
      <p>
        It has been argued that a paradigm-shift [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] has taken place over the last few
years in the behaviour of Twitter bots where more sophisticated social bots have
been identi ed that have been able to fool traditional bot detection techniques
as well as human annotators. These novel social spambots appear to behave
as genuine users when looked at individually, however they exhibit patterns of
similar activity when considered as part of a group. A methodology titled \Social
Fingerprinting" was proprosed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to identify such sophisticated bots not at
individual level but rather as groups of accounts that exhibit similar behaviour.
The technique which is inspired by bioinformatics encoded the behaviour of an
account as a sequence of bases, represented using a prede ned set of alphabet
of nite cardinality consisting of letters such as A, C, G and T. This sequence is
said to be the digital DNA sequence for an account. One such encoding scheme
assigns the type of post (tweet, reply or retweet) made on Twitter to a unique
character. The tweets, replies and retweets posted by an account are mapped
to characters A, C and T respectively to generate the digital DNA string for the
account.
      </p>
      <p>Bt3ype =
8 A
&lt; C
: T
tweet; 9
reply; =
retweet ;</p>
      <p>
        An approach similar to digital DNA was proposed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to study the temporal
evolution of users in online forums. A feature space was de ned to visually
represent chronological user events (posts and replies in forum threads) as paths
and to model the distribution of inter-event times in order to de ne a clustering
of various online forums.
      </p>
      <p>
        The authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] de ne the length of the longest common substring (LCS)
or the length of the k-common substring in the digital DNA sequences as a
measure of similarity for the two categories of bot and genuine user accounts
on Twitter. For a group of accounts consisting of genuine users, the length of
the LCS of the digital DNA strings was found to be very short. However, for a
group of bot accounts, the length of the LCS was much longer. This idea was
then used to identify bot accounts in a mixed group of user accounts which
contained accounts for both genuine users and bots by nding the k-common
substring for all those accounts and then nding a cut-o , the number of accounts
which share a long common substring and classifying all such accounts which
share that substring as bots and other accounts with shorter common substrings
were classi ed as genuine human users. Based on this idea, two techniques, one
supervised and another unsupervised were further discussed to nd of groups of
similarly-behaved accounts among a mixed set of accounts of automated bots
and genuine human users, both of which showed promising results and were
found to be better than the existing state-of-the-art techniques for the task of
bot detection.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Dataset</title>
      <p>
        We use the dataset5 created in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to analyze our approach for detecting bot
accounts on Twitter. This dataset contains a random sample of genuine human
user accounts on Twitter along with a variety of bot accounts. The data was
collected over a period of a few months in 2014 and contains pro le information
for 11017 accounts with more than 6 million posts (tweets, retweets and replies)
made by those accounts. The bot accounts in this dataset are broadly
categorized into three categories, namely, social spambots, traditional spambots and
fake followers. Table 1 contains the description of the accounts across the various
categories. The authors in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] created two test sets, Mixed1 and Mixed2 from this
dataset to study their approach of Digital DNA and Social Fingerprinting for
identifying bot accounts on Twitter. Mixed1 contains a random sample of
genuine user accounts and a subset a social spambot accounts. These bot accounts
were found to be created to target Mayoral elections in Rome, Italy. Around
1000 automated accounts were found to be created by a social media marketing
rm to promote and publicize the campaign and policies for one of the
candidates. Mixed2 also contains a random sample of genuine user accounts and a
di erent subset of social spambot accounts. The bot accounts in Mixed2 were
found to be created primarily for promotion and advertisement of products on
the popular E-commerce platform amazon.com. The bot accounts categorized as
social spambots in Table 1 exhibited properties which di ered vastly from other
known traditional bot accounts. These accounts appeared much more
sophisticated than traditional bot accounts and appeared to be operated by genuine
human users. The pro les of these social spambot accounts displayed detailed
information such as pro le pictures, bio, location etc., although most of it was
found to be either fake or stolen from other accounts. We evaluate our approach
of bot detection using DNA compression on both Mixed1 and Mixed2 test sets.
5 http://mib.projects.iit.cnr.it/dataset.html
In this section, we describe the methodology used to model past activity and
behaviour of Twitter users as a DNA sequence. We also explain our
assumptions and reasoning behind the idea of using string compression on these DNA
sequences. Furthermore, we describe our strategy to visually represent genuine
user accounts and bot accounts based on the compression statistics and how
we use this information to classify the accounts using a simple classi cation
algorithm like logistic regression.
We start with the assumption that the behaviour of a bot account is more
predictable and less random than the behaviour of a genuine human-operated
account. To measure the predictability in the behaviour, we make use of digital
DNA sequences. The digital DNA sequence for a user account can be created
using an alphabet set of cardinality 3 with symbols A, C and T corresponding to a
tweet, a retweet and a reply respectively. This provides a way of representing an
account's posting history by encoding the information of past posting activity
in a compact and concise fashion. The DNA sequence can be generated for an
account by starting with an empty string and scanning through the post history
of the account in chronological order, appending the relevant alphabet to the
string based on the type of post made by the account. For example, a series
of 10 consecutive posts made by an account can be represented by the vector,
s = hA; C; T; C; A; T; T; T; T; Ai or alternatively, by the string s = \ACTCATTTTA".
      </p>
      <p>
        In Information Theory, the concept of entropy is used as a measure of
randomness in a data signal. It is well known that data with high entropy cannot
be compressed and decompressed with high e ciency using a lossless
compression technique [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. We compress the digital DNA sequences with a lossless
compression algorithm and obtain compression statistics such as the size of the
compressed string and the compression ratio which serve as metrics to measure
the predictability of an account's behaviour.
      </p>
      <p>The digital DNA strings are compressed using the zlib compression library
in the Python programming language with default parameters. It is essential to
convert a string of ASCII characters into a bytes object in memory before calling
the zlib.compress() compression function. The sys module in Python o ers a
built-in function getsizeof() which returns the size in bytes of an object stored
in memory. The value of the compression ratio for a string is then computed as
the ratio of the size (in bytes) of the uncompressed bytes object and the size (in
bytes) of the compressed bytes object in memory.</p>
      <p>Table 2 shows the mean and standard deviation for the values of the
compression statistics for the two groups of bots and genuine users as well as the
combined statistics for all accounts in Mixed1 and Mixed2 datasets. The average
value of the compression ratio for the DNA sequences of bot accounts across
the two datasets is 10:53, while for the DNA sequences corresponding to the
accounts of genuine users, the mean value of compression ratio is 4:43. This is
in agreement with our assumption that the behaviour of a bot account is more
predictable and less random compared to the behaviour of a genuine human
user, hence the higher value of compression ratio.
4.2</p>
      <sec id="sec-3-1">
        <title>Visualizing Compression Statistics</title>
        <p>Our approach of using string compression on Digital DNA strings presents us
with a way to easily visualize the behaviour for a group of accounts using a simple
two dimensional scatterplot, as shown in Figure 1. Each point in the plot
represents an account from Mixed1 and Mixed2 with the X-coordinate representing
the size of the original DNA string in bytes and the Y-coordinate representing
the value of the compression ratio.
4.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Classi cation with Logistic Regression</title>
        <p>It can be seen in Figure 1 that a line of separation exists between the bot accounts
and genuine human-operated accounts in the dataset under study. The value of
compression ratio is typically smaller than 10 for genuine users while bots appear
to have much higher value, especially for longer DNA sequences. This implies
that a linear classi cation algorithm can be trained on some labelled data to
estimate the line separating the accounts into two classes of bot and genuine
user accounts. We split both Mixed1 and Mixed2 in the ratio of 50:50, where
50
40
o
it
a
R
in30
o
s
rs
e
p
m
o
C20
10
0
0
500
1000</p>
        <p>
          1500 2000
Original DNA Size
2500
3000
we use 50% of the accounts in the dataset for training and the remaining 50%
for evaluation. We use the sklearn [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] implementation of logistic regression for
binary classi cation with the default parameters to train two di erent models,
each with two features. We keep the original DNA size as one of the features
in both classi ers and use the size of the compressed DNA as the other feature
in one of the models and use the compression ratio as the second feature in the
other classi cation model.6
5
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>
        We compare the results obtained with our approach to the results in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] on
the same test sets and using the same evaluation metrics - Accuracy, Precision,
Recall, F-Measure (F1), Matthews Correlation Coe cient (MCC) and Speci city
(True Negative Rate). We randomly pick 50% of the accounts from Mixed1
and Mixed2, which we use to train a binary logistic regression classi er with
compression statistics as the features and test on the remaining 50% of the
accounts in the dataset. We repeat this 1000 times and report the average value
for each evaluation metric in Table 3.
      </p>
      <p>Our simple approach to model account behaviour using a compression
algorithm on the digital DNA sequence allows us to accurately distinguish between
bot accounts and genuine users. For the Mixed1 test set, the classi er which uses
6 Our code is available at https://github.com/pasricha/bot-dna-compression.
the compressed DNA size as one of the features, slightly outperforms both the
supervised and the unsupervised versions of the k-common substring technique
with respect to accuracy, recall, F1 and MCC scores. The other classi er which
uses compression ratio and the size of the original uncompressed DNA as the
two features, is signi cantly better than the k-common substring technique for
all metrics except recall. In the case of the Mixed2 dataset, the classi er with
the size of the compressed DNA performs slightly worse than the k-common
substring techinque. However, the other classi er is better in terms of 3 of the
6 evaluation metrics, namely, accuracy, F-Measure and MCC. In terms of
complexity and scalability, our approach provides a fast and simple way for assessing
the predictability in the behaviour of Twitter accounts because compression of
a string can be performed in linear time and constant memory overhead.
(b) Mixed2
Fig. 2: Results with di erent values for maximum sequence length L.</p>
      <p>Evaluation Metrics</p>
      <p>Accuracy
F1Score
MCC
Precision
Recall
Specificity</p>
      <p>Evaluation Metrics</p>
      <p>Accuracy
F1Score
MCC
Precision
Recall
Specificity</p>
      <p>We perform additional experiments to study the e ect of the length of the
DNA sequence on the nal result. For the accounts in the dataset under study,
the average length of the DNA sequence is 1050:29 with the longest sequence of
length 3250. These numbers might not be representative of the real world
situation. On the one hand, newer accounts may only have a few dozen posts, on the
other hand, accounts that have been operational for many years may have
thousands of posts. We set di erent limits L = f10; 25; 50; 100; 200; 500; 1000; 2000g
for the maximum length of the DNA sequence and for each account in the
dataset, we pick a sub-sequence of random length between 1 and L. Figure 2
shows the value for all evaluation metrics, averaged over 100 executions, with
di erent values of L when using the original DNA size and the compression ratio
as the two features for training the logistic regression classi er. For short DNA
sequences, say L 25, the model performs poorly in classifying the Twitter
accounts and we see low scores for all evaluation metrics with MCC score as
low as 0.5. However, there are great improvements in the performance as the
DNA sequence length increases. We believe this is due to the fact that a longer
DNA sequence potentially allows the compression algorithm to look for longer
repeating patterns that can be utilised for e cient compression, hence leading
to higher compression ratio for the DNA sequences of bot accounts, compared
to the compression ratio for the DNA sequences of the genuine user accounts.</p>
      <p>We also analyze the e ectiveness of our approach against evading techniques
by performing an experiment where we apply random permutations to the digital
DNA sequences. We again pick 50% of the accounts in the dataset for training
and the keep the remaining 50% for evaluation. However, before evaluating the
classi er, we apply random permutations to the DNA sequences in the test
set. We repeat this experiment 100 times and present the average value for the
evaluation metrics in Table 4. We notice that the string compression technique
still performs well and there is only a slight decrease in the performance with
respect to the evaluation metrics.</p>
      <p>
        In both Mixed1 and Mixed2, about 40-50% accounts are labelled as bots.
However, the actual proportion of bots on Twitter is expected to be much smaller
than this. According to a report submitted by Twitter to the United States
Securities and Exchange Commission [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], less than 5% of monthly active users
(MAU) are bots. It has been estimated in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] that bots make up about 9-15% of
accounts on Twitter. We address this issue by keeping a xed number of genuine
1.0
0.9
0.8
0.7
0.6
CC0.5
M
0.4
0.3
0.2
0.1 FSeizaetuorfeCompressedDNA
      </p>
      <p>CompressionRatio
0.00.00 0.02 0.04 P0.r0o6por0t.i0o8n of0B.1o0ts i0n.1t2he D0a.1t4aset0.16 0.18 0.20
(a) Mixed1
(b) Mixed2
user accounts in the test set and vary the proportion of bot accounts from 1%
to 20%. Figure 3 shows the MCC score evaluated over the di erent proportions
of bot accounts in both Mixed1 and Mixed2 datasets averaged over 100 di erent
executions. Our approach appears to be reliable even with unbalanced test sets
and performs slightly worse only when the test set is extremely unbalanced with
bot accounts making up only 1-2% of the accounts.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>In this paper we presented an approach to detect bot accounts on the popular
online social network Twitter by building on top of the already existing
technique of Digital DNA. We extended the idea of using a digital DNA sequence to
model the past activity of Twitter account by employing the technique of string
compression on such a sequence. By leveraging the compression statistics, we
were able to accurately model and visually represent the behaviour of Twitter
accounts. This approach of modelling account behaviour with digital DNA is
fast and can be scaled to handle large number of accounts with the advantage
of being language and content independent. Our experiments suggest that this
technique is also robust to potential bot evading techniques and can work with
unbalanced datasets too. A future path of work can look at incorporating other
user-pro le and content-based features into the digital DNA sequence and apply
this technique to identify automated accounts on other social media platforms
such as Facebook and Reddit.</p>
      <p>Acknowledgement. This publication has emanated from research supported
in part by a research grant from Science Foundation Ireland (SFI) under Grant
Number SFI/12/RC/2289P2.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abulaish</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A generic statistical approach for spam detection in online social networks</article-title>
          .
          <source>Computer Communications</source>
          <volume>36</volume>
          (
          <fpage>10</fpage>
          -
          <lpage>11</lpage>
          ),
          <volume>1120</volume>
          {
          <fpage>1129</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bessi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
          </string-name>
          , E.:
          <article-title>Social bots distort the 2016 u.s. presidential election online discussion</article-title>
          .
          <source>First Monday</source>
          <volume>21</volume>
          (
          <issue>11</issue>
          ) (
          <year>2016</year>
          ), https://firstmonday.org/ojs/index. php/fm/article/view/7090
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Broniatowski</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jamison</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>AlKulaib</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benton</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quinn</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Weaponized health communication: Twitter bots and russian trolls amplify the vaccine debate</article-title>
          .
          <source>American Journal of Public Health</source>
          <volume>108</volume>
          (
          <issue>10</issue>
          ),
          <volume>1378</volume>
          {
          <fpage>1384</fpage>
          (
          <year>2018</year>
          ), https://doi.org/10.2105/AJPH.
          <year>2018</year>
          .304567
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gianvecchio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jajodia</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Detecting automation of twitter accounts: Are you a human, bot, or cyborg</article-title>
          ?
          <source>IEEE Transactions on Dependable and Secure Computing</source>
          <volume>9</volume>
          (
          <issue>6</issue>
          ),
          <volume>811</volume>
          {
          <fpage>824</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cresci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pietro</surname>
            ,
            <given-names>R.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrocchi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spognardi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tesconi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Social ngerprinting: Detection of spambot groups through dna-inspired behavioral modeling</article-title>
          .
          <source>IEEE Transactions on Dependable and Secure Computing</source>
          <volume>15</volume>
          (
          <issue>4</issue>
          ),
          <volume>561</volume>
          {
          <fpage>576</fpage>
          (
          <year>July 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cresci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di</surname>
            <given-names>Pietro</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Petrocchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Spognardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Tesconi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>The paradigmshift of social spambots: Evidence, theories, and tools for the arms race</article-title>
          .
          <source>In: Proceedings of the 26th International Conference on World Wide Web Companion</source>
          . pp.
          <volume>963</volume>
          {
          <fpage>972</fpage>
          . WWW '17 Companion, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, Switzerland (
          <year>2017</year>
          ), https://doi.org/10.1145/3041021.3055135
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gilani</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochmar</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crowcroft</surname>
          </string-name>
          , J.:
          <article-title>Classi cation of twitter accounts into automated agents and human users</article-title>
          .
          <source>In: Proceedings of the 2017 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining</source>
          <year>2017</year>
          . pp.
          <volume>489</volume>
          {
          <fpage>496</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hogan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bailey</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leckie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A time decoupling approach for studying forum dynamics</article-title>
          .
          <source>World Wide Web</source>
          <volume>16</volume>
          (
          <issue>5</issue>
          ),
          <volume>595</volume>
          {620 (Nov
          <year>2013</year>
          ), https://doi.org/10.1007/s11280-012-0169-1
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eo</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caverlee</surname>
          </string-name>
          , J.:
          <article-title>Seven months with the devils: A long-term study of content polluters on twitter</article-title>
          .
          <source>In: Fifth International AAAI Conference on Weblogs and Social Media</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>MacKay</surname>
            , D.J.,
            <given-names>Mac</given-names>
          </string-name>
          <string-name>
            <surname>Kay</surname>
            ,
            <given-names>D.J.:</given-names>
          </string-name>
          <article-title>Information theory, inference and learning algorithms</article-title>
          . Cambridge university press (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <volume>2825</volume>
          {
          <fpage>2830</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Shao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciampaglia</surname>
            ,
            <given-names>G.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varol</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>K.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The spread of low-credibility content by social bots</article-title>
          .
          <source>Nature communications 9(1)</source>
          ,
          <volume>4787</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Twitter:
          <article-title>Quarterly report pursuant to section 13 or 15(d) of the securities exchange act of 1934 (</article-title>
          <year>2014</year>
          ), https://www.sec.gov/Archives/edgar/data/ 1418091/000156459014003474/twtr-10q_
          <fpage>20140630</fpage>
          .htm
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. Twitter:
          <article-title>Quarterly results (</article-title>
          <year>Jul 2019</year>
          ), https://investor.twitterinc.com/ financial-information/quarterly-results/
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Varol</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menczer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flammini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Online humanbot interactions: Detection, estimation, and characterization</article-title>
          .
          <source>In: Eleventh international AAAI conference on web and social media</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>