<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identification Of Bot Accounts In Twitter Using 2D CNNs On User-generated Contents</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Marco Polignano, Marco Giuseppe de Pinto</institution>
          ,
          <addr-line>Pasquale Lops, and Giovanni Semeraro</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Bari ALDO MORO via E. Orabona 4</institution>
          ,
          <addr-line>70125, Bari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>The number of accounts that autonomously publish contents on the web is growing fast, and it is very common to encounter them, especially on social networks. They are mostly used to post ads, false information, and scams that a user might run into. Such an account is called bot, an abbreviation of robot (a.k.a. social bots, or sybil accounts). In order to support the end user in deciding where a social network post comes from, bot or a real user, it is essential to automatically identify these accounts accurately and notify the end user in time. In this work, we present a model of classification of social network accounts in humans or bots starting from a set of one hundred textual contents that the account has published, in particular on Twitter platform. When an account of a real user has been identified, we performed an additional step of classification to carry out its gender. The model was realized through a combination of convolutional and dense neural networks on textual data represented by word embedding vectors. Our architecture was trained and evaluated on the data made available by the PAN Bots and Gender Profiling challenge at CLEF 2019, which provided annotated data in both English and Spanish. Considered as the evaluation metric the accuracy of the system, we obtained a score of 0.9182 for the classification Bot vs. Humans, 0.7973 for Male vs. Female on the English language. Concerning the Spanish language, similar results were obtained. A score of 0.9156 for the classification Bot vs. Humans, 0.7417 for Male vs. Female, has been earned. We consider these results encouraging, and this allows us to propose our model as a good starting point for future researches about the topic when no other descriptive details about the account are available. In order to support future development and the replicability of results, the source code of the proposed model is available on the following GitHub repository: https://github.com/marcopoli/Identification-ofTwitter-bots-using-CNN Copyright c 2019 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). CLEF 2019, 9-12 September 2019, Lugano, Switzerland.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A bot can be considered as an automatic system capable of creating contents on the web
independently, i.e., without a human being involved. They are often used as link
aggregators or as repositories of copyrighted contents such as movies, songs or images. In the
past, they have also been used to alter public opinion about famous people through the
inclusion of highly targeted comments, especially during political elections [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In other
cases, they have been used for manipulating the stock market or to spread fake
theories. Each bot is different and it can complete different tasks. Some are more complex
and produce contents very similar to those of a human being. Others simply retweet
or answer simple questions. Bots that provide a legal and potentially useful service are
also present on the web, but before using them, it is always necessary to be aware that
you do not have a human being on the other side and that such data could be reused for
other purposes. A recent analysis by the Pew Research Center 1 shows that automated
systems generate two-thirds of the links shared on Twitter and their usage by real users
is continuously growing due to their extreme efficiency and availability. If on the one
hand, these systems provide the real user with a service of content sharing, on the other
hand they represent a severe risk to the privacy of the user. It is clear that any interaction
with them is stored and after a sufficient number of iterations is sufficient to profile the
user in order to mislead him with fake advertising, fake websites, phishing, and other
malicious activities.
      </p>
      <p>The identification of bots on social networks is a challenge that, in recent years, has
involved many researchers in creating an accurate classification system with real-time
response times. A bot detection system can, therefore, help to maintain the stability of
the network and ensure the safety of users. State of the art systems have relied on the
use of numerous features of the user profile of the accused person such as the number of
followers, the number of tweets and retweets, the length of the name, the age and so on.
In our solution, we use only a set of one hundred tweets per user without knowing any
information about the account from which they come. This situation can make the task
of classification much more complex and utterly different from what is already present
in the literature. This choice, which could be considered as limiting, presents a broader
application scenario. In this way, it is possible to identify bots even in contexts where
little information is available about the account, such as in possible situations of privacy,
a topic that has become increasingly important in recent years. Our solution, therefore,
wants to be able to work correctly with the least amount of information possible, trying
to focus on what are the particularities of the writing style of an automatic system and
a real user. In this regard, it may, therefore, be interesting to learn more about any
differences in the style of writing used by men and women. Consequently, we wanted
to deepen the theme, trying to carry out a further step of classification male vs. female
in case the system was able to classify the account as human. This task, known in the
literature as author profiling, has been a factor of scientific interest for many years,
demonstrating how this information can be effectively identified simply by analyzing
the content produced on social media.</p>
      <p>
        The rest of the article is organized as follow: we start with an overview of the state of
the art techniques used to address the classification task in Sec. 2. In Sec. 3 we describe
and detail the model used in the challenge PAN Bots and Gender Profiling challenge
at CLEF 2019 [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], while in Section 4, we discuss the experimental analysis carried
out on the available training and validation dataset. In Sec. 5, conclusions and future
      </p>
    </sec>
    <sec id="sec-2">
      <title>1 https://www.pewinternet.org/2018/04/09/bots-in-the-twittersphere/</title>
      <p>developments close the article inviting to download the available code of the proposed
model in order to allow everyone to continue its implementation for future studies.
2</p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>
          The research area concerning authorship attribution to short texts is the one that well
encloses the goal of our model: the classification of social media accounts into a bot
or a human being using the user-generated textual contents. In this field of research,
many classification strategies have been presented in the last years. One of the first
systems of the interest of community was the one proposed by Lee [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. It was based
on the concept of honeypots defined as fake websites used as traps to identify the main
characteristics of online spammers. The collected data were used at a later stage of
classification through standard machine learning algorithms such as logistic regression and
J48 to successfully detect them. Approaches based on the use of lexical features such as
characters, words, and n-grams, part of speech (pos) tags, number of punctuations and
more, have been proposed for numerous tasks of authorship attribution [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] as an
example for author verification [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], plagiarism detection [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], author profiling or
characterization [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. However, approaches based on the manual definition of features engineering,
as the previous, are often high time consuming and inaccurate. As a consequence, it has
become always more common to identify approaches based on a representation of text
in a vectorial space automatically learned from a neural network [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the author exposes the concept of word embedding that can be summarized
as a "learned distributed feature vector to represent the similarity between words". This
concept has been exploited by Mikolov [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] through word2vec, a tool for
implementing work embeddings, but also by for Pennington [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] that proposed GloVe and by
Bojanowski [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] who implemented its n-gram based vector space named FastText. The
benefits of using these vectorial representations are about the property that every vector
has in this space. In particular, two words that are semantically related as an example
"sport" and "football" results very similarly in that space. This property supports the
encoding of a sentence in a new structure able to preserve semantic relations among
words. Approaches based on neural networks that use word embeddings as
representations of sentences have proved to be very promising and effective for text classification
tasks even in domains close to the authorship attribution [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. In 2017 Shrestha [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]
presents an authorship attribution model for short texts using convolutional neural
networks (CNNs)[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The architecture uses n-grams as input, suitably transformed into
semantic embeddings through a vector space of size 300 learned directly from the
training data using the word2vec approach. The results described by Shrestha in her work
show the good predictive ability of deep neural models, especially when a very large
amount of data is available for the training phase. Strategies based on the use of
recurrent neural networks (RNNs)[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] were used also by Kudugunta [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for the task of
tweet-level bot detection. The authors use GloVe for the encoding of the words
constituting the Tweets so that an RNN followed by more dense networks enriched by a
context vector could be applied to the data. Account-level Twitter bots account detection is,
instead, performed by more classical machine learning strategies such as Logistic
Regression and Random Forest Classifiers. The excellent results obtained by the authors
confirm again the usefulness of deep learning approaches for dealing with the task.
        </p>
        <p>
          It is not difficult to find in literature examples of approaches that use neural
networks even for the task of identifying the gender of the author of short texts, such as
those published on social media. The overview about the multimodal gender
identification challenge at PAN 2018 [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], well describe how the close totality of the submitted
approaches are based on CNNs, RNNs and, more in general, a combination of neural
networks. Following the wave of positive results obtained by such approaches based
on neural networks, in this work, we combined the contribution of word embeddings
with that of CNNs and Dense neural networks. In particular, unlike the works analyzed,
we decided to use word embeddings pre-trained on English and Spanish and 2D CNNs
able to work on the tensorial representation of content generated by each account. The
particularity of our model lies in the idea to work with one single model directly on
all the contents generated by the user considering them as a single learning example
of our net. Moreover, we were able to use the same network architecture for both the
two different classification tasks (bot vs. human, male vs. female) by varying only the
training data obtaining by the way good results.
3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>The Proposed Model</title>
        <p>The model of classification of Twitter accounts in bot or human and later in male or
female proposed in this work is based on a neural architecture that through
convolutional layers and dense layers is able to capture distant relationships among words. In
particular it focus not only on word relations in the same short text but also it is able
to detect relations among word in different pieces of text. The key idea behind the
proposed approach is that in order to exploit the stylistic information contained in a set
of contents produced by a user, it is necessary to work on them at the same time. It is
immediately clear that you can imagine the account of a user u 2 U as a set of contents
C = cu1; cu2; cu3; :::; cun aggregated in the form of a matrix. It is possible to place each
sentence as a row of the matrix and to tokenize them into words for creating the grid
structure. This allows us to draw a structured representation of the profile of the user
and to work on them with approaches, as CNNs, that are more commonly used for data
in a grid-like topology (i.e. images).</p>
        <p>
          Fig. 1 describes very generally the architecture of the model designed for facing the
bot identification task at user-level. In order to feed the network, we started with the
encoding of each user-generated content in a list of word embedding vectors. Starting
from a word embedding matrix S 2 IR e jV j, where e is the size of the word embedding
vectors and jV j the cardinality of the dictionary pre-trained, we encoded each of the first
50 words wjk that composed the j-th user tweet tuj 2 Tu as word embedding vector.
Tweets with a number of words higher than 50 have been truncated. Words not found
in the vector space are transformed into word embeddings through a random vector
selected from the entire dictionary, as proposed by Zhang [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. We obtained a matrix of
encoded user tweets Eu 2 IR m 50 e where m is the number of tweets considered for
the user, 50 is the fixed number of word of each tweet we decided to consider during
the encoding, e is the size of the word embedding vectors. We used Eu as input of our
architecture. Obviously, during the training phase, the matrix will already labeled with
the class value. During the forecast phase, instead, the class will be properly estimated.
The first three layers of our net are composed by 2D convolutional networks mediated
by three max-pooling 2 x 2 operations. These max-pooling operations have the purpose
of allowing the network to focus on low-level feature blocks and to map them in a
higher-level feature space, more descriptive than the previous, through the convolutional
operations. In this way, CNNs allow us to efficiently detect local relations shared among
words closed in positions in user’s tweets.
        </p>
        <p>
          The details about the number of parameters for each layer and its shape are
reported in Fig. 2. The figure allows us to observe the configuration of parameters used
in particular for the first three CNNs and max-pooling operations. Starting from a word
embedding matrix of size 100 x 50 x 300, we performed the first 2D CNN with a 5 x
5 kernel using a ReLu activation function [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. The rectified linear unit (ReLu) can be
formalized as g(z) = maxf0; zg and it allows to obtain a large value in case of
activation by applying this function as a good choice to represent hidden units. As commonly
performed, after a CNN able to perform convolutional operations on data, we applied a
max-pooling operation of size 2 x 2 transforming each squared portion of our previous
output as their max value. Pooling can, in this case, help the model to focus only on
principal characteristics of every portion of data making them invariant by their
position. We have run the process described twice more by simply changing the kernel size
to 5 x 4 and finally to 3 x 3. Following we have flatten the output to make it compatible
with the next layers.
        </p>
        <p>
          We applied three Dense layers (fully connected layers) with a tanh activation
function, varying the output size from 800 elements to 400, 200, and finally 100. This step
allows the network to put in relations all the intermediate results obtained by the
previous layers on the different sections of the input, discovering hidden correlations that
can contribute to the prediction. The tanh function defined as g(z) = 2 (2z) 1 has
an S-Shape and produces values among -1 and 1 making layer output more center to
the 0. Moreover, it produces a gradient larger than sigmoid function helping to speed
up the convergence [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Finally, another dense layer with a soft-max activation function
has been applied for estimating the probability distribution of each of the classes.
4
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Model Evaluation</title>
        <p>The evaluation of the proposed model has been carried out in order to be able to answer
two different research questions. First, we want to investigate whether the model
produces results of accuracy that are good enough to identify Twitter bot accounts and to
detect the gender of the user in case of human being. Secondly, we want to understand
if the choice of word embedding vector space influences the final performance of the
model. Such a result may produce interesting considerations to be used in future work
on the subject by providing a detailed starting point.
4.1</p>
        <sec id="sec-2-3-1">
          <title>Datasets, Baselines and Metrics</title>
          <p>
            The dataset used for the training and validation phase of the model is the one
provided by the PAN Bots and Gender Profiling challenge at CLEF 2019 [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ]. It has been
provided in two languages English and Spanish and it is composed by tweets labeled
with a bot or to a male/female human class. The detailed statistics about the dataset are
available in Tab. 1 . The test dataset has been not released yet.
          </p>
          <p>We used as validation metric the Accuracy, the same used in the PAN challenge. In
particular, for each language, individual accuracies are calculated. Firstly, we calculated
the accuracy of identifying bots vs. human. Then, in case of humans, we calculated the
accuracy of identifying males vs. females. Finally, the average of the accuracy values
per language are used for obtain the final score.</p>
          <p>The baseline considered for our task is the majority vote classifier that, as
consequence of the homogeneous classes distribution in the dataset has an accuracy score
equal to 0.5.
4.2</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Tweets Preprocessing and Data Enrichment</title>
          <p>The tweets available in the dataset were supplied without any cleaning or pre-processing.
In this regard, it was decided to keep as much information as possible within the model.
Each sentence is divided into tokens through the TweetTokenizer class of the NLTK
library. After that, we used Ekphrasis preprocessor library to annotate mentions, URL,
email, numbers, dates, amount of money, and make word spelling correction and
hashtag unpacking if necessary. As an example "14/01/2018" is translated into the word
"date" and the hashtag "#lovefootball" is translated into "love football". This process
allows to correctly apply the translation of words in word embedding without
excluding the previously mentioned elements.</p>
          <p>
            In order to obtain a more general and accurate model, we have decided to
enrich the reference dataset with further data regarding the contents produced by men
or women. We followed the strategy used in the GitHub project
Gender-Classificationusing-Twitter-Feeds 2 based on the work of Liu [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. For the English language, we
developed a module able to collect tweets localized in the USA through official Twitter
API. Messages collected between the 15/03/2019 to the 30/04/2019 have been filtered
on the base of the official first name written in the account description. In particular,
if the name was one of them reported by the USA census 3 as used for a male subject
we labeled it as a consequence. Similarly, we labeled female tweets4. As far as Spanish
is concerned, the strategy was symmetrical. First of all, we collected the geolocalized
tweets in Spain, then we filtered and categorized them following a list of male and
female names available on the web5.
          </p>
          <p>The previously filtered tweets were then aggregated into groups of 100, randomly
selecting them from the corresponding reference class. We obtained 1000 additional
examples of annotated accounts for each of the classes and for each of the two languages.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2 https://github.com/vjonnala/Gender-Classification-using-Twitter-Feeds 3 http://www2.census.gov/topics/genealogy/1990surnames/dist.male.first 4 http://www2.census.gov/topics/genealogy/1990surnames/dist.female.first 5 http://www.20000-names.com/male_spanish_names.htm</title>
      <p>4.3</p>
      <sec id="sec-3-1">
        <title>Word Embedding Pre-trained Vectors</title>
        <p>Word embeddings could be trained directly on "training data" of the domain of
application, but this strategy can lack generalization. When a new sentence to classify is
provided as input many words in it could be not possible to be translated making
impossible the correct classification. For this reason, we decided to use a common practice
of transfer learning in NLP tasks i.e. the use of vector spaces word embeddings already
pre-calculated on different domains. This allows us to cover an extensive variety of
terms by reducing the computational cost of the model and including information about
terms that are independent of their domain of use. We decided to compare the results of
the model obtained varying three different pre-trained word embeddings:
– Google word embeddings (GoogleEmb) 6: 300 dimensionality word2vec vectors,
case sensitive, composed by a vocabulary of 3 million words and phrases that are
obtained from roughly 100 billion of tokens extracted by a huge dataset of Google
News;
– GloVe (GloVeEmb)7: 300 dimensionality vectors, composed by a vocabulary of
2.2 million words case sensitive obtained from 840 billion of tokens and trained on
data crawled from generic Internet web pages;
– FastText (FastTextEmb)8: 300 dimensionality vectors, composed by a vocabulary
of 2 million words and n-grams of the words, case sensitive and obtained from
600 billion of tokens trained on data crawled from generic Internet web pages by
Common Crawl nonprofit organization;
4.4</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experimental Runs And Results</title>
        <p>
          The model has been trained using the categorical cross entropy loss function [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and
Adam optimizer for 20 epochs and best models have been used for the classification
phase. The number of batches is set as 64 for optimization reason.
        </p>
        <p>In order to evaluate the influence that the different pre-trained word embeddings
have on the final performances of the model, we have performed the training phase
keeping constant all the parameters except the word embedding matrix used to encode
the user’s tweets. The portion of the dataset used for the evaluation is equal to 20% of
the training set using 42 as a random seed. As a consequence of the not availability of
GoogleEmb and GloVe pre-trained vectors for the Spanish language, we evaluate the
model performances only on English data.</p>
        <p>
          The results in Tab. 2 showed quite better performances obtained by FastText word
embedding vector space. The differences among results are not statistically significant
but this allow us to decide to use FastText in our final model. Moreover, its availability
in both the languages of the interest of the PAN Bots and Gender Profiling challenge at
CLEF 2019 [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] (English and Spanish) has emphasized and confirmed our choice.
        </p>
        <p>
          The final model trained four times (Eng - Bot vs. Human, Eng - Male vs. Female,
Esp - Bot vs. Human, Esp - Male vs. Female) has been deployed on Tira.io 9 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6 https://goo.gl/zQFRx3 7 https://nlp.stanford.edu/projects/glove/ 8 https://fasttext.cc/docs/en/english-vectors.html 9 https://www.tira.io</title>
      <p>for participating to the PAN 2019 Bots and Gender Profiling challenge. The first run
on a preliminary released test set showed an accuracy score of 0.9470 and 0.8181
respectively for the Bot vs. Human and Male vs. Female tasks in English language. In
Spanish language we obtained 0.9611 and 0.7778 respectively for the Bot vs. Human
and Male vs. Female tasks. Finally we run the model also on the final test set earning
an accuracy score of 0.9182 and 0.7973 respectively for the Bot vs. Human and Male
vs. Female tasks in English language. In Spanish language we obtained 0.9156 and
0.7417 respectively for the Bot vs. Human and Male vs. Female tasks.</p>
      <p>
        A comparison with the baselines proposed by the challenge authors after the end
of the competition [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is showed in Fig. 3. In particular, it is interesting to observe
that our model is better than all the proposed baselines for the Spanish language. For
English, a simple strategy based on n-grams and Random Forest can overcome our
results. Probably this anomaly has been obtained as a consequence of the ability of
ngrams to capture words misspelled or unknown in a standard vocabulary like the one
used in FastText embedding space. In any case, it is performed more better than the
LDSE baseline [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. A possible future work can try to overcome these limits using one
of the newest language models, as BERT [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], as a base for our proposed network.
5
      </p>
      <sec id="sec-4-1">
        <title>Conclusion</title>
        <p>In this work, we presented an architecture based on deep neural networks for the
classification of Twitter accounts in a bot or human, carrying out, where possible, a further
refinement in man or woman. The model is based on the representation of the user
account as a list of 100 tweets appropriately transformed into tensorial form through the
encoding of words in the equivalent word embedding FastText form. Unlike state of
the art, no additional information about the Twitter account has been used allowing our
approach to be used even when this information is not available. On these
representations of input data, we learned a deep neural network that uses 2D CNNs, max-pooling
operations, and Dense Layers to estimate the probability of distribution of annotation
classes correctly. The approach was tested on the data provided by the PAN Bots and
Gender Profiling challenge at CLEF 2019, which provided appropriately annotated data
in both English and Spanish. As a preliminary step, the model was validated by varying
the pre-trained word embedding space used in the training phase. The embeddings
provided by Google and trained on News, GloVe trained on Tweets and FastText trained on
web pages collected on the web were tested. The final choice fell on FastText following
the good results obtained and its effectiveness in the tasks of classification of the text
demonstrated in the literature. The final model learned was submitted to the
competition and obtained a score of accuracy equal to: 0.9182 and 0.797 respectively for the
Bot vs. Human and Male vs. Female tasks in English language, 0.9156 and 0.7417 for
Spanish data.</p>
        <p>The good results obtained suggest that such approaches based on deep neural
networks are an excellent basis for solving the task. In particular, they are general enough
to work in a real application context correctly. Since the result obtained can only be
considered a starting point for further extensions, modifications, and improvements of the
proposed approach, we have decided to make the product code available to the whole
reference community with the hope that it will be useful for future work in the field.
The code that implements the presented model can be found at the following GitHub
repository: https://github.com/marcopoli/Identification-of-Twitter-bots-using-CNN
6</p>
      </sec>
      <sec id="sec-4-2">
        <title>Acknowledgment</title>
        <p>This work is partially funded by project "Electronic Shopping &amp; Home delivery of
Edible goods with Low environmental Footprint" (ESHELF), under the Apulian
INNONETWORK programme, Italy. Moreover, it is partially funded by project
"DECiSION" codice raggruppamento: BQS5153, under the Apulian INNONETWORK
programme, Italy.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ducharme</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vincent</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jauvin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A neural probabilistic language model</article-title>
          .
          <source>Journal of machine learning research 3(Feb)</source>
          ,
          <fpage>1137</fpage>
          -
          <lpage>1155</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the ACL 5</source>
          ,
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. zu Eissen,
          <string-name>
            <given-names>S.M.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Kulig</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Plagiarism detection without reference collections</article-title>
          .
          <source>In: Advances in Data Analysis, Proceedings of the 30th Annual Conference of the Gesellschaft für Klassifikation</source>
          e.V., Freie Universität Berlin, March 8-
          <issue>10</issue>
          ,
          <year>2006</year>
          , pp.
          <fpage>359</fpage>
          -
          <lpage>366</lpage>
          (
          <year>2006</year>
          ), https://doi.org/10.1007/978-3-
          <fpage>540</fpage>
          -70981-7_
          <fpage>40</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Deep learning</article-title>
          ,
          <source>vol. 1</source>
          . MIT press Cambridge (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woolley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          : Algorithms, bots, and
          <article-title>political communication in the us 2016 election: The challenge of automated political communication for election law and administration</article-title>
          .
          <source>Journal of information technology &amp; politics 15(2)</source>
          ,
          <fpage>81</fpage>
          -
          <lpage>93</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shimoni</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Automatically categorizing written texts by author gender</article-title>
          .
          <source>Literary and linguistic computing 17(4)</source>
          ,
          <fpage>401</fpage>
          -
          <lpage>412</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
          </string-name>
          , J.:
          <article-title>Authorship verification as a one-class classification problem</article-title>
          .
          <source>In: Proc. of the twenty-first international conference on Machine learning</source>
          . p.
          <fpage>62</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kudugunta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferrara</surname>
          </string-name>
          , E.:
          <article-title>Deep neural networks for bot detection</article-title>
          .
          <source>Information Sciences</source>
          <volume>467</volume>
          ,
          <fpage>312</fpage>
          -
          <lpage>322</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>LeCun</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.:
          <article-title>Deep learning</article-title>
          .
          <source>nature</source>
          <volume>521</volume>
          (
          <issue>7553</issue>
          ),
          <volume>436</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>LeCun</surname>
          </string-name>
          , Y., et al.:
          <article-title>Generalization and network design strategies</article-title>
          . Connectionism in perspective pp.
          <fpage>143</fpage>
          -
          <lpage>155</lpage>
          (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caverlee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Webb</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Uncovering social spammers: social honeypots+ machine learning</article-title>
          .
          <source>In: Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval</source>
          . pp.
          <fpage>435</fpage>
          -
          <lpage>442</lpage>
          . ACM (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruths</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>What's in a name? using first names as features for gender inference in twitter</article-title>
          .
          <source>In: 2013 AAAI Spring Symposium Series</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Nair</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.E.:
          <article-title>Rectified linear units improve restricted boltzmann machines</article-title>
          .
          <source>In: Proceedings of the 27th international conference on machine learning (ICML-10)</source>
          . pp.
          <fpage>807</fpage>
          -
          <lpage>814</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          . pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>TIRA Integrated Research Architecture</article-title>
          . In: Ferro,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <article-title>Information Retrieval Evaluation in a Changing World - Lessons Learned from 20 Years of</article-title>
          CLEF. Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franco-Salvador</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A low dimensionality representation for language variety identification</article-title>
          .
          <source>In: International Conference on Intelligent Text Processing and Computational Linguistics</source>
          . pp.
          <fpage>156</fpage>
          -
          <lpage>169</lpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 7th author profiling task at pan 2019: Bots and gender profiling</article-title>
          . In: Cappellato L.,
          <string-name>
            <surname>Ferro</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MÃijller</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Losada</surname>
            <given-names>D</given-names>
          </string-name>
          . (Eds.)
          <article-title>CLEF 2019 Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR Workshop Proceedings. CEUR-WS.org</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes-y Gómez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 6th author profiling task at pan 2018: multimodal gender identification in twitter</article-title>
          .
          <source>Working Notes Papers of the CLEF</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Rumelhart</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          :
          <article-title>Learning internal representations by error propagation</article-title>
          .
          <source>Tech. rep., California Univ San Diego La Jolla Inst for Cognitive Science</source>
          (
          <year>1985</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Shrestha</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sierra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solorio</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks for authorship attribution of short texts</article-title>
          .
          <source>In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume</source>
          <volume>2</volume>
          ,
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          . pp.
          <fpage>669</fpage>
          -
          <lpage>674</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Stamatatos</surname>
          </string-name>
          , E.:
          <article-title>A survey of modern authorship attribution methods</article-title>
          .
          <source>Journal of the American Society for information Science and Technology</source>
          <volume>60</volume>
          (
          <issue>3</issue>
          ),
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tepper</surname>
          </string-name>
          , J.:
          <article-title>Detecting hate speech on twitter using a convolution-gru based deep neural network</article-title>
          .
          <source>In: European Semantic Web Conference</source>
          . pp.
          <fpage>745</fpage>
          -
          <lpage>760</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>