<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bot and Gender Identification: Textual Analysis of Tweets</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Rodrigo Ribeiro Oliveira, Cláudio Moisés Valiense de Andrade, José Solenir Lima Figuerêdo, João B. Rocha-Junior, Rodrigo Tripodi Calumby, Iago Machado da Conceição Silva</institution>
          ,
          <addr-line>Almir Moreira da Silva Neto</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Feira de Santana Feira de Santana</institution>
          ,
          <addr-line>Bahia</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>In this paper, we describe the participation of the Advanced Data Analysis and Management (ADAM) group of the University of Feira de Santana in the Bots and Gender Profiling Task organized by PAN@CLEF 2019. We used Support Vector Machines (SVM) optimized through nested cross-validation. In bot detection, we used features related to behavior of the account, sentiment and variety of posts, in gender detection function words and emoticons. These features were evaluated both individually and in groups. Before starting the training phase, we preprocessed the data to better adjust it. For bot detection, our method reached approximately 0:9057 for English and 0:8767 for Spanish. For gender detection, 0:7696 for English and 0:7150 for Spanish. Although the results for Spanish are poorer than the ones for English, they are above the random baseline (50%).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Social media companies employ mobile and web-based technologies to create highly
interactive platforms through which individuals and communities share, cocreate,
discuss, and modify user-generated content [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. These services have changed the way we
see the world and how information is disseminated. A prominent example of such
services is Twitter1. Twitter is a popular microblogging service. Microblogging is a form
of communication in which users must describe their current status in short posts
distributed by instant messages, mobile phones, email or the Web [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this kind of social
media, users follow others or are followed. However, unlike other social networks,
Twitter demands no reciprocity in the following-followed binomial. Thus, when following
a particular user, that user is not required to follow you back. In this kind of
application, users write about different aspects of their life, sharing a variety of subjects, and
generating heterogeneous discussion.
      </p>
      <p>
        Twitter is used in multiple contexts. It is mainly considered as an information
dissemination tool, but also as a source of data that may support studies in different areas
of knowledge. This feature is specially interesting considering it offers an Application
Programming Interface (API)2 that allows crawling and collecting data. In this context,
a subject that has attracted the attention of researchers is the so-called Author profiling,
in which information like age, cultural background, gender, native language, and
personality can be inferred through textual analysis of users’ posts. This type of analysis
enables numerous applications, such as business intelligence, digital forensics,
psychological profiling, brand reputation monitoring, etc. With regard to forensic applications,
bot detection has gained attention, especially due to bots’ self-controlled ability to
disseminate political, extremist or misinformation material that may negatively influence
a massive amount of users [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        In this context, this paper describes the participation of the ADAM team in the
Bots and Gender Profiling Task [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] organized by PAN@CLEF 2019. In this edition,
different from previous years, in which various aspects of the author’s profile in social
media (age and gender, also along with personality, gender and variety of languages,
and gender from a multimodal perspective) were investigated, this edition included, in
addition, investigation whether the author of a Twitter feed is a bot or a human. The
gender profiling was maintained as a task. The analysis, as in other editions, followed
a multilingual perspective, with English and Spanish being the chosen languages. The
main contributions of this paper are:
      </p>
      <p>We define a set of features with discrimative power for bot and gender detection;
We evaluate the power of optimizing a model through cross-validation;
We analyze the effectiveness of each group of features in the task.</p>
      <p>The remainder of this paper is organized as follows. Section 2 presents the related
work. Section 3 presents the experimental validation setup. The results are discussed in
Section 4. Finally, Section 5 presents our conclusions and directions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The use of bots has brought problems in many collaborative systems, e.g., Wikipedia
and OpenStreetMaps. This scenario motivates the study of strategies to identify bots in
such collaborative systems, examining the contributions done by users [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        For textual classification, information about function words, n-grams (at word and
character level), quantitative features, orthographic features, part of speech (POS) tags,
and vocabulary richness features are usually used [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In a given country, the texts
produced in a language may vary depending on the culture of the region of origin [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Language variety identification is a popular research topic of natural language
processing [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The regional influence has an impact on the extraction of features, for example,
      </p>
      <sec id="sec-2-1">
        <title>2 https://developer.twitter.com/en/docs.html</title>
        <p>in sentiment analysis, the weight of the dictionary used may vary according to location
of the author.</p>
        <p>
          In the gender identification task, previous works show that women use question tags
more frequently, more emoticons, and less profanities [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. The work in [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] identifies
messages from men as often related to themes such as money, sports, and work, while
women refer more frequently about family, friends, and food. In addition, some subjects
are predominantly approached by men (e.g. gaming) and women (e.g. shopping).
        </p>
        <p>
          The work in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] explored some techniques for identifying twitter messages
produced by bots. In the experiment, results are presented concerning the use of neural
networks (MLP) and Random Forest. The Random Forest algorithm presented superior
performance, reaching an accuracy of 92%. In comparison to the dataset of this article,
the dataset of Braz el al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] contains additional information about the user accounts
(e.g. number of users the account follows, number of followers), impacting the results
achieved, because there is data beyong the textual aspect.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], a purely textual approach to bot detection is used. Three features are build
from the text: dissimilarity between pairs of tweets of a user; word introduction decay
rate (a measure of new unique words a user introduced over time) and average number
of URLs per tweet. The classification used 10-fold cross validation and achieved a rate
of 90.32% in detecting bots.
        </p>
        <p>
          Similar research is related to gender identification in e-mail messages [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Some
features used are those that express emotions through a sequence of characters (e.g.
“ur ssoooo kooool”, “ihaaa”) and emoticons (e.g. “:D”, “:(” ). Another feature used
by the authors in gender identification is the number of words ending with a sequence
of characters (e.g. ‘less’). The paper suggests through various research that women
make frequent use of adverbs and emotionally intensive adjectives and terms related to
questions, personal orientation and support.
        </p>
        <p>
          The work in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] compares several classifiers and states that SVM is appropriate for
classifying text, some of the reasons being the high dimensionality in the input set (large
amount of features), and having overfitting protection.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setup</title>
      <p>In this section, we present the dataset (Section 3.1), the preprossessing phase
(Section 3.2), the feature extraction method (Section 3.3) and the classification model
(Section 3.4).
3.1</p>
      <sec id="sec-3-1">
        <title>Dataset Description</title>
        <p>In order to better adjust the posts to the experiments, we executed three independent
operations:</p>
        <sec id="sec-3-1-1">
          <title>Conversion from multiple to single whitespaces; Lower-casing of all text; Removal of non-alphanumeric characters. All the operations were performed using the Natural Language Toolkit (NLTK) [13].</title>
          <p>3.3</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Feature Extraction</title>
        <p>Multiple classes of features were used for bot and gender detection. The following
sections describe each one individually, according to the associated task.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Bot detection</title>
        <p>Many of the features used in previous work to detect bots in twitter, as usernames,
geodata, tweet intervals and following data, is unavailable in this case. Therefore, the
features are limited to the twitter text.</p>
        <p>Twitter features: these features are related to the twitter profile of the users.
Frequency of hashtags (#), frequency of mentions of users (@), frequency of retweets
(occurrence of the string “rt”) and frequency of links (occurrence of “http”). The
rationale behind the first two is that bots tend to try to increase their reach inserting
trending hashtags in their posts or mentioning multiple users to call their attention.
Bots use to retweet content as a way to easily build a profile, and constant posting
of links is typical behavior of spam bots.</p>
        <p>
          Sentiment features: in the English corpus, the VADER library3 was used to
identify the sentiments of the corpus. The VADER generates the mean of the four
sentiment metrics: positive, negative, neutral and compound. An additional feature was
also computed, the sentiment flip, defined in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] as the number of sentiment
inversions (positive to negative and vice-versa) between two adjacent posts normalized
by the total number of tweets authored by the user. In VADER, the compound value
has the range [
          <xref ref-type="bibr" rid="ref1">-1, 1</xref>
          ], so is the one used to calculate the sentiment flip, the point of
inversion being 0. Getting the sentiment features for the Spanish corpus was
somewhat difficult. The method recommended by the VADER developers is to translate
        </p>
        <sec id="sec-3-3-1">
          <title>3 https://github.com/cjhutto/vaderSentiment</title>
          <p>
            the texts to English automatically and use the library on the resulting text. We ended
using a machine learning based solution4, which generated only one output in the
range [
            <xref ref-type="bibr" rid="ref1">0, 1</xref>
            ]. The point of sentiment inversion was fixed on 0.5.
          </p>
          <p>Variety features: these features measure how varied is the content generated by
the user, under the assumption that bots tend to repeat content in their posts. Two
features belong to this group: the ratio between number of words used and total
number of words in all posts; and the cleanliness: the ratio between the number of
characters after and before preprocessing.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Gender detection</title>
        <p>Emoticons: frequency of each term in a list of emoticons, meant to describe a range
of emotions.</p>
        <p>Function words: frequency of function words: pronouns, determiners, modals and
conjunctions in English and in Spanish, conjunctions and determiners.</p>
        <p>Sentiment features: the same sentiment features used for bot detection.
3.4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Classification Model</title>
        <p>
          For conducting the experiments we used a SVM classifier. In order to optimize the
model learned, a 5-fold nested cross-validation [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] was done. In nested cross-validation,
after a variation of parameters in order to find the optimal model, one more
crossvalidation is done to evaluate the model found.
        </p>
        <p>This process is done only in the train dataset. After the model is chosen, it is used to
classify the dev dataset, giving a notion of how the model will perform in unseen data,
i.e., this set validates the model. To make this last step statistically significant, during
the learning process nothing of the dev dataset was used.</p>
        <p>
          In the cross-validation, both linear and RBF kernels were used varying their
respecting hyperparameters (C and in RBF, gamma). To do this, we used the scikit-learn
library [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <p>The training and validation steps were done to each feature group individually at first.
After that, the same was done using all the features together, so the impact of each
feature group on the overall result could be evaluated.</p>
      <p>
        The final evaluation was performed with the official PAN@CLEF 2019 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] test set
using the TIRA platform [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The results were grouped according to their associated
task, i.e, bot or gender detection. For each task we present the assessment of the
classifier for the three datasets provided: Train; Validation; and Test.
4.1
      </p>
      <sec id="sec-4-1">
        <title>Bot detection</title>
        <sec id="sec-4-1-1">
          <title>4 https://github.com/aylliote/senti-py</title>
          <p>In this paper we described the participation of the ADAM team in the Bots and Gender
Profiling Task organized by PAN@CLEF 2019. In this task, focused on Twitter posts,
should be determined if the author of a Twitter feed was a bot or human. Moreover, in
case of a post from a human, the challenge was to identify the gender. We used a set of
features in SVM with cross-validation.</p>
          <p>The final outcome suggests that our proposal, in general, achieves good results when
compared to the random baseline. The best results were achieved for Bot detection in
English. Although the accuracy for gender detection is inferior to the accuracy in Bot
detection, it also presents promising results.</p>
          <p>As future work, we suggest trying to extract metadata from the tweet text, for
example, building networks of citations through user mentions. Another approach is to
detect how original are the posts of a user within the corpus, since bots are known to
replicate human content to fake authenticity. In addition, it is interesting to experiment
different machine learning approaches, such as deep learning. Better tools for features
in non-English languages are needed, like the Sentiment ones, which are scarce.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Braz</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldschmidt</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          :
          <article-title>Redes neurais convolucionais na detecção de bots sociais: Um método baseado na clusterização de mensagens textuais</article-title>
          . In: Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais (SBSeg),
          <source>SBSeg 2018</source>
          . pp.
          <fpage>323</fpage>
          -
          <lpage>336</lpage>
          . SBC (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galbraith</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danforth</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dodds</surname>
            ,
            <given-names>P.S.:</given-names>
          </string-name>
          <article-title>Sifting robotic from organic text: a natural language approach for detecting automation on twitter</article-title>
          .
          <source>Journal of Computational Science</source>
          <volume>16</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Corney</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Vel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohay</surname>
          </string-name>
          , G.:
          <article-title>Gender-preferential text mining of e-mail discourse</article-title>
          .
          <source>In: 18th Annual Computer Security Applications Conference (ACSAC)</source>
          ,
          <year>2002</year>
          . Proceedings. pp.
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          . IEEE (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kestemont</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manjavancas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Specht</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tschuggnall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zangerle</surname>
          </string-name>
          , E.: Overview of PAN 2019:
          <article-title>Author Profiling, Celebrity Profiling, Cross-domain Authorship Attribution and Style Change Detection</article-title>
          . In: Crestani,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Rauber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Heinatz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Cappellato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.)
          <source>Proceedings of the Tenth International Conference of the CLEF Association (CLEF</source>
          <year>2019</year>
          ). Springer (Sep
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dickerson</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kagan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subrahmanian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Using sentiment to detect bots on twitter: Are humans more opinionated than bots?</article-title>
          <source>In: Proceedings of the 2014 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining</source>
          . pp.
          <fpage>620</fpage>
          -
          <lpage>627</lpage>
          . IEEE Press (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ferrara</surname>
          </string-name>
          , E.:
          <article-title>Disinformation and social bot operations in the run up to the 2017 french presidential election</article-title>
          .
          <source>First Monday</source>
          <volume>22</volume>
          (
          <issue>8</issue>
          ) (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Terveen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halfaker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Bot detection in wikidata using behavioral and other informal cues</article-title>
          .
          <source>Proceedings of the ACM on Human-Computer Interaction 2(CSCW)</source>
          ,
          <volume>64</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Java</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tseng</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Why we twitter: Understanding microblogging usage and communities</article-title>
          .
          <source>In: Proceedings of the 9th WebKDD and 1st SNA-KDD 2007 Workshop on Web Mining and Social Network Analysis</source>
          . pp.
          <fpage>56</fpage>
          -
          <lpage>65</lpage>
          . ACM, New York, NY, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Joachims</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Text categorization with support vector machines: Learning with many relevant features</article-title>
          .
          <source>In: European conference on machine learning (ECML)</source>
          . pp.
          <fpage>137</fpage>
          -
          <lpage>142</lpage>
          . Springer (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kietzmann</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hermkens</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>I.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silvestre</surname>
            ,
            <given-names>B.S.:</given-names>
          </string-name>
          <article-title>Social media? get serious! understanding the functional building blocks of social media</article-title>
          .
          <source>Business Horizons</source>
          <volume>54</volume>
          (
          <issue>3</issue>
          ),
          <fpage>241</fpage>
          -
          <lpage>251</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shimoni</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Automatically categorizing written texts by author gender. Literary and linguistic computing (LLC) 17(4</article-title>
          ),
          <fpage>401</fpage>
          -
          <lpage>412</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lakoff</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Language and woman's place</article-title>
          .
          <source>Language in society 2(1)</source>
          ,
          <fpage>45</fpage>
          -
          <lpage>79</lpage>
          (
          <year>1973</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Loper</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Nltk: the natural language toolkit</article-title>
          .
          <source>arXiv preprint cs/0205028</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , et al.:
          <article-title>Scikit-learn: Machine learning in python</article-title>
          .
          <source>Journal of machine learning research (JMLR) 12(Oct)</source>
          ,
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehl</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niederhoffer</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          :
          <article-title>Psychological aspects of natural language use: Our words, our selves</article-title>
          .
          <source>Annual review of psychology 54(1)</source>
          ,
          <fpage>547</fpage>
          -
          <lpage>577</lpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>TIRA Integrated Research Architecture</article-title>
          . In: Ferro,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <surname>C</surname>
          </string-name>
          . (eds.)
          <article-title>Information Retrieval Evaluation in a Changing World - Lessons Learned from 20 Years of</article-title>
          CLEF. Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franco-Salvador</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A low dimensionality representation for language variety identification</article-title>
          . In: Gelbukh,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (ed.)
          <source>Computational Linguistics and Intelligent Text Processing</source>
          . pp.
          <fpage>156</fpage>
          -
          <lpage>169</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of the 7th Author Profiling Task at PAN 2019: Bots and Gender Profiling</article-title>
          . In: Cappellato,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Müller</surname>
          </string-name>
          , H. (eds.)
          <article-title>CLEF 2019 Labs and Workshops, Notebook Papers</article-title>
          .
          <source>CEUR-WS.org (Sep</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.:</given-names>
          </string-name>
          <article-title>Effects of age and gender on blogging</article-title>
          . In: AAAI spring symposium:
          <article-title>Computational approaches to analyzing weblogs</article-title>
          .
          <source>vol. 6</source>
          , pp.
          <fpage>199</fpage>
          -
          <lpage>205</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Stone</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Cross-validatory choice and assessment of statistical predictions</article-title>
          .
          <source>Journal of the Royal Statistical Society: Series B (Methodological)</source>
          <volume>36</volume>
          (
          <issue>2</issue>
          ),
          <fpage>111</fpage>
          -
          <lpage>133</lpage>
          (
          <year>1974</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>