<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CapetownMilanoTirana for GxG at Evalita2018. Simple n-gram based models perform well for gender prediction. Sometimes. (Short Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Angelo Basile</string-name>
          <email>angelo.basile@symanto.net</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth Dwyer</string-name>
          <email>garethdwyer@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Rubagotti</string-name>
          <email>chiara.rubagotti@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CoGrammar</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Independent Researcher</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Symanto Research</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe our participation in the Evalita 2018 GxG crossgenre/domain gender prediction shared task for Italian. Building on previous results obtained on in-genre gender prediction, we try to assess the robustness of a linear model using n-grams in a crossgenre setting. We show that performance drops significantly when the training and testing genres differ. Furthermore, we experiment with abstract features in trying to capture genre-independent features. We achieve an average F1-score of 0.55 on the official in-genre test set - being thus ranked first out of five submissions - and 0.51 on the cross-genre test set.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In questo articolo presentiamo il nostro
contributo per lo shared task GxG di
Evalita 2018 per l’analisi predittiva del
genere di un autore su corpora di dominˆı
diversi. Usiamo un modello basato su
una macchina a vettori supporto con
kernel lineare che usa gli n-grammi come
feature. Il modello, che ha ottenuto in
passato risultati eccellenti nella
predizione di genere quando allenato e
valutato all’interno di un unico dominio
(Twitter), crolla significativamente nella
performance in questo lavoro, anche quando
usato in combinazione con una serie di
feature astratte. Sul test set ufficiale il
nostro modello raggiunge una F1-score pari
a 0.55 (permettendoci di piazzarci in cima
alla classifica) nel contesto in-genre e 0.51
nel contesto cross-genre.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Gender prediction is the task of profiling authors
to infer their gender based on their writing. This
task has been carried out so far with success
within one-genre/single-domain data sets,
reaching state-of-the-art accuracy of 85% on English
tweets
        <xref ref-type="bibr" rid="ref1">(Basile et al., 2017)</xref>
        . Cross-domain
gender classification, on the other hand, has proven
to be more difficult, with state-of-the-art accuracy
halting at 60%
        <xref ref-type="bibr" rid="ref1 ref6">(Medvedeva et al., 2017)</xref>
        .
      </p>
      <p>
        The theoretical assumption behind gender
prediction from text is language variation: the same
meaning can be expressed in different forms and
this variation can be explained in terms of social
variables, such as social status, personality, age
and, indeed, gender
        <xref ref-type="bibr" rid="ref12 ref3 ref5">(Labov, 2006; Verhoeven et
al., 2016; Johannsen et al., 2015)</xref>
        .
      </p>
      <p>
        Proof of significant variation in language use
between men and women has been found at the
morpho-syntactic level
        <xref ref-type="bibr" rid="ref3">(Johannsen et al., 2015)</xref>
        and indeed syntactic features have been used
effectively for gender attribution
        <xref ref-type="bibr" rid="ref10">(Sarawgi et al.,
2011)</xref>
        . Using syntax for attribution tasks has the
benefit of modelling the problem in a space which
is more resilient to topic and genre effects.
However, we do not experiment here with deep
syntactic features, but instead we try to leverage surface
and frequency-based features.
      </p>
      <p>
        In this paper we use a model built, trained,
and tested for gender prediction on a single
domain (i.e. Twitter). Instead of experimenting with
new techniques, our aim is to sound the
existing model’s resilience in the different context of
a cross-genre train and test setting (RQ1). Since
we expect our model to fail on this task, we set
out to design an experiment using a set of
abstract features that has recently been applied for
cross-lingual gender prediction
        <xref ref-type="bibr" rid="ref11">(van der Goot et
al., 2018)</xref>
        : this way we want to investigate if
surface features can be used to mitigate topic and
genre effects (RQ2).
      </p>
      <p>We organise this work as follows: in Section 2
we present an overview of the data set released by
the organisers; in Section 3 we describe the
experimental set up, the model and the features used; in
Section 4 we give an overview of the results and
finally we conclude our work in Section 5.</p>
      <p>The contributions of this work for the GxG task
for EVALITA 2018 are the following:
we test if a gender prediction model
achieving state-of-the-art performances when
trained and tested on the same domain can
also achieve good results when tested in a
cross-domain setting
we experiment with abstract features in order
to factor out domain-dependent effects
we release all the code for further
reproducibility at https://github.com/
anbasile/gxg-partecipation</p>
      <p>Task description The GxG task is a document
classification task. Given a document belonging
to a given genre, we have to predict the gender of
the author. Thus the task is a binary classification
task. The task is composed by two sub-tasks:
ingenre prediction and cross-genre prediction. In the
first case, the training set and the test set belong
to the same genre. In the cross-genre sub-task, on
the other hand, models will be trained on four
genres and then tested on the single genre which they
have not been exposed to during training.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>We use only the data released by the task
organisers: that is, texts from five different genres. The
genres are as follows.</p>
      <sec id="sec-3-1">
        <title>YouTube comments</title>
      </sec>
      <sec id="sec-3-2">
        <title>Tweets</title>
      </sec>
      <sec id="sec-3-3">
        <title>Children’s writing</title>
      </sec>
      <sec id="sec-3-4">
        <title>News</title>
      </sec>
      <sec id="sec-3-5">
        <title>Personal diaries</title>
        <p>From the data description given by the
organisers we know that one author can possibly have
authored multiple documents. We provide an
overview of the data set in Table 1.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Data set</title>
      </sec>
      <sec id="sec-3-7">
        <title>Children</title>
        <p>
          Diaries
Journalism
Twitter
Youtube
In this section we describe the feature extraction
process and the model that we built. We develop
one single model and train it in ten different
configurations, as required by the task assignment.
We train and test on five different domains,
indomain and across domains.
We decide not to pre-process the data in any way,
since we have no linguistic (nor non-linguistic)
reasons for doing so. As a tokenization strategy
for the lexicalized models we simply split on all
white space tokens. For building the abstract
feature representation we use spaCy’s Italian
tokenizer
          <xref ref-type="bibr" rid="ref2 ref3">(Honnibal and Johnson, 2015)</xref>
          .
We build a sparse linear model for approaching
this task.
        </p>
        <p>
          As features we use n-grams extracted at the
word level as well as at the character level. We
use 3-10 n-grams and binary TF-IDF. We feed
these features to a Support Vector Machine (SVM)
model with a linear kernel; we use the
implementation included in scikit-learn
          <xref ref-type="bibr" rid="ref7">(Pedregosa et
al., 2011)</xref>
          . This model in this same
configuration has achieved excellent results during the PAN
2017 evaluation campaign
          <xref ref-type="bibr" rid="ref1 ref9">(Potthast et al., 2017;
Basile et al., 2017)</xref>
          .
        </p>
        <p>
          Furthermore, we experiment with feature
abstraction: we follow the bleaching approach
recently proposed by
          <xref ref-type="bibr" rid="ref11">(van der Goot et al., 2018)</xref>
          .
First, we transform each word into a list of
symbols that 1) represents the shape of the
individual characters and 2) abstracts from meaning by
still approximating the vowels and characters that
compose the word; then, we compute the length
of the word and its frequency (while taking care
of padding the first one with a zero in order to
avoid feature collision); finally, we use a boolean
label for explicitly distinguishing words from
nonalphanumeric tokens (e.g. emojis). Table 3 shows
an example of this feature abstraction process.
        </p>
        <p>
          <xref ref-type="bibr" rid="ref11">(van der Goot et al., 2018)</xref>
          proposed this
bleaching approach for successfully modelling gender
across languages, by leveraging the
languageindependent nature of these features: here, we test
if this approach is sound for mitigating the genre
effect on the model.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation and Results</title>
      <p>Since the data set labels are evenly distributed
across the two classes, we use accuracy to evaluate
our model. First, we report results obtained via a
10-fold cross-validation on the training set; then,
we report results from the official test set, whose
labels have been released.
4.1</p>
      <sec id="sec-4-1">
        <title>Development Results</title>
        <p>We report the development results obtained by
using different text representations. Table 2 presents
an overview of these results. Overall, we see that
all the different feature representation formats lead
to comparable results; the combination of words
and characters seems to be the best combination.
This is the same combination that we use for the
bleached representation.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Test Results</title>
        <p>We present official test results in Table 4. We
submitted only one run. For one genre (Diaries)
we obtained exactly the same score in both the
in- and cross-genre settings: this outcome is
extremely unlikely, however we ran the models
several times and inspected the code for bugs and yet
the results remained identical. In the cross-genre
setting, we did not tune the hyper-parameters of
our model considering the target genre.</p>
        <p>
          The overview of the results is puzzling. First,
compared to related in-domain work
          <xref ref-type="bibr" rid="ref9">(Potthast et
al., 2017)</xref>
          , the overall performance is
considerably lower, even taking into account domain
variance. Second, the drop in performance from the
in-genre to the cross-genre setting is not as high as
expected. Third, even in the in-genre setting the
difference in performance between genres is not
trivial and it seems to be independent from
training corpus’s size: the two social network domains
are considerably bigger in size and yet the testing
scores are lower when compared to other genres.
        </p>
        <p>GENRE</p>
        <p>IN</p>
        <p>CROSS</p>
        <sec id="sec-4-2-1">
          <title>Diaries</title>
          <p>YouTube
Twitter
Children
Journalism
Average
acc.
0.635
0.547
0.545
0.615
0.480
We presented our participation to the GxG
crossgenre gender prediction task and we obtained good
results using a simple system. On top of that, we
experimented with abstract features and got
suboptimal results.</p>
          <p>
            Based on our experiment with a
genderprediction model which obtained state-of-the-art
in-domain performance in the past, we conclude
that genre plays a crucial role in gender
prediction: not only genre, but the notion of variety
space, as instructed by
            <xref ref-type="bibr" rid="ref12 ref8">(Plank, 2016)</xref>
            , should be
taken into consideration for a fuller account of
social variability and for building more robust
systems (RQ1). We then attempted to improve our
stock model using abstract, delexicalized features,
but we failed to demonstrate any substantial
improvement (RQ2). Furthermore, from our results
it emerges that not all the examined genres pose
the same challenges for gender prediction
purposes: in some genres, namely journalism, the
personality of the author, whether male or female, is
hedged by the domain style, which favours
objectivity and neutrality over self-expression and
abandonment in writing. Therefore we suspect
that genre-inherent style elements might make it
harder for the model to carry out effective
profiling of the author (be it for gender or for other
social variables).
          </p>
          <p>
            Recently, Variational Auto Encoders (VAE)
            <xref ref-type="bibr" rid="ref4">(Kingma and Welling, 2013)</xref>
            are emerging as a
good tool for properly modelling language in
presence of latent variables: we plan to investigate the
effectiveness of VAEs in predicting gender while
modelling genre as a latent variable.
          </p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Gareth Dwyer, Maria Medvedeva, Josine Rawee, Hessel Haagsma, and
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>N-gram: New groningen author-profiling model</article-title>
          .
          <source>arXiv preprint arXiv:1707</source>
          .
          <fpage>03764</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Honnibal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Johnson</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>An improved non-monotonic transition system for dependency parsing</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1373</fpage>
          -
          <lpage>1378</lpage>
          , Lisbon, Portugal, 9. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Anders</given-names>
            <surname>Johannsen</surname>
          </string-name>
          , Dirk Hovy, and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Cross-lingual syntactic variation over age and gender</article-title>
          .
          <source>In Proceedings of the Nineteenth Conference on Computational Natural Language Learning</source>
          , pages
          <fpage>103</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Diederik P Kingma and Max Welling</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Auto-encoding variational bayes</article-title>
          .
          <source>arXiv preprint arXiv:1312</source>
          .
          <fpage>6114</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>William</given-names>
            <surname>Labov</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The social stratification of English in New York city</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Medvedeva</surname>
          </string-name>
          , Hessel Haagsma, and
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>An analysis of cross-genre and in-genre performance for author profiling in social media</article-title>
          .
          <source>In International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , pages
          <fpage>211</fpage>
          -
          <lpage>223</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          :
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>What to do about non-standard (or non-canonical) language in nlp</article-title>
          .
          <source>arXiv preprint arXiv:1608</source>
          .
          <fpage>07836</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          , Francisco M. Rangel Pardo, Michael Tschuggnall, Efstathios Stamatatos, Paolo Rosso, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Overview of pan'17 - author identification, author profiling, and author obfuscation</article-title>
          .
          <source>In CLEF.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Ruchita</given-names>
            <surname>Sarawgi</surname>
          </string-name>
          , Kailash Gajulapalli, and
          <string-name>
            <given-names>Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Gender attribution: tracing stylometric evidence beyond topic and genre</article-title>
          .
          <source>In Proceedings of the Fifteenth Conference on Computational Natural Language Learning</source>
          , pages
          <fpage>78</fpage>
          -
          <lpage>86</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Rob van der Goot</surname>
          </string-name>
          , Nikola Ljubesˇic´,
          <string-name>
            <surname>Ian</surname>
            <given-names>Matroos</given-names>
          </string-name>
          , Malvina Nissim, and
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bleaching text: Abstract features for cross-lingual gender prediction</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <surname>Short</surname>
            <given-names>Papers)</given-names>
          </string-name>
          , volume
          <volume>2</volume>
          , pages
          <fpage>383</fpage>
          -
          <lpage>389</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Ben</given-names>
            <surname>Verhoeven</surname>
          </string-name>
          , Walter Daelemans, and
          <string-name>
            <given-names>Barbara</given-names>
            <surname>Plank</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Twisty: a multilingual twitter stylometry corpus for gender and personality profiling</article-title>
          .
          <source>In Proceedings of the 10th Annual Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          )/Calzolari, Nicoletta [edit.]; et al., pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>