<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Quanti anni hai? Age Identification for Italian</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandra Maslennikova</string-name>
          <email>a.maslennikova@studenti.unipi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Labruna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Cimino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <email>felice.dellorlettag@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita` di Pisa Istituto di Linguistica Computazionale “Antonio Zampolli” (ILC-CNR) ItaliaNLP Lab -</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. We present the first work to our knowledge on automatic age identification for Italian texts. For this work we built a dataset consisting of more than 2.400.000 posts extracted from publicly available forums and containing authorship attribution metadata, such as age and gender. We developed an age classifier and performed a set of experiments with the aim of evaluating the possibility of assigning the correct age of an user and which information is useful to tackle this task: lexical or linguistic information spanning across different levels of linguistic descriptions. The performed experiments show the importance of lexical information in age classification, but also that exists writing style that relates to the age of an user.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. In questo articolo presentiamo
il primo lavoro a nostra conoscenza sul
riconoscimento automatico dell’eta` per la
lingua italiana. Per condurre il lavoro
abbiamo costruito un dataset composto da
piu` di 2.400.000 di post estratti da
forum pubblici e associati a informazioni
rispetto all’eta` e al genere degli autori.
Abbiamo sviluppato un sistema di
classificazione dell’eta` dello scrittore di un
testo e condotto una serie di esperimenti
per valutare se e` possibile definire l’eta` e
attraverso quali informazioni estratte dal
testo: lessicali o di descrizione
linguistica a diversi livelli. I risultati ottenuti
dimostrano l’importanza del lessico nella
classificazione, ma anche l’esistenza di
uno stile di scrittura correlato all’eta`.</p>
      <p>Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Social media platforms such as Facebook,
Twitter and public forums allow users to communicate
and share their opinions and to build social
relations. The proliferation of such platforms allowed
the scientific community to study many
communication phenomena such as the analysis of the
sentiment
        <xref ref-type="bibr" rid="ref7">(Pak et al., 2010)</xref>
        or irony
        <xref ref-type="bibr" rid="ref3">(Herna´ndez
Far´ıas et al, 2016)</xref>
        . Another related research field
is the ”author profiling” one, where the features
that allow to discriminate age, gender, or native
language of a person are analyzed. These studies
are conducted both for forensic and marketing
reasons, since the classification of these
characteristics allow companies to better focus their
marketing campaigns. In the author profiling scenario,
many are the studies conducted by the scientific
community, that were generally focused on
English and Spanish language. The majority of these
studies were performed in PAN 1
        <xref ref-type="bibr" rid="ref5">(Rangel et al.,
2016)</xref>
        , a lab at CLEF 2 that holds each year and
in which many shared tasks related to the
”authorship attribution” research topic are run. In these
shared tasks participants were asked to identify the
gender or the age using manually annotated
training data from social media platforms. Among the
most successful approaches proposed by
participants the ones that achieved the best results
        <xref ref-type="bibr" rid="ref6">(op
Vollenbroek et al., 2016)</xref>
        ,
        <xref ref-type="bibr" rid="ref4">(Modaresi et al., 2016)</xref>
        are based on SVM classifiers exploiting a wide
variety of lexical and linguistic features, such as
word n–grams, part–of–speech, and syntax. Only
recently deep learning based approaches were
proposed and have showed very good results
especially when dealing with multi–modal data, i.e.
text and images posted on Twitter
        <xref ref-type="bibr" rid="ref9">(Takahashi et
al., 2018)</xref>
        .
      </p>
      <p>In the present work we tackle a specific
author1https://pan.webis.de/
2http://www.clef-initiative.eu/
association/steering-committee
ship attribution task: the age detection for the
Italian language. To our knowledge, this is the first
time that such task is performed on Italian. For this
reason, we built a multi–topic corpus, developed a
classifier which exploits a wide range of
linguistic features, and conducted several experiments to
evaluate both the newly introduced corpus and the
classifier.</p>
      <p>The main contributions of this work are: i) an
automatically built corpus for the age detection
task for the Italian language; ii) the development
of an age detection system; iii) the study of the
impact of linguistic and lexical features.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Dataset construction</title>
      <p>With the aim of building an automatic dataset from
the web, we needed a set of Italian texts with the
age of authors publicly available. Nowadays
collecting this information is a challenging task, since
the majority of the available platforms, for the sake
of privacy, prefer not to make the user’s age public.
So, first-of-all, we had to find a website with such
data. We choose the ForumFree platform3 which
allows users to create their own forums without
any coding skills, using an existing template.
Having all the forums based on the same templates
makes them perfect for automated crawling. We
extracted all the posts of the users that decided to
show publicly their age. We tried to collect the
data from the top 200 most active forums. Not all
the forums had users with all the user information
filled and, in the end of the processes, we fetched
messages from 162 different forums. Since our
goal was to build a corpus with author profiling
purposes, and such task is very difficult with very
small comments, we selected only posts with a
minimum length of 20 words.</p>
      <p>Another problem we faced is that users are not
age-balanced in the forums: for example, anime
dedicated forum have mostly users aged under
35. Another example are cars dedicated forums,
where usually users are more mature with respect
to anime forums. Only a couple of forums have
very balanced information, which usually is the
best data for training machine learning based
classifiers. For this reason, we decided to group the
forums by their topics, because in this scenario
it is more probable to gather enough textual data
for each age gap. We manually looked the
content of all forums and assigned the topic for each
3https://www.forumfree.it/?wiki=About
one of them. We didn’t have a preassigned settled
list of possible topics. Instead, we were adding
them in the process. For example, if we have an
entire forum which discusses about only watches,
we wouldn’t assign some general ”Hobby” tag, but
we would create a special group ”Watches”
specifically for this forum.</p>
      <p>At the and of the collection process, we
collected 2.445.012 posts from 7.023 different users
and 162 forums, that we divided in 30 different
topic groups. All the information regarding the
dataset are shown in Table 1.
3</p>
    </sec>
    <sec id="sec-4">
      <title>The Age classifier</title>
      <p>
        We implemented a document age classifier that
operates on morpho–syntactically tagged and
dependency parsed texts. The classifier exploits
widely used lexical, morpho-syntatic and
syntactic features that are used to build the final
statistical model. This statistical model is finally used
to predict the age range of unseen documents.
We used linear SVM implemented in
LIBLINEAR
        <xref ref-type="bibr" rid="ref8">(Rong-En et al., 2008)</xref>
        as machine learning
algorithm. The input documents were
automatically POS tagged by the Part–Of–Speech tagger
described in
        <xref ref-type="bibr" rid="ref2 ref4">(Cimino and Dell’Orletta, 2016)</xref>
        and
dependency–parsed by the DeSR parser
        <xref ref-type="bibr" rid="ref1">(Attardi et
al., 2009)</xref>
        .
3.1
      </p>
      <sec id="sec-4-1">
        <title>Features</title>
        <p>Raw and Lexical Text Features
Word n-grams, calculated as presence or absence
of a word n-gram in the text.</p>
        <p>Lemma n-grams, calculated as the frequency of
each lemma n-gram in the text and normalized
with respect to the number of tokens in the text.
Morpho–syntactic Features</p>
      </sec>
      <sec id="sec-4-2">
        <title>Coarse and fine grained Part-Of-Speech n</title>
        <p>grams, calculated as the logarithm of the
frequency of each coarse/fine grained PoS n-gram in
the text and normalized with respect to the number
of tokens of the text.</p>
        <p>Syntactic Features</p>
      </sec>
      <sec id="sec-4-3">
        <title>Linear dependency types n-grams, calculated as</title>
        <p>the frequency of each dependency n-gram in the
text with respect to the surface linear ordering of
words and normalized with respect to the number
of tokens in the text.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Hierarchical dependency types n-grams calcu</title>
        <p>lated as the logarithm of the frequency of each
hierarchy dependency n-gram in the text and
nor</p>
        <p>Cars
Bicycles</p>
        <p>Smoking
Anime/Manga
Role playing</p>
        <p>Gaming</p>
        <p>Spirituality
Aesthetic medicine</p>
        <p>Sport
Culinary</p>
        <p>Pets
Celebrities</p>
        <p>Politics
Different topics</p>
        <p>Fishing
Institution community
Rail transport modelling</p>
        <p>Culture
Tourism</p>
        <p>Sexuality
Metal Detecting</p>
        <p>Music</p>
        <p>Parenting
Technologies</p>
        <p>Nature
Religion</p>
        <p>Films
Psychology
Gambling
Watches</p>
        <p>Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
Users
Posts
21-30
41-50</p>
        <p>Table 1: Distribution of number of users and posts per age gap in different topics in the corpus
malized with respect to the number of tokens in
the text. In addition to the dependency
relationship, the feature takes into account whether a node
is a left or a right child with respect to its parent.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>
        In order to test the corpus and the classifier, we
performed a set of experiments. The experiments
were devised in order to test real-word scenarios
where 1) we were interested to classify a set of
posts written by a single user rather then a
single post; 2) we always classified unseen users, i.e.
no training data was available for such users. For
these reasons, we merged all the posts of a
single user in the original corpus in a single
document. We then considered only the users that
wrote a minimum of 200 tokens and limited the
final merged document to a ’soft’ limit of 1000
tokens for each user. When the soft limit was
exceeded, we included the whole post that exceeded
the soft limit. The described procedure allows
training and test splits to never contain the same
user. For the age detection tasks, similarly as in
        <xref ref-type="bibr" rid="ref5">(Rangel et al., 2016)</xref>
        , we considered age-splits as
the classification classes. More precisely, we took
into account two different age group splits: the
first one, which we will refer with the name 5–
class, in which we split the documents in 5
different age groups: 20-29, 30-39, 40-49, 50-59,
6069. The second age group split, which we will
refer with the name 2–class, is composed by the
following age group splits: 29, 50-69
(excluding all the documents written by users that did not
belong to these age groups). We conducted two
different kind of experiments. In the first
experiment (in–domain), we evaluated the performance
of the classifier on in-domain texts, more precisely
we selected three different topics starting from the
main corpus and on each of the topics we trained
the classifier on the 80% of the data, and
evaluated the performance of the classifier on the
remaining 20%. For this experiment we choose the
the following domains: Sports, Watches and Cars.
In the second experiment (out–domain) we trained
the classifier on the all the 3 topics used for the
in–domain experiments and evaluated the
performance of the classifier on other 3 different topics
(Smoking, Celebrities, Metal Detecting).
      </p>
      <p>In addition, we devised 3 different machine
learning models based on 3 different sets of
features. The first one (Lexicon), which uses only
word and lemmas features, the second one
(Syntax), which uses only the morpho–syntactic and
syntactic features. Finally, the last model (All),
which uses both the lexical, morpho–syntactic and
syntactic features. We considered as baseline
model a classifier which predicts always the most
frequent class.
4.1</p>
      <sec id="sec-5-1">
        <title>Results</title>
        <p>Tables 2 and 3 report the results achieved by the
classifier for the in–domain and out–domain
experiments respectively. For what concerns all
the experiments, we can notice that the results
achieved by our classifier are higher than the
baseline results, showing that there are features that are
able to discriminate among the considered classes.
The in–domain results show that the lexical
features are the ones that have the most
discriminative power with respect to the syntax ones. The
f-score achieved by the lexicon model is 3-4 times
better than the baseline in the 5–class setting, and
2 times better in average in the 2–class setting.
The syntax model shown very good results but, as
expected, lower than the results achieved by the
lexicon model. This is an important result since
it shows that syntax and morpho–syntax are
relevant characteristics in each age-group, both in
the 5–class and 2–class settings. Surprisingly, the
All model didn’t show in any experiment an
increase in classification performance. The
classification patterns revealed in the in–domain
experiments are similarly shown also in the out–domain
experiments. The results achieved in this setting as
expected are lower than results achieved in the in–
domain settings. The 5–class experiments show a
drop in performance achieved by the considered
learning models of 8-10% f–score points in
average w.r.t. to the in–domain experiments. When
we move to the 2–class experiments, no significant
drop in performance is noticed. This shows that
in case of domain shifting, the machine learning
models are still able to well discriminate between
young and aged people.</p>
        <p>Figures 1 and 2 report the confusion matrices
of the in–domain and out–domain experiments
using the 5-class age-groups. More precisely, the
in–domain confusion matrix is obtained by
training the All model on all the three training in–
domain topics and testing the model on the
respective testset (f–score: 0.47). Similarly, the
outdomain confusion matrix is obtained by training</p>
        <p>5-class
Lexicon
0.45
0.43
0.54</p>
        <p>2-class
Lexicon
0.74
0.85
0.87
the All model on all the in-domain topics
(including the test-sets), and testing the model on the
outdomain documents of the selected 3 topics. As
it can be seen, the errors both on the in–domain
and out–domain experiments show very good
performances of the classifier, i.e., in case of errors,
usually it makes a mistake of a range of 10
years. Such results show also that the
automatically built corpus is a very useful resource for the
age classification task. Finally, it is interesting
to notice that the most correct predicted classes
are the ranges 20-29 and 40-49, both in the in–
domain and out–domain settings, while the worst
predicted class in both experiments is the 60-69
age range, most probably because is the most
underrepresented class in the training set.
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>We presented the first automatically built corpus
for the age detection task for the Italian language.
By exploiting the publicly available information
on the FreeForum platform, we built a corpus
consisting of more than 2.400.000 posts and 7.000
different users containing the user’s age
information. The first experiments performed through
a machine learning based classifier that uses a
wide range of linguistic features showed
promising results in two different range classification
tasks both in the in–domain and out–domain
settings. The conducted experiments show that
lexicon plays a fundamental role in the age
classification task both in in–domain and out–domain
scenarios. Lastly, the experiments shown that the
corpus, even though if automatically generated, is
suitable for real–world applications. We plan to
release the full corpus as soon as privacy and legal
issues will be fully investigated.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was partially supported by the 2-year
project ARTILS, Augmented RealTime Learning
for Secure workspace, funded by Regione Toscana
(BANDO POR FESR 2014-2020).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Attardi</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Maria</given-names>
            <surname>Simi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Turian</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Accurate dependency parsing with a stacked multilayer perceptron</article-title>
          .
          <source>In Proceedings of the 2nd Workshop of Evalita</source>
          <year>2009</year>
          . December, Reggio Emilia, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cimino and Felice Dell'Orletta</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Building the state-of-the-art in POS tagging of italian tweets</article-title>
          .
          <source>In Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          . Final Workshop (EVALITA),
          <source>December 5-7.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Delia</given-names>
            <surname>Irazu</surname>
          </string-name>
          ´
          <article-title>Herna´ndez Far´ıas, Viviana Patti and Paolo Rosso 2016</article-title>
          .
          <article-title>Irony Detection in Twitter: The Role of Affective Content</article-title>
          .
          <source>In ACM Transactions on Internet Technology (TOIT)</source>
          , Volume
          <volume>15</volume>
          , number 3.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Pashutan</given-names>
            <surname>Modaresi</surname>
          </string-name>
          , Matthias Liebeck and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Conrad</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Exploring the Effects of Cross-Genre Machine Learning for Author Profiling in PAN 2016</article-title>
          . In Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , E´vora, Portugal,
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          September,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Manuel Rangel Pardo</surname>
          </string-name>
          , Paolo Rosso, Ben Verhoeven, Walter Daelemans,
          <source>Martin Potthast and Benno Stein</source>
          .
          <year>2016</year>
          .
          <article-title>Overview of the 4th Author Profiling Task at PAN 2016: Cross-Genre Evaluations</article-title>
          . In Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , E´vora, Portugal,
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          September,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Mart</given-names>
            <surname>Busger op Vollenbroek</surname>
          </string-name>
          , Talvany Carlotto, Tim Kreutz, Maria Medvedeva, Chris Pool, Johannes Bjerva, and Hessel Haagsma and
          <string-name>
            <given-names>Malvina</given-names>
            <surname>Nissim</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Gronup: Groningen user profiling In Working Notes of CLEF 2016 - Conference and Labs of the Evaluation forum</article-title>
          , E´ vora, Portugal,
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          September,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Pak</surname>
          </string-name>
          and
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Paroubek</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Twitter as a corpus for sentiment analysis and opinion mining</article-title>
          .
          <source>In Proceedings of the International Conference on Language Resources and Evaluation</source>
          ,
          <string-name>
            <surname>LREC</surname>
          </string-name>
          <year>2010</year>
          ,
          <volume>17</volume>
          -
          <fpage>23</fpage>
          May
          <year>2010</year>
          , Valletta, Malta
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Fan</given-names>
            <surname>Rong-En</surname>
          </string-name>
          , Chang Kai-Wei, Hsieh Cho-Jui,
          <source>Wang Xiang-Rui and Lin Chih-Jen</source>
          .
          <year>2008</year>
          .
          <article-title>LIBLINEAR: A library for large linear classification</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>9</volume>
          :
          <fpage>1871</fpage>
          -
          <lpage>1874</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Takumi</given-names>
            <surname>Takahashi</surname>
          </string-name>
          , Takuji Tahara, Koki Nagatani, Yasuhide Miura, Tomoki Taniguchi and
          <string-name>
            <given-names>Tomoko</given-names>
            <surname>Ohkuma</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Text and image synergy with feature cross technique for gender identification</article-title>
          .
          <source>In Working Notes of CLEF</source>
          <year>2018</year>
          <article-title>- Conference and Labs of the Evaluation forum</article-title>
          , Avignon, France,
          <fpage>10</fpage>
          -
          <lpage>14</lpage>
          September,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>