<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the Author Profiling Task at PAN 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco Rangel</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moshe Koppel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Efstathios Stamatatos</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giacomo Inches</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Autoritas Consulting</institution>
          ,
          <addr-line>S.A.</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Science, Bar-Illan University</institution>
          ,
          <country country="IL">Israel</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dept. of Information and Communication Systems Engineering, University of the Aegean</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Information Retrieval Group, Faculty of Informatics, University of Lugano</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Natural Language Engineering Lab, ELiRF, Universitat Politècnica de València</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This overview presents the framework and results for the Author Profiling task at PAN 2013. We describe in detail the corpus and its characteristics, and the evaluation framework we used to measure the participants performance to solve the problem of identifying age and gender from anonymous texts. Finally, the approaches of the 21 participants and their results are described.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        In classical authorship attribution, we are given a closed set of candidate authors
and are asked to identify which one of them is the author of an anonymous text.
Author profiling, on the other hand, distinguishes between classes of authors, rather
than individual authors. Thus, for example, profiling is used to determine an author’s
gender, age, native language, personality type, etc. Author profiling is a problem of
growing importance in a variety of areas, including forensics, security and marketing.
For instance, from a forensic linguistics perspective, being able to determine the
linguistic profile of the author of a suspicious text solely by analyzing the text could
be extremely valuable for evaluating suspects. Similarly, from a marketing viewpoint,
companies may be interested in knowing, on the basis of the analysis of blogs and
online product reviews, what types of people like or dislike their products. Here we
consider the problem of author profiling in social media, with particular focus on the
use of everyday language and how this reflects basic social and personality processes.
Our starting point is the seminal work of Argamon et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where it was shown that
statistical analysis of word usage in documents could be used to determine an author’s
gender, age, native language and personality type.
      </p>
      <p>
        In PAN 20131 we consider the gender and age aspects of the author profiling
problem, both in English and Spanish. So far research work in computational linguistics
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and social psychology [26] has been carried out mainly for English. We believe it
is interesting to investigate gender and age classification task in a language other than
1 http://www.uni-weimar.de/medien/webis/research/events/pan-13/pan13-web/index.html
      </p>
      <sec id="sec-1-1">
        <title>English, therefore considering Spanish, too.</title>
        <p>In Section 2 we present the state of the art, describing related work and how the
task has been approached. In Section 3 we describe the details of the collection used
and the evaluation measures. In Section 4 we present the authors’ approaches and we
discuss the results in Section 5, concluding the overview in Section 6.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        The study of how certain linguistic features vary according to the profile of their authors
is a subject of interest for several different areas such as psychology, linguistics and,
more recently, natural language processing. Pennebaker et al. [27] connected language
use with personality traits, studying how the variation of linguistic characteristics
in a text can provide information regarding the gender and age of its author.
Argamon et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] analyzed formal written texts extracted from the British National Corpus,
combining function words with part-of-speech features and achieving approximately
80% accuracy in gender prediction. Other researchers (Holmes and Meyerhoff[13],
Burger and Henderson[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) have also investigated obtaining age and gender information
from formal texts.
      </p>
      <p>With the rise of the social media, the focus is on other kind of writings, more
colloquial, less structured and formal, like blogs or fora. Koppel et al. [16] studied the
problem of automatically determining an author’s gender by proposing combinations
of simple lexical and syntactic features, and achieving approximately 80% accuracy.
Schler et al. [30] studied the effect of age and gender in the style of writing in
blogs; they gathered over 71,000 blogs and obtained a set of stylistic features like
non-dictionary words, parts-of-speech, function words and hyperlinks, combined with
content features, such as word unigrams with the highest information gain. They
obtained an accuracy of about 80% for gender identification and about 75% for age
identification. They demonstrated that language features in blogs correlates with age, as
reflected in, for example, the use of prepositions and determiners. Goswami et al. [12]
added some new features as slang words and the average length of sentences, improving
accuracy to 80.3% in age group detection and to 89.2% in gender detection.</p>
      <p>It is to be noted that the previously described studies were conducted with texts of
at least of 250 words. The effect of data size is known, however, to be an important
factor in machine learning algorithms of this type. In fact, Zhang and Zhang [33]
experimented with short segments of blog post, specifically 10,000 segments with 15
tokens per segment, and obtained 72.1% accuracy for gender prediction, as opposed
to more than 80% in the previous studies. Similarly, Nguyen et al. [22] studied the
use of language and age among Dutch Twitter users, where the documents are really
short, with an average length of less than 10 terms. They modelled age as a continuous
variable (as they had previously done in [21]), and used an approach based on logistic
regression. They also measured the effect of the gender in the performance of age
detection, considering both variables as inter-dependent, and achieved correlations up
to 0.74 and mean absolute errors between 4.1 and 6.8 years.</p>
      <p>One common problem when investigating the author profiling problem is the need
to obtain labelled data for the authors, for example, to obtain their age and gender.
Studies in classical literature deals with a small number of well-known authors, where
manual labelling can easily be applied, however for the dimensions of the actual social
media data this is a more difficult task, which should be automated. In some cases,
researchers manually label the collection [22] with some risk of bias. In other cases,
as in the vast majority of the aforementioned studies, researchers took into account
information provided by the authors themselves. For example, in blog platforms, the
contributors self-specify their profiles. This is the case for Peersman et al. [25] who
retrieved a dataset from Netlog2,where authors report their gender and exact age, and
Koppel et al. [16], who retrieved the dataset from Blogspot3. In these cases we have
to be aware of a common issue, the use of these media (mainly blogs) to promote web
positions in search engines through the use of false profiles. This is likely to introduce
noise to the evaluation corpus, but it also reflects the realistic state of the available data.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation framework</title>
      <p>In this section we describe the data collection obtained for the task, its properties,
challenges and novelties as well the evaluation measures.
3.1</p>
      <sec id="sec-3-1">
        <title>Data collection</title>
        <p>We built the corpus with thousands of blog posts taking into account that:
– The variety of themes provides a wide spectrum of topics, making the task of
determining age and gender more realistic. The ample diversity of topics allows to
investigate standard cliches, for example, men speaking a lot about beer or football
and women about nails or shopping, for breaking or reinforcing them.
– Blog posts are used daily for search engine optimization and can be automatically
generated by robots or be advertisements (chatbots).
– People may use social media to talk also about sex and few can also break the line
and use these systems to misbehave and engage in conversations that may result
into sexual harassment. For this reason and due to the importance of unveiling
fake profiles, we decided to test the robustness of the author profiling approaches
including in our collection some texts from last year PAN task on sexual predator
identification.
– We wanted to carry out the task in a multilingual setting, therefore, in addition to
English we included a Spanish part in our collection. Spanish and English are two
of the most used languages in the world4.</p>
        <sec id="sec-3-1-1">
          <title>2 http://www.netlog.com</title>
          <p>3 http://blogspot.com
4 http://www.internetworldstats.com/stats7.htm
en</p>
          <p>es
We looked online for open and public repositories such as Netlog with posts labelled
with author demographics such as gender and age. Once found, we decided to group
posts by author, selecting those authors with at least one post, and chunking in different
files those authors with more than 1,000 words in their posts. We also included authors
with very few and possibly short posts in order to maintain a realistic evaluation
framework. We divided the collection into the following parts: training, early bird
evaluation and final testing. Authors were randomly split into these parts, making
sure that each author is included in exactly one part. For age detection, we followed
what was previously done in [30] and considered three classes: 10s (13-17), 20s
(23-27) and 30s (33-47). The collection is balanced by gender and imbalanced by
age group. Additionally, trying to preserve a real-world scenario5 , we incorporated a
small number of samples from conversations of sexual predators [14] together with
samples from adult-adult conversations about sex. In Table 1 we illustrate the statistics
of English and Spanish collections.</p>
          <p>In the training part of the English collection, numbers inside parentheses for male
20s and 30s correspond to the number of samples of sexual predator conversations
5 E.g. There are statistics of about 200 tweets per hour in English from sexual predators
(http://www.mirror.co.uk/news/uk-news/paedophiles-using-twitter-to-find-victims-1253833).
Twitter issued about 200 million tweets per day
(https://blog.twitter.com/2011/200million-tweets-day) in 2011, achieving 400 million tweets per day in 2013
(http://www.webpronews.com/twitter-turns-7-boasts-400m-tweets-per-day-2013-03). This is
about 0.0012%
while numbers inside parenthesis for female 20s correspond to the adult-adult sexual
conversation samples. We provided these samples for training purposes. In the
collection for early bird evaluation, we did not include any sample of this kind. The
final collection was built adding a 20% of samples over the early bird dataset this
time including samples from sexual predator conversations for male 20s and 30s, and
samples from adult-adult conversations for female 20s.</p>
          <p>As can be seen, there are significant differences between the two languages.
More than 80% of Spanish posts are about 15-word long (e.g. greetings, especially
for teenagers). On the other hand, English speakers seem to describe situations,
experiences or thoughts, but in a more elaborated way.
3.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Performance measures</title>
        <p>For evaluating participants’ approaches we have used accuracy. Concretely, we
calculated the ratio between the number of authors correctly predicted by the total
number of authors. We calculated separately the accuracy for each language, gender,
and age group. Moreover, we combined accuracy for the joint identification of age and
gender. The final score used to rank the participants is the average for the combined
accuracies for each language.</p>
        <p>We also calculated the total number of correctly identified gender and age for
predator samples, in order to determine what approaches are more robust to this kind
of outliers. Finally, we calculated the total time needed to process the test data, in order
to investigate the difficulties of processing big volumes of data in the framework of a
real-world application.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Overview of the participants’ approaches</title>
      <p>We received 21 submissions for the task of Author Profiling and a total of 18 notebook
papers: 8 long papers and 10 short papers. We present the analysis of the 18 approaches
we received a description of.</p>
      <p>
        Pre-processing . Only few participants preprocessed the data. Various participants
[23][20][19][32][24] cleaned HTML to obtain plain text, one participant [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] deleted
those documents containing at least 0.1% of spam words and another participant [17]
used Principal Component Analysis to linearly reduce the dimensionality. During the
training phase, some participants [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ][
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][20][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][29] selected a subset from the training
data in order to reduce dimensionality. Only one participant [19] tried to discriminate
between human-like posts and spam-like posts or chatbots.
      </p>
      <p>
        Features . Many participants [17] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [24] [23] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] [19] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [28] used stylistic
features such as frequencies of punctuation marks, capital letters, quotations, and so
on, together with POS tags [17] [19] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [28] or HTML-based features as image urls
or links [28] [29] [19]. Readability features has been widely used in several approaches
[23] [17] [19] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [32] [11]. In the last approach readability features were the only
ones used. Emoticons were used by two participants [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and discarded from one
participant [29].
      </p>
      <p>
        Different content features (e.g. Latent Semantic Analysis, bag of words, TF-IDF,
dictionary-based words, topic-based words, entropy-based words, and so on) were also
used by many participants [29] [23] [17] [31] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [19] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [28] [24] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Different
participants considered named entities [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], sentiment words [23], emotion words [19],
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and slang, contractions and words with character flooding [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        A different approach based on information retrieval was presented by one
participant [32]. In such approach, the text to be identified was used as a query for a
search engine. One participant [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] introduced a high variety of corpus statistics to
build unsupervised features and four participants [19] [15] [20] [29] used n-grams
models. Finally, one participant [19] introduced advanced linguistic features such as
collocations and another participant [18] used second order representation based on
relationships between documents and profiles.
      </p>
      <p>
        Classification approaches . All the approaches used supervised machine learning
methods. The vast majority of them [28] [23] [31] [11] [32] used decision trees. Three
approaches [17] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [29] used Support Vector Machines, two approaches [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] used
logistic regression, and the rest used Naïve Bayes [19], Maximum Entropy [24],
Stochastic Gradient Descent [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and random forest [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation of the participants’ approaches and discussion</title>
      <p>We divided the evaluation in two steps, an early bird option for those who wanted to
test their approaches before the final submission in order to have some feedback, and
the final evaluation. There were 5 early bird submissions and 21 for final evaluation.
We could not evaluate one early bird submission due to runtime errors on the TIRA6
platform. A baseline was provided in order to compare the different approaches with.
This baseline was programmed as two random classifiers for each variable (gender
and age group), obtaining 50% of accuracy for gender identification and 33% for age
identification, and 16.5% for joint identification.</p>
      <p>In Table 2 the performance of early bird submissions is shown. In Table 3 the final
ranking for each language is presented. We show the accuracy for gender and age
group and the accuracy for the joint identification. The difficulty of the task is reflected
in the low values of such measure, especially for gender identification with close to the
baseline. In addition, the joint identification shows a dramatic decrease in the result,
highlighting the even greater difficulty of the joint identification.</p>
      <sec id="sec-5-1">
        <title>6 http://tira.webis.de</title>
        <p>
          It is difficult to establish a correlation between the used features in the different
approaches and the obtained results, due mainly to the amount of shared features
in all of them. It is to be noted the usage of second order representations based on
relationships between documents and profiles by the winner of the task [18] and the
use of collocations for the winner of the English task [19], features that do not seem
to be as good for Spanish (or maybe more difficult to tune). Stylistic and content
features were used for the vast majority of approaches and the obtained values for
accuracy show results in different positions of the ranking. POS features were used in
five different approaches, e.g. by systems in the first position for English [19] and in
the first position for the Spanish [28], with values under the median of the ranking for
the rest of the approaches. Such features seem to improve the performance on the task.
Readability is another feature widely used for the vast majority of the approaches. We
can compare the performance of this feature with the rest because there is an approach
[11] based only on such feature, achieving the 8th position in English and the 13th in
Spanish. Except one approach [19], those which used n-gram features did not achieve
very good results, all of them over the median of the ranking. The use of sentiment
words [23] and emotion words [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] does not seem to improve the accuracy, in the
same manner than the use of slang words [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], although these approaches
used many other features and it is difficult to establish a correlation.
        </p>
        <p>Regarding employing some kind of preprocessing, it is interesting that except two
cases [19] [17] the rest get worse performance, although it may be probably due to the
features used not to the preprocessing itself.</p>
        <p>In Table 4 the identification of fake profiles for sexual predators is shown. The
first group of columns shows the number of correctly identified profiles for adult-adult
sexual conversations and the second group shows the number of correctly identified
fake profiles for sexual predators. In brackets the ratio is shown.</p>
        <p>
          The vast majority of participants identified correctly cases of adult-adult sexual
conversations but what is more surprising is that all the participants identified the
right age and gender of many predator samples. At least 7 participants identified more
than 50% of such cases, 10 participants identified gender for more than 95% of the
cases and 7 participants identified age for more than 50% of them. Best results were
obtained by 3 participants who combined content and stylistic features [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] [19] [24],
one participant who used n-grams [15] and one participant who used a content-based
approach improved with specific dictionaries (slang, contractions...) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The approach
based on Information Retrieval techniques [32] also obtained top results. The approach
based only on the readability features [11] obtained 42% of accuracy, meaning that
such features have an important impact on detecting such cases.
        </p>
        <p>Finally, in Table 5 we show the time each participant needed to finish the task,
reversely ordered by runtime. Runtime is shown in milliseconds. The differences
between the fastest (10.26 minutes) [11] and the slowest (11.78 days) [31] is enormous.
The fastest [11] approached the task only with the readability features, obtaining the
8th position in English and 13th in Spanish. The slowest [31] approached the task with
content features, obtaining the 3rd. position in English and 21th in Spanish. The vast
majority of approaches took a few hours. The slowest participants used collocations
[19], POS [17], n-grams [20] and performed preprocessing such as html removal [20]
[19], detection of chatbots [19] and Principal Component Analysis [17].
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper we present the results of the 1st International Author Profiling Task at
PAN-2013 within CLEF-2013. Given a large and realistic collection of blog posts and
chat logs, the 21 participants of the task had to identify gender and age of anonymous
authors.</p>
      <p>Participants used several different features to approach the problem, being able
to be grouped into content-based (bag of words, named entities, dictionary words,
slang words, contractions, sentiment words, emotion words, and so on), stylistic-based
(frequencies, punctuations, POS, HTML use, readability measures and many different
statistics), n-grams based, IR-based and collocations-based. Results show the difficulty
of the task, mainly for the gender identification and for the joint identification of gender
and age.</p>
      <p>We introduced some conversations from sexual predators in order to check the
robustness of the approaches, and we were pleasantly surprised by the high amount of
such cases correctly identified by all the participants.</p>
      <sec id="sec-6-1">
        <title>Acknowledgements</title>
        <p>The author profiling task @PAN-2013 was an activity of the WIQ-EI IRSES project
(Grant No. 269180) within the FP 7 Marie Curie People Framework of the European
Commission. We want to thank the Forensic Lab of the Universitat Pompeu Fabra
Barcelona for sponsoring the award for the winner team. The work of the first author
was partially funded by Autoritas Consulting SA and by Ministerio de Economía y
Competitividad de España under grant ECOPORTUNITY IPT-2012-1220-430000. The
work of the second author was in the framework the DIANA-APPLICATIONS-Finding
Hidden Knowledge in Texts: Applications (TIN2012-38603-C02-01) project, and the
VLC/CAMPUS Microcluster on Multimodal Interaction in Intelligent Systems. The
work of fifth author was funded in part by the Swiss National Science Foundation
(SNF) project "Mining Conversational Content for Topic Modelling and Author
Identification (ChatMiner)" under grant number 200021_130208.
[10] Pamela Forner, Roberto Navigli, and Dan Tufis, editors. CLEF 2013 Evaluation
Labs and Workshop – Working Notes Papers, 23-26 September, Valencia, Spain,
2013.
[11] Lee Gillam. Readability for author profiling?—Notebook for PAN at CLEF
2013. In Forner et al. [10].
[12] Sumit Goswami, Sudeshna Sarkar, and Mayur Rustagi. Stylometric analysis of
bloggers’ age and gender. In Eytan Adar, Matthew Hurst, Tim Finin, Natalie S.
Glance, Nicolas Nicolov, and Belle L. Tseng, editors, ICWSM. The AAAI Press,
2009.
[13] Janet Holmes and Miriam Meyerhoff. The Handbook of Language and Gender.</p>
        <p>Blackwell Handbooks in Linguistics. Wiley, 2003. ISBN 9780631225027.
[14] Giacomo Inches and Fabio Crestani. Overview of the International Sexual
Predator Identification Competition at PAN-2012. In Pamela Forner, Jussi
Karlgren, and Christa Womser-Hacker, editors, CLEF 2012 Evaluation Labs and
Workshop – Working Notes Papers, 17-20 September, Rome, Italy, September
2012.
[15] Magdalena Jankowska, Vlado Keselj, and Evangelos Milios. CNG Text
Classification for Authorship Profiling Task—Notebook for PAN at CLEF 2013.</p>
        <p>In Forner et al. [10].
[16] Moshe Koppel, Shlomo Argamon, and Anat Rachel Shimoni. Automatically
categorizing written texts by author gender, 2003.
[17] Wee Yong Lim, Jonathan Goh, and Vrizlynn L. L. Thing. Content-Centric Age
and Gender Profiling—Notebook for PAN at CLEF 2013. In Forner et al. [10].
[18] A. Pastor Lopez-Monroy, Manuel Montes-Y-Gomez, Hugo Jair Escalante, Luis
Villasenor-Pineda, and Esau Villatoro-Tello. INAOE’s Participation at PAN’13:
Author Profiling task—Notebook for PAN at CLEF 2013. In Forner et al. [10].
[19] Michal Meina, Karolina Brodzinska, Bartosz Celmer, Maja Czokow, Martyna
Patera, Jakub Pezacki, and Mateusz Wilk. Ensemble-based Classification for
Author Profiling Using Various Features—Notebook for PAN at CLEF 2013. In
Forner et al. [10].
[20] Erwan Moreau and Carl Vogel. Style-based Distance Features for Author</p>
        <p>Profiling—Notebook for PAN at CLEF 2013. In Forner et al. [10].
[21] Dong Nguyen, Noah A. Smith, and Carolyn P. Rosé. Author age prediction from
text using linear regression. In Proceedings of the 5th ACL-HLT Workshop on
Language Technology for Cultural Heritage, Social Sciences, and Humanities,
LaTeCH ’11, pages 115–123, Stroudsburg, PA, USA, 2011. Association for
Computational Linguistics.
[22] Dong Nguyen, Rilana Gravel, Dolf Trieschnigg, and Theo Meder. "how old do
you think i am?"; a study of language and age in twitter. Proceedings of the
Seventh International AAAI Conference on Weblogs and Social Media, 2013.
[23] Braja Gopal Patra, Somnath Banerjee, Dipankar Das, Tanik Saikh, and Sivaji
Bandyopadhyay. Automatic Author Profiling Based on Linguistic and Stylistic
Features—Notebook for PAN at CLEF 2013. In Forner et al. [10].
[24] Aditya Pavan, Aditya Mogadala, and Vasudeva Varma. Author Profiling Using
LDA and Maximum Entropy—Notebook for PAN at CLEF 2013. In Forner et al.
[10].
[25] Claudia Peersman, Walter Daelemans, and Leona Van Vaerenbergh. Predicting
age and gender in online social networks. In Proceedings of the 3rd international
workshop on Search and mining user-generated contents, SMUC ’11, pages
37–44, New York, NY, USA, 2011. ACM.
[26] James W. Pennebaker. The Secret Life of Pronouns: What Our Words Say About</p>
        <p>Us. Bloomsbury USA, 2013. ISBN 9781608194964.
[27] James W. Pennebaker, Mathias R. Mehl, and Kate G. Niederhoffer.</p>
        <p>Psychological aspects of natural language use: Our words, our selves. Annual
review of psychology, 54(1):547–577, 2003.
[28] K Santosh, Romil Bansal, Mihir Shekhar, and Vasudeva Varma. Author
Profiling: Predicting Age and Gender from Blogs—Notebook for PAN at CLEF
2013. In Forner et al. [10].
[29] Upendra Sapkota, Thamar Solorio, Manuel Montes-Y-Gomez, and Gabriela
Ramirez-De-La-Rosa. Author Profiling for English and Spanish Text—Notebook
for PAN at CLEF 2013. In Forner et al. [10].
[30] Jonathan Schler, Moshe Koppel, Shlomo Argamon, and James W. Pennebaker.</p>
        <p>Effects of age and gender on blogging. In AAAI Spring Symposium:</p>
        <p>Computational Approaches to Analyzing Weblogs, pages 199–205. AAAI, 2006.
[31] Mechti Seifeddine, Jaoua Maher, and Hadrich Belghith Lamia. Author Profiling
Using Style-based Features—Notebook for PAN at CLEF 2013. In Forner et al.
[10].
[32] Edson Weren, Viviane P. Moreira, and Jose Oliveira. Using Simple Content
Features for the Author Profiling Task—Notebook for PAN at CLEF 2013. In
Forner et al. [10].
[33] Cathy Zhang and Pengyu Zhang. Predicting gender from blog posts. Technical
report, Technical Report. University of Massachusetts Amherst, USA, 2010.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Yuridiana</given-names>
            <surname>Aleman</surname>
          </string-name>
          , Nahun Loya, Darnes Vilarino Ayala, and
          <string-name>
            <given-names>David</given-names>
            <surname>Pinto</surname>
          </string-name>
          .
          <article-title>Two Methodologies Applied to the Author Profiling Task-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          , Moshe Koppel, Jonathan Fine, and Anat Rachel Shimoni. Gender, genre, and
          <article-title>writing style in formal written texts</article-title>
          . TEXT,
          <volume>23</volume>
          :
          <fpage>321</fpage>
          -
          <lpage>346</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          , Moshe Koppel, James W. Pennebaker, and
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Schler</surname>
          </string-name>
          .
          <article-title>Automatically profiling the author of an anonymous text</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>52</volume>
          (
          <issue>2</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          ,
          <year>February 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>John D. Burger</surname>
          </string-name>
          , John Henderson, George Kim, and
          <string-name>
            <given-names>Guido</given-names>
            <surname>Zarrella</surname>
          </string-name>
          .
          <article-title>Discriminating gender on twitter</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP '11</source>
          , pages
          <fpage>1301</fpage>
          -
          <lpage>1309</lpage>
          , Stroudsburg, PA, USA,
          <year>2011</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Fermin</given-names>
            <surname>Cruz</surname>
          </string-name>
          , Rafa Haro, and
          <string-name>
            <given-names>Javier</given-names>
            <surname>Ortega</surname>
          </string-name>
          .
          <source>ITALICA at PAN</source>
          <year>2013</year>
          :
          <article-title>An Ensemble Learning Approach to Author Profiling-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Maria</given-names>
            <surname>De-Arteaga</surname>
          </string-name>
          , Sergio Jimenez, George Duenas, Sergio Mancera, and
          <string-name>
            <given-names>Julia</given-names>
            <surname>Baquero</surname>
          </string-name>
          .
          <article-title>Author Profiling Using Corpus Statistics, Lexicons and Stylistic Features-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Andres</given-names>
            <surname>Alfonso Caurcel Diaz</surname>
          </string-name>
          and Jose Maria Gomez Hidalgo.
          <article-title>Experiments with SMS Translation and Stochastic Gradient Descent in Spanish Text Author Profiling-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Delia</given-names>
            <surname>Irazu Hernandez Farias</surname>
          </string-name>
          , Rafael Guzman-Cabrera, Antonio Reyes, and Martha Alicia Rocha.
          <article-title>Semantic-based Features for Author Profiling Identification: First insights-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Lucie</given-names>
            <surname>Flekova</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <article-title>Can We Hide in the Web? Large Scale Simultaneous Age and Gender Author Profiling in Social Media-Notebook for PAN at CLEF 2013</article-title>
          . In Forner et al. [
          <volume>10</volume>
          ].
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>