<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Bag-of-Words and Psycho-Linguistic Features For MAPonSMS?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Asmara Safdar</string-name>
          <email>asmarasafdar@cuilahore.edu.pk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Osama Akhter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Osama Inayat</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abdullah Khalid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bag-of-Words</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Psycho-</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>COMSATS University Islamabad, Lahore Campus</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the use of Bag-of-Words( BoW) and Psycho-Linguistic( P-L) approaches based upon the demographic trends in modeling multilingual( Roman-Urdu and English) SMS text( Short Message Service) for gender and age prediction. The data set1 was provided as a standard source to work for the multilingual author pro ling task in the contest FIRE'18-MAPonSMS2. The proposed approaches, as compared to the baseline results, adequately classify the test set to age and gender separately.</p>
      </abstract>
      <kwd-group>
        <kwd>Author pro ling Linguistic</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Authorpro ling is a task in automatic authorship identi cation that nds
characteristics, particularly: demographic, of the author of a document. Having known
the pro le of an author can help in resolving many issues, such as, crime
investigation( e.g., by identifying the linguistic pro le of a suspected message),
developing a recommendation system to recommend di erent products to di
erent users( by nding the demographic features of authors through his reviews
on a product) etc.</p>
      <p>MAPonSMS task is about the prediction of gender and age of the authors
based on the multilingual i.e. English and Roman-Urdu SMS data set of 350
documents.</p>
      <p>
        In this paper, we present our proposed approaches and their contribution
towards the classi cation of gender and age for the contest. We have applied some
stylometric( lexical, syntactic, structural features) and content based( based
upon the textual content rather than the metadata) approaches to perform the
task. In stylometric approaches, features based upon Psycho-Linguistic have
proposed as taken from [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and we refer such features as P-L in this write-up. As
? https://lahore.comsats.edu.pk/cs/MAPonSMS/index.html
1 https://lahore.comsats.edu.pk/cs/MAPonSMS/de.html
2 \Forum for Information Retrieval Evaluation-Multilingual Author Pro ling on SMS"
https://lahore.comsats.edu.pk/cs/MAPonSMS/index.html
content based features, we have devised some sets of Psycho-Linguistic
content based words as a representation of the document. We call these features
as Psycho-Linguistic Bag-of-Words features and write as P-BoW in the whole
discussion that follows. The proposed features outperformed the baseline
accuracy results both for age and gender prediction. The software submitted for the
contest can be downloaded from https://github.com/Osama081/MapOnSMS.
      </p>
      <p>In the sections those follow, we present the related work in authorpro ling in
section 2. In section 3 we describe the approaches we used to generate features
of the given dataset. Section 3.4 provides an overview of the results for both
prediction tasks separately. Section 4 concludes the paper and suggests potential
improvements.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Literature Survey</title>
      <p>
        In the literature-world of automatic authorpro ling, a lot of work has performed
on datasets mainly collected from social media sites and blogs while SMS, as
data-source, remains neglected. Besides, the multilingual datasets are in the
languages which are spoken in developed countries while a little work( such as
given by Fatima et al. in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) has done on multilingual datasets with
RomanUrdu as one of the languages.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Chen et al. did their studies based on a dataset of gendered usage of
emojis containing 134,419 Android-smartphone users across 183 countries, in 58
languages to analyze various aspects of emoji usage and nd out that the people
of di erent genders tend to use emojis of slightly di erent categories.
      </p>
      <p>
        Fatima et al. in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] provide a standard multilingual resource of 810 SMS
based user pro les annotated with 7 demographic traits including age and
gender. They have applied stylometric and content based features for the gender
identi cation task. However, the classi cation of other demographic features has
not yet explored for the corpus.
      </p>
      <p>
        Cheng et al. propose Psycho-Linguistic and gender preferential cues along
with stylometric features for gender prediction in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. They performed the
experiments for short length, multi-genre, content-free English text. The empirical
studies show an accuracy up to 85.1%.
      </p>
      <p>
        In third authorpro ling task at PAN 2015 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Rangel et al. organized the
tasks of age, gender, and personality recognition. The given dataset was collected
from Twitter and consisted of English, Spanish, Dutch and Italian languages.
The participants used content-based features( including bag of words and
ngrams) and style-based features ( frequencies, punctuations and some Twitter
speci c such as hash-tags).
      </p>
      <p>
        Rangel et al. in the author pro ling task at PAN 2013 contest [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] describe
the identi cation of age and gender using multilingual dataset, consisting of
English and Spanish, collected from social media sites. The approaches used by
the participants include content-based, stylistic-based, n-grams based, IR-based
and collocations-based features. Empirical studies show the di culty of the task
especially in gender prediction and collective prediction of gender and age.
      </p>
      <p>
        Researchers have been providing pieces of evidence for the last few decades
that a person's physical and mental health is strongly correlated with the words
he/she uses. Gottschalk et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Rosenberg et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] discuss the di erent
factors and theoretical bases of psychological states. We have presented our work
using some existing and rest by modi cation of the existing approaches that
have applied to various types and genre of the dataset in the past. Our proposed
approaches are mainly related to psycholinguistic features on the given SMS
based multilingual dataset in English and Roman-Urdu.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Authorpro le Experiments</title>
      <p>We conducted feature extraction on the dataset provided for age and gender
prediction separately. The results show that both of the proposed approaches
have proved better than the baseline approaches.
3.1</p>
      <sec id="sec-3-1">
        <title>Dataset</title>
        <p>The given dataset was a training set to work for gender and age classi cation
task for the contest FIRE18'-MAPonSMS. It is an SMS based multilingual
corpus containing 350 total documents each from a di erent user. Each document
contains multiple text messages and annotated with age groups and the gender
such that the instances of all groups are balanced. For gender classi cation, it
has 60% and 40% documents written by male and female authors respectively.
For age classi cation, there are 31% documents categorized in age group 15-19,
50% in age group 20-25 while 19% in the group 25-xx.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Approaches Used</title>
        <p>
          We applied 1) Stylometry and 2) Content-based approaches for gender and age
prediction tasks among the renowned methods for authorpro ling i.e. stylometry,
content-based and topic-based [
          <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Feature Extraction For Gender Classi cation For gender classi cation</title>
        <p>task, we used 67 stylometric features in groups of three namely: character based(
Table1), vocabulary richness( Table 2) and word based( Table3). Under the group
of word based features are introduced some P-L as well. In content based( Table
4) both features are P-BoW.</p>
        <p>
          For character based approaches, there are 43 features in total( as shown in
Table1) most of which have been employed by [
          <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
          ] for prediction of
demographic features of the author.
        </p>
        <p>
          For vocabulary richness, we used total 9 features as given in Table2). These
vocabulary richness features have used by many researchers for age and gender
prediction problem such as in [
          <xref ref-type="bibr" rid="ref14 ref3 ref5 ref8">14, 8, 3, 5</xref>
          ].
        </p>
        <p>Feature Description Feature Description
F1 Total number of all characters(C) F26 Count of underscore
F2-F12 Punctuation marks(. , ? etc less m-dsh3) F27-34 Count of @, &amp;, *, $, =, /, %, + sign
F13 Percentage of punctuation marks to C F35 Count of all sorts of brackets
F14-F15 Opening and closing curly braces F36 Count of white spaces
F16-F17 Opening and closing square brackets F37 Percentage of of white spaces to C
F18-F19 Opening and closing parenthesis F38 Percentage of letters to C
F20-F21 Opening and closing angle brackets F39 Percentage of upper case letters to C
F22 Count of white spaces F39 Percentage of upper case letters to C
F23 Count of vertical lines F41 Percentage of white spaces to C
F24 Count of uppercase letters F42 Percentage of digits to C
F25 Count of digits F43 Percentage of tabs to C
Intuition for proposing word ending with i/I ( F7 in Table3) and word ending
with a/A( F6 in Table3) is the fact that in Roman-Urdu, many words are gender
speci c( that discriminate a masculine noun4 from a feminine noun5). The ending
of a word( things, abstract nouns, participles), if a vowel, usually helps in this
gender classi cation. Words, ending with a are usually masculine whereas if a
word ends with i or ii, it is usually a feminine6 word. For example, \Answer
ne kr saki mein" is written in Roman-Urdu that means \I couldn't answer"( as
written by a female author). The word \saki" means could that is a feminine
version of participle while the same word is written as \saka" if referred by
a male. There are many such words in Roman-Urdu those are used with the
slight change of i and a letter, in the end, to refer to female and male author
respectively. Other examples for such words are \khata-khati"( eat in English),
\ata-ati"( come in English), \karta-karti"( do in English), \sota-soti"( sleep in
English) and so on. Limitation of this approach is the fact that there are many
neutral words( with no gender) that might have been counted as a masculine or
a feminine. Additionally, a male author may refer to many feminine words and
vice versa.</p>
        <p>
          Two features( F9 and F10 in Table3) are related to the use of emojis and
smilies in the SMS document by each user. Emojies are combinations of di
erent characters to express emotions( emojis are also called emoticons, winks or
smileys). Chen et al. claim in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] that women are more likely to use emojis than
men. We got a resource for a number of text-emojis from7. One feature is the
4 all male human beings,animals and plants those are considered \masculine" are
masculine in Roman-Urdu
5 all female human beings,animals and plants those are considered \feminine" are
feminine in Roman-Urdu
6 https://en.wikibooks.org/wiki/Urdu/Nouns
        </p>
        <p>Last Visited: 05, 08, 2018
7 http://cool-smileys.com/text-emoticons</p>
        <p>Last Visited: 05, 08, 2018
count of emojis( F9 in 3) and the other is the average number of emojis per
message( F10 in 3).</p>
        <p>
          As the given multilingual dataset has been generated having collected from
di erent mobile users in Pakistan, another approach for proposing P-L features
is to see the tendency of authors to use English words in the multilingual dataset.
Urdu is Pakistan's o cial language yet English is used equally in o ces especially
for writing many o cial documents. Moreover, many text displays on di erent
banners, billboards, organization's name boards and many other activities use
English. Besides English is the medium of education in almost all of the
educational setups and institutes in Pakistan8. Studies show that some demographic
features a ect the language one uses [
          <xref ref-type="bibr" rid="ref4 ref9">4, 9</xref>
          ]. Keeping this in view, the proportions
of English to Roman-Urdu contents sounds a potential feature to contribute
substantially in predicting the demographic features like age and gender of authors.
To see its e ect on the given classi cation tasks, we proposed 4 such features(
F11-F14 in Table3) for gender classi cation. We used the standard English word
dictionary, used in Linux, as a resource to match the English words in the given
dataset.
We added some P-BoW features as well: 1) Percentage of Assent words to total
words, and 2) Percentage of Negation words to total words given in Table 10.
Such categories of P-BoW are to see the e ect of count-based representation of
the document based on the correlation of the linguistic factors and psychological
aspects of an author. Cheng et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] propose several P-L features to build the
feature space for gender prediction. In our case, where the data set provided is
multilingual, we identi ed the group of some psycholinguistic words as given by
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and added some Roman-Urdu words in the selected categories. One feature
related to P-BoW words is the percentage of assent words to total words. English
assent words we selected are ok, agree, alright, right, yes, yup, yeah. The same
category also included Roman-Urdu words as sai ( ok or alright in English), k9
ok9, han, haan, sae, h9.
8 https://en.wikipedia.org/wiki/Pakistani English
        </p>
        <p>Last Visited: 05,08,2018
9 one or more characters</p>
        <p>
          Second group in P-BoW is the Negation Words, from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. It contains no,
never, not, na ,ni , nae, niii, nahi. We proposed some Roman-Urdu words mostly
used for negation in this category. Note that all non-English( Roman-Urdu)
words in this group are variants of no in English.
        </p>
        <p>As Roman-Urdu lacks standard lexicon, many spelling variations exist for a
given word most of the times. For example, nahi, ni, nae, nii are all variations
of a single word in Roman-Urdu that means no in English. So, it is important
to mention here that any group of these P-BoW is not exhaustive because of the
inherent inconsistency in the representation of the Roman-Urdu text.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Feature Extraction For Age Classi cation For age classi cation task,</title>
        <p>total 75 features were generated out of which 70 are stylometric and 5 are content
based.</p>
        <p>In stylometric features, we introduced total 44 character based features as
listed in Table 5 and 8 vocabulary richness features as given by Table 6.</p>
        <p>
          Feature Description Feature Description
F1 Total number of all characters(C) F25 Percentage of digits
F2-F12 Punctuation marks(. , ? etc less m-dsh10) F26-36 Count of @, &amp;, *, $, =, /, %, +,
F13 Percentage of punctuation marks to C F37 Count of all sorts of brackets
F14-F15 Opening and closing curly braces F38 Count of white spaces
F16-F17 Opening and closing square brackets F39 Percentage of letters to C
F18-F19 Opening and closing parenthesis F40 Count of upper case letters to C
F20-F21 Opening and closing angle brackets F41 Percentage of digits to C
F22 Percentage of white spaces F42 Percentage of tabs to C
F23 Count of vertical lines F43 Count of tabs
F24 Percentage of uppercase letters to C F44 Percentage of special characters to C
Percentage of certainty( F5) given in Table 8. As studies show that age strongly
a ects the use of language [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], knowing this, we proposed the feature of slang(
F3 and F4 in Table8). We categorized slang words as the words those don't
relate either to English or Roman-Urdu and are not used in formal speaking
or writing. 16 such words were identi ed from the dataset. Some of the slang
words we selected are lol, plz, btw, k, idk, jigar, oye, oyee, yar, yr. We identi ed
a few words of certainty( F5 in Table8) having taken idea from [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Words
of certainty12, that we selected, include always, hamesha, hmesha, never, ever,
kabi, kabhi, kbhi, kbi, forever.13
        </p>
        <p>Table 8. Content based features for age classi cation .</p>
        <p>Feature Description Feature Description
F1 Percentage of assent(P-BoW) F4 Percentage of slang(P-BoW)
F2 Percentage of negation(P-BoW) F5 Percentage of certainty(P-BoW)
F3 Count of slang(P-BoW)
3.3</p>
      </sec>
      <sec id="sec-3-5">
        <title>Classi ers Used</title>
        <p>As the training dataset was annotated, the gender and age prediction tasks are
supervised machine learning problems with Gender prediction a binary classi
cation task( class attributes as male or female) whereas Age prediction a multiple
classi cation task( class attributes as 15-19, 20-24, or 25-xx ). We used two
classi ers: Random Forest( RF) and Meta Bagging( MB) from Meta class. Both
of these algorithms are ensemble machine learning algorithms and are closely
related. Note that Bagging was used with its default settings and REP Tree as
its component classi er14.</p>
        <p>We used 10-fold cross validation to evaluate the prediction models and
reported accuracy as a measure to evaluate the performance because the dataset
is balanced. Accuracy is the percentage ratio of correctly classi ed instances to
incorrectly classi ed instances.
3.4</p>
      </sec>
      <sec id="sec-3-6">
        <title>Results and Analysis</title>
        <p>Gender Classi cation Table 5 shows the accuracy measure of di erent groups
of features as reported by RF and MB. MB and RF gave accuracies of 60.8% and
52.85% for All P-BoW features respectively. All P-L and P-BoW features
combined gave 66.85% accuracy for MB and 65.7% for RF. All word based features
collectively gave an accuracy of 70.2% by MB and 71.4% by RF. All
character based features combined resulted in 78.28% accuracy for MB and 77.14% for
RF. Then the fth set of features i.e. combination of all features gave the highest
accuracy of 80.29% by MB.
12 \Certainty is something that is certain or sure"</p>
        <p>https://www.merriam-webster.com/dictionary/certainty Last Visited: 05, 08, 2018
13 P-BoW features of certainty and use of slangs proved more useful to discriminate
age groups than the gender.
14 Although we also selected other Decision Tree algorithms as component classi ers
for Bagging but REP Tree gave the best results.</p>
        <p>It is evident that overall results were best reported by MB as 80.29% for
combination of all features. The best group of features, if analyzed group-wise,
was character based that gave an accuracy of 78.28% for MB. There is also an
interesting fact about using P-L and P-BoW features combined. There were total
8 such features for gender classi cation which collectively gave an accuracy of
66.85% with MB. Of these 8 features, 4 were related to use of English or
RomanUrdu words for which RF gave 65.7% accuracy. This shows that all P-BoW and
P-L approaches contributed substantially to train gender classi er.</p>
        <p>Table 9. Gender classi cation results using di erent groups of features.
Feature Classi er Accuracy Feature
All P-BoW MB 60.8% All character based</p>
        <p>RF 52.85%
All P-L and P-BoW MB 66.85%</p>
        <p>RF 65.7%
All word based 15 MB 70.2%</p>
        <p>RF 71.4%</p>
        <p>Classi er Accuracy
MB 78.28%</p>
        <p>RF 77.14%
All features combined MB 80.29%</p>
        <p>RF 78.57%
Age Classi cation The accuracy scores for di erent sets of features with the
names of classi ers are given in Table10 for Age classi cation. MB and RF gave
accuracies of 49.14% and 44.85% for All P-BoW features respectively. All
PBoW and P-L is the group of 10 psycholinguistic features. This combined group
showed an accuracy of 55.7% for RF and 53.7% by MB. Then the combination
of all word based features gave an accuracy result of 58% and 53.7% for RF
and MB respectively. Group of all character based features when combined gave
accuracy results of 55.14% by RF and 50.57% for MB. The combination of all
stylometric and content based features gave the maximum accuracy of 60% for
RF and 52% by MB</p>
        <p>For age classi cation, results show that RF performed better than MB. All
P-L and P-BoW features gave a combined accuracy of 55.7% for RF that is
greater than 55.14% - the accuracy reported by RF for all character based( 44
in total) features combined. This shows a strong contribution of the P-L and
P-BoW approaches( total 12 such features). We can infer from the results that
best results( i.e., 60%) are reported by the set of all features combined using RF
Decision Tree algorithm.</p>
        <p>The accuracy values given by the classi ers for any set of features could not
go beyond 60% for age classi cation, unfortunately. One reason of the proposed
features for not being able to classify the age groups adequately can be due to the
fact that the division of age groups is so closely related( 15-19, 20-24, 25-xx) in
terms of many demographic traits such as education level, income background
and even the type of educational institute( university) that the authors have
many overlapping traits.</p>
        <p>Table 10. Age classi cation results using di erent groups of features.
Feature Classi er Accuracy Feature Classi er Accuracy
All P-BoW RF 44.85% All character based RF 55.14%</p>
        <p>MB 49.14% MB 50.57%
All P-L and P-BoW RF 55.7% All word based16 RF 58%</p>
        <p>MB 53.7 % MB 53.7%
All features combined RF 60%</p>
        <p>MB 52%</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion And Future Work</title>
      <p>
        In this paper, we presented our approaches for gender and age prediction tasks
on the training data that is SMS based multilingual dataset containing
English and Roman-Urdu text of 350 documents each from a di erent user. We
used stylometric and content based approaches to extract the features and
reported 80.29% accuracy for gender and 60% accuracy for age prediction task
for MAPonSMS contest. The trained prediction models, when used to
predict the test set containing 150 multilingual documents, outperform the
baseline approaches. The improvement in accuracy for gender prediction goes from
0.60%( baseline) to 0.69% and for age prediction from 0.51%( baseline) to 0.53%.
The joint result accuracy improvement is from 0.32%( baseline) to 0.35%. 17
The task of authorpro ling for multilingual text displays a great cushion for
improvement, particularly for the gender classi cation task. To improve the
results some approaches, such as, preprocessing the dataset to normalize the
Roman-Urdu text( as discussed by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]), introducing topic based features [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
and devising methods for word-sense dis-ambiguity to di erentiate English and
Roman-Urdu text can be implied.
17 https://lahore.comsats.edu.pk/cs/MAPonSMS/de.html
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alvarez-Carmona</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez-Monroy</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes-y Gomez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Villasen~orPineda,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Meza</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          :
          <article-title>Evaluating topic-based representations for author pro ling in social media</article-title>
          .
          <source>In: Ibero-American Conference on Arti cial Intelligence</source>
          . pp.
          <volume>151</volume>
          {
          <fpage>162</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ai</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mei</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Through a gender lens: Learning usage patterns of emojis from large-scale android users</article-title>
          .
          <source>In: Proceedings of the 2018 World Wide Web Conference on World Wide Web</source>
          . pp.
          <volume>763</volume>
          {
          <fpage>772</fpage>
          .
          <string-name>
            <surname>International World Wide Web Conferences Steering Committee</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. Cheng, N.,
          <string-name>
            <surname>Chandramouli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbalakshmi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Author gender identi cation from text</article-title>
          .
          <source>Digital Investigation</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <volume>78</volume>
          {
          <fpage>88</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Collier</surname>
            ,
            <given-names>V.P.</given-names>
          </string-name>
          :
          <article-title>Age and rate of acquisition of second language for academic purposes</article-title>
          .
          <source>TESOL quarterly 21(4)</source>
          ,
          <volume>617</volume>
          {
          <fpage>641</fpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fatima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anwar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naveed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arshad</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NAWAB</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.M.A.</given-names>
            ,
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Masood</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Multilingual sms-based author pro ling: Data and methods</article-title>
          .
          <source>Natural Language</source>
          Engineering pp.
          <volume>1</volume>
          {
          <issue>30</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gottschalk</surname>
            ,
            <given-names>L.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gleser</surname>
            ,
            <given-names>G.C.</given-names>
          </string-name>
          :
          <article-title>The measurement of psychological states through the content analysis of verbal behavior</article-title>
          . Univ of California Press (
          <year>1969</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Hu aker,
          <string-name>
            <given-names>D.A.</given-names>
            ,
            <surname>Calvert</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.L.</surname>
          </string-name>
          :
          <article-title>Gender, identity, and language use in teenage blogs</article-title>
          .
          <source>Journal of computer-mediated communication 10</source>
          (
          <issue>2</issue>
          ),
          <source>JCMC10211</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kubat</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milicka</surname>
          </string-name>
          , J.:
          <article-title>Vocabulary richness measure in genres</article-title>
          .
          <source>Journal of Quantitative Linguistics</source>
          <volume>20</volume>
          (
          <issue>4</issue>
          ),
          <volume>339</volume>
          {
          <fpage>349</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gravel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trieschnigg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meder</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>" how old do you think i am?" a study of language and age in twitter</article-title>
          .
          <source>In: ICWSM</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inches</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the author pro ling task at pan 2013</article-title>
          .
          <source>In: CLEF Conference on Multilingual and Multimodal Information Access Evaluation</source>
          . pp.
          <volume>352</volume>
          {
          <fpage>365</fpage>
          .
          <string-name>
            <surname>CELCT</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Rangel</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Celli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          :
          <article-title>Overview of the 3rd author pro ling task at pan 2015</article-title>
          . In:
          <article-title>CLEF 2015 Evaluation Labs</article-title>
          and Workshop Working Notes Papers. pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rosenberg</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tucker</surname>
            ,
            <given-names>G.J.:</given-names>
          </string-name>
          <article-title>Verbal behavior and schizophrenia: The semantic dimension</article-title>
          .
          <source>Archives of General Psychiatry</source>
          <volume>36</volume>
          (
          <issue>12</issue>
          ),
          <volume>1331</volume>
          {
          <fpage>1337</fpage>
          (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sharf</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahman</surname>
            ,
            <given-names>S.U.</given-names>
          </string-name>
          :
          <article-title>Lexical normalization of roman urdu text</article-title>
          .
          <source>INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND NETWORK SECURITY</source>
          <volume>17</volume>
          (
          <issue>12</issue>
          ),
          <volume>213</volume>
          {
          <fpage>221</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fakotakis</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kokkinakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Automatic text categorization in terms of genre and author</article-title>
          .
          <source>Computational linguistics 26(4)</source>
          ,
          <volume>471</volume>
          {
          <fpage>495</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>