<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MAPonSMS'18: Multilingual Author Pro ling using Combination of Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ramsha Imran</string-name>
          <email>ramsha.rabbiya@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muntaha Iqbal</string-name>
          <email>muntaha.iqbal@vu.edu.pk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>COMSATS University Islamabad, Lahore Campus</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Virtual University of Pakistan</institution>
          ,
          <addr-line>Lahore</addr-line>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Author pro ling is used to identify di erent personality traits of an author from the given text. Di erent applications of author pro ling are marketing, security, forensics, spam detection etc. In this paper, we describe the methodology used in author pro ling for the FIRE'18MAPonSMS shared task. Our main objective is to identify age and gender of an author from the multilingual (English and Roman Urdu) text. We incorporated di erent style, vocabulary and emoticon based features to build our proposed system using the training data provided by the task organizers. Furthermore, we used a combination of di erent extracted features and investigated di erent ML algorithms (i.e. Random Forest, Nave Bayes and Support Vector Machines) for the classi cation problem. Results show that our proposed system achieved 73% accuracy in gender identi cation and 53% accuracy in age prediction tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Gender Identi cation</kwd>
        <kwd>Age Prediction</kwd>
        <kwd>Author Pro ling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Author pro ling task is to identify di erent personality traits of an author i.e.
age, gender, personality type, native language etc. from their social dialect. In
other words, author pro ling is to detect di erent attributes of an author's
personality from the written text. The task can be further categorized into
monolingual and multilingual author pro ling. In monolingual, the pro le text is in one
language while in multilingual, the pro le text is in more than one language.
Author pro ling has vast applications in marketing, security and for the detection
of fake pro les.</p>
      <p>In this research work, we used style, vocabulary and emoticon based features
to detect age and gender of a person from a multilingual author pro ling corpus
provided for the MAPonSMS'18 task. Our major contribution are as follows:
{ We used 83 distinct features from the multilingual author pro ling corpus,
grouped into three main categories (1) style based features (2) vocabulary
based features and (3) emoticon based features.
{ We used an ensemble approach to combine di erent features for the classi
cation task.
{ For classi cation, we investigated three Machine Learning (ML) algorithms
i.e. Random Forest (RF), Support Vector Machine (SVM), and Nave Bayes
(NB).
{ We performed better than the baseline and achieved 73% accuracy in
gender identi cation, 53% accuracy in age prediction and an overall 38% joint
accuracy.</p>
      <p>The rest of paper is organized as follows. Section 2 summarizes existing work
on author pro ling. Section 3 explains the multilingual corpus and its statistics.
We then present the proposed system and details of the three types of features in
Section 4. Section 5 discusses results and their analysis and Section 6 concludes
the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>With the evolution of mobile technology, the Short Message Service (SMS) is
considered as one of the most widely used source of communication. Extracting
useful information from SMS data is very crucial task and it attracted many
researches.</p>
      <p>
        In recent past, a lot of work has been done on author pro ling using data
from social media i.e. Twitter and Facebook data, but less work has been done
on multilingual author pro ling using SMS. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], authors proposed stylistic and
content based features for multilingual author pro ling using Facebook data.
      </p>
      <p>
        Argamon et al. in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] also investigated the task of author pro ling and
categorized the features into two main classes (1) style based features and (2) content
based features.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], authors used dictionary of emotions (consists of 344 entries),
dictionary of contractions (consists of 65 entries) and dictionary of unrecognized or
misspelled words for the task of author pro ling.
      </p>
      <p>
        Hernandez et al. in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed linguistic markers, slangs, emotions and
semantic similarity for the identi cation of author's age and gender from the
multilingual text.
      </p>
      <p>
        Cruz et al. in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], De-Arteaga et al. in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Flekova and Gurevych in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Lim
et al. in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Meina et al. in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Patra et al. in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Santosh et al. in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] make
use of di erent stylistic features for the task of multilingual author pro ling.
      </p>
      <p>
        As can be seen from the above discussion, researchers have used Facebook
and Twitter data for the author pro ling task but less work has been done
for multilingual author pro ling using SMS data. Fatima et al. in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], recently
developed a corpus for multilingual author pro ling using SMS data called
SMSAP-18. SMS-AP-18 corpus consist of a total 810 pro les of authors with di
erent demographics i.e age, gender, native language native city, personality type,
quali cation and occupation. From these 810, 610 are male pro les (with 64,968
messages) and 200 female pro les (with 19,726 messages). These pro les were
categorized in three age groups i.e. 15-19, 20-24 and 25-xx. They used 64
stylistics features and 12 content based features and reported 97.5% joint accuracy
for gender identi cation task.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Corpus</title>
      <p>
        For the task of FIRE'18-MAPonSMS organizers provided a total of 500 author
pro les from the SMS-AP-18 corpus [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Out of these 500 pro les, 350 were for
training and 150 for testing. The training data was distributed into two gender
groups i.e. male and female, and three age groups i.e. 15-19, 20-24 and 25-xx.
The training data consists of 141 female and 209 male pro les. Moreover, 108
pro les were from age group of 15-19, 176 from age group of 20-24 and 66 were
from age group of 25-xx. The test data contains 150 pro les, which was provided
by the organizers to test our system and report accuracies.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>System Description</title>
      <p>We propose an author pro ling system to detect age and gender from the
multilingual text for the MAPonSMS'18 task. Our system has three main steps 1)
feature extraction, 2) combination of features and 3) classi cation.
4.1</p>
      <sec id="sec-4-1">
        <title>Feature Extraction</title>
        <p>In the feature extraction step, we extracted three types of features i.e. style
based features, vocabulary based features and emoticon based features. Using
these three types, a total of 83 features were extracted 3. In the following sections,
we describe each of the feature types in detail.</p>
        <p>Style based Features We used a total of 24 style based features which were
extracted from the training corpus. This set of 24 features has 17 special characters
features and 7 count of occurrences features. The details are given below.
{ number of tokens
{ number of distinct words
{ average words per line
{ number of lines
{ number of capital letters
{ number of small letters
{ number of digits used
{ number of commas [,]
{ number of full stops [.]
{ number of commercial at [@]
{ number of start braces [(]
{ number of end braces [)]
{ number of exclamation marks [!]
{ number of dashes [-]
3 Before extracting these features, the text was preprocessed by converting it into
lower case.
{ number of question marks [?]
{ number of percentages [%]
{ number of ampersands [&amp;]
{ number of underscores [ ]
{ number of hashes [#]
{ number of equal-tos [=]
{ number of colons [;]
{ number of semicolons [;]
{ number of spaces [] ]
{ number of forward slashes [/]
Vocabulary Based Features We also used a number of vocabulary based
features i.e. abbreviations, academic terms, contractions and slang words. Each
of these are explained in the next sections.</p>
        <p>Word
aoa
w.s
w8
ty
tc
msg
y
ia
os
lol
wow
pls
btw
Abbreviations Abbreviations are de ned as shortened form of a word. In SMS
conversations, people prefer to use short form of words. Considering this fact, we
used a list of Roman Urdu and English abbreviations to extract 14 features from
the training corpus. As the corpus data comes from the SMS messages, people
often use di erent forms of a single word (abbreviations) in SMS communication.
Therefore, we further normalized these word variations to their single form. For
example, \aoa" is abbreviated in many ways i.e. \a.o.a", \a.a.w.w" and \a.w".
We replaced all of these variants with a single abbreviation \aoa".</p>
        <p>Table 1 lists the abbreviations along with their English translations that we
extracted from the training corpus.</p>
        <p>Academic terms We also used the academia related terms as features for the
classi er. Academic terms such as mam (madam), sir, assignment, class, quiz,
o ce, teacher, test, uni (university) and clg (college) were extracted from the
training corpus. Similar to the abbreviations (See Section 4.1) di erent spelling
variants of a word were normalized to their single word form e.g. mam was used
for ma'am, maam and ma'm.</p>
        <p>Contractions Contractions are shorten form of words excluding some internal
letters. We used the most commonly used 15 contractions as features from the
training corpus. The list of contractions is shown below.</p>
        <p>{ I'm { I am
{ you're { You are
{ we're { We are
{ they're { They are
{ Who're { Who are
{ I'll { I shall
{ She'll { She will
{ It'll { It will
{ Can't {Can not
{ Don't { Do not
{ Isn't { Is not
{ Aren't { Are not
{ Doesn't { Does not
{ Wasn't { Was not
{ Didn't { Did not
Slang Words Slangs are those informal words that people use in their
conversation especially on the Web and di erent social media forums. We developed a
small dictionary of these commonly used slang words in Roman Urdu language
and their English translations. Some of the examples from the dictionary are
shown in below.</p>
        <p>{ chuss { Unpleasant
{ scene { Situation
{ chill { Enjoyment
{ pagal { Mad
{ miss { Remembering
{ mood { Temper
{ burger { Impersonate
{ bhai { Brother
{ drama { Acting
{ t { Healthy
{ chawal { Stupid
{ yar { Friend
{ dfa { Get lost</p>
        <p>It should be noted that our dictionary also converts di erent variations of
these Roman Urdu words into their normalized forms.</p>
        <p>Emoticon based Features: Emoticons are small images or icons that are
used to express moods, expressions and feelings in day-to-day communication.
We collected seven main types of emoticons. The following list shows all types
of emoticons extracted from the training corpus.</p>
        <p>{ Happy { :)
{ Sad { :(
{ Cry { :'(
{ Unsure { :/
{ Squint { -
{ Kiss { :*
{ Wink { ;)
4.2</p>
        <p>Combination of Features
{ All style features: In this combination set, we combined all style based
features together (See Section 4.1). The set consists of 24 features which
were used together to train the classi er for prediction of age and gender.
{ All vocabulary features: For the second set, we combined all vocabulary
based features (See Section 4.1). The set consists of 14 abbreviations, 10
academic terms, 13 slang words and 15 contractions.
{ All emoticon features: The third set is the combination of all emoticon
features which are 7 in number (See Section 4.1).
{ All features combined: In this combination set, we combined all style,
vocabulary and emoticon features together for the classi cation task. The
combination set contains all 83 features.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Classi cation:</title>
        <p>For the classi cation task of predicting age and gender from the multilingual
text, we explored di erent ML classi cation algorithms including Random Forest
(RF), Support Vector Machine (SVM) and Nave Bayes (NB). The feature sets
from Section 4.2 were provided as input to train the classi er.
Based on the feature sets (See Section 4.2) and ML algorithms (See Section 4.3),
we trained separate models for both gender identi cation and age prediction
tasks. We applied 10-fold cross validation and evaluated the performance on the
basis of accuracy.</p>
        <p>From Table 2 it can be seen that RF classi er has outperformed in gender
identi cation with `All features' set with the highest accuracy of 78.57%. This
shows that combining all distinct features has resulted in better predicting
gender of an author from the multilingual text. The gender prediction also highlights
that the max contribution to the best accuracy is by 'Style' based features
(accuracy = 75%). On the other hand, for age prediction, SVM performs overall
best with the highest accuracy of 55.42% again using `All features'.
Furthermore, all the three individual sets of features used in the age prediction task has
reported almost the same results. Conclusively, we can say that the ensemble
approach that we have used to combine the di erent individual features together
has resulted in the better performance of di erent classi ers.</p>
        <p>We submitted our system to the MAPonSMS'18 task and it was evaluated
on the test dataset. Our system performed better than the baseline with the
achieved results are 73% accuracy in gender identi cation task while 53%
accuracy in age prediction task.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper we used style, vocabulary and emoticon based features for gender
identi cation and age prediction tasks on the multilingual author pro ling corpus
provided by the MAPonSMS'18 task. We trained a number of ML classi ers i.e.
RF, SVM and NB using di erent combinations of the above mentioned sets
of features. Though our system (gender = 73.57%, age = 53.42% and joint
38%) beat the baseline accuracies (gender = 0.60, age = 0.51 and joint 0.32), in
comparison to the top rank system, there is still much room for improvement.
If the task is o ered in the future, we will explore more sophisticated methods
to improve our results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Fatima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anwar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
          </string-name>
          , R.-M.
          <article-title>-</article-title>
          A.:
          <article-title>Multilingual author pro ling on Facebook</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>53</volume>
          (
          <issue>4</issue>
          ),
          <volume>886</volume>
          {
          <fpage>904</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.-W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,:
          <article-title>Automatically pro ling the author of an anonymous text</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>52</volume>
          (
          <issue>2</issue>
          ),
          <volume>119</volume>
          {
          <fpage>123</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Aleman</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loya</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ayala</surname>
            ,
            <given-names>D.</given-names>
            -V, Pinto, D.
          </string-name>
          :
          <article-title>Two Methodologies Applied to the Author Pro ling Task</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          , Citeseer, (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hernandez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guzman-Cabrera</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reyes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rocha</surname>
          </string-name>
          , M.-A.:
          <article-title>Semantic-based Features for Author Pro ling Identi cation: First insights</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>F.-L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rafa</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>-R</surname>
          </string-name>
          .,
          <string-name>
            <surname>Ortega</surname>
            ,
            <given-names>F. J.</given-names>
          </string-name>
          :
          <source>ITALICA at PAN</source>
          <year>2013</year>
          :
          <article-title>An Ensemble Learning Approach to Author Pro ling</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>De-Arteaga</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mancera</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baquero</surname>
          </string-name>
          , J.:
          <article-title>Author pro ling using corpus statistics, lexicons and stylistic features</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Flekova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Can we hide in the web? large scale simultaneous age and gender author pro ling in social media</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lim</surname>
          </string-name>
          , W.-Y.,
          <string-name>
            <surname>Goh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thing</surname>
          </string-name>
          , V.-L.:
          <article-title>Content-centric age and gender pro ling</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Meina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brodzinska</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Celmer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Czokow</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pezacki</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Ensemble-based classi cation for author pro ling using various features</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Patra</surname>
            ,
            <given-names>B.-G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saikh</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bandyopadhyay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Automatic author pro ling based on linguistic and stylistic features</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Santosh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bansal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shekhar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varma</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Author pro ling: Predicting age and gender from blogs</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          , (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Fatima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anwar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naveed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arshad</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
          </string-name>
          , R.-M.
          <article-title>-</article-title>
          A.,
          <string-name>
            <surname>Iqbal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masood</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Multilingual SMS-based author pro ling: Data and methods</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>1</volume>
          {
          <fpage>30</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>