<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MAPonSMS - Overview of the Multilingual SMS-based Author Pro ling Task at FIRE'18</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Muhammad Sharjeel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehwish Fatima</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saba Anwar</string-name>
          <email>sabaanwarg@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rao Muhammad Adeel Nawab</string-name>
          <email>adeelnawabg@cuilahore.edu.pk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>COMSATS University Islamabad, Lahore Campus</institution>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the overview of 1st International shared task of Multilingual Author Pro ling on SMS (MAPonSMS) at Forum for Information Retrieval Evaluation (FIRE'18). The aim of the MAPonSMS task is to identify the author's gender and age for a given multilingual (Roman Urdu and English) SMS messages pro le, where each pro le consists of an aggregation of SMS messages from a single author. This paper provides the details of the dataset and its distribution, overview of the submitted approaches and the evaluation framework used for measuring the performance of the submitted multilingual author pro ling systems.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>Multilingual Author Pro ling</kwd>
        <kwd>SMS Corpus</kwd>
        <kwd>Roman Urdu</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Authorship pro ling is a task where the objective is the identi cation of author's
demographic traits (age, gender, native language, etc.) by analyzing the author's
written text. The subject of author pro ling is bene cial in many domains such
as digital forensic analysis [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], marketing intelligence for business [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
sentiment analysis and classi cation for social and physiological behaviors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
pro le of an author can be either: (1) monolingual, or (2) multilingual. In the
former case, the entire author pro le is written in one language, while in the
latter case, a single author pro le will contain text in two or more languages.
Author pro ling on pro le containing text in two or more languages is known as
Multilingual Author Pro ling [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ].
      </p>
      <p>With the advancement of technology and the Internet, people can interact
globally via different mediums (text messaging, social media, blogs, etc.). The
phenomenon of multilingualism emerged due to the communications among
various nationalities having different native languages. Because, people usually use
a common language (like English) for global communication, but somehow, they
have an inclination to their native language(s). Thus, these global connections
have in uenced not only how the languages are being used among different
communities, but also morphing the vocabularies of different languages. Moreover,
multilingualism has also affected the texting trends (SMS messaging, chatting
applications) in the past few years. It might be because a multilingual person
tends to opt vocabulary from multiple known languages during a spontaneous
speech or typing process. So, the research on multilingual text is attracting the
attention of the research community due to the rapid growth of multilingual text.</p>
      <p>The development and evaluation of automatic author pro ling techniques
demand standard evaluation resources in various languages and genres. Because
the selection of language and genre in uences the structure, style and content
of the document. The example of structure based attributes is document length,
sentence length, etc., While the example of style and content based attributes is
vocabulary/ word construction, grammatical forms, punctuation choices, use of
special characters/ emojis, etc. The impact of language and genre can be
comprehended with the following cases. The SMS messages are considered short,
having informal language, including slang and emojis. While, Facebook posts
are regarded as a different genre (length can be short to long), also having
informal language containing emojis. On the other hand, a book chapter or scienti c
article is classi ed as a different genre usually having long to very long document
length where language is formal having dense vocabulary with proper grammar
form. Due to this, the feature extraction process from different genres and
languages would be very important for the training of the author pro ling systems.
Therefore, the selection of language and genre is very important because it can
affect the robustness of the author pro ling system.</p>
      <p>
        In previous research studies, different genres (Twitter, Facebook, blogs, web
forums) have been considered for mostly English and other European languages
in monolingual setting [
        <xref ref-type="bibr" rid="ref24 ref32 ref38 ref4">4, 24, 38, 32</xref>
        ]. The SMS genre has been neglected for
author pro ling regardless of its global popularity, ease of use and access. The
most probable reason of this negligence is its challenging and time consuming
collection as a standard resource. In short, the problem of author pro ling has
not been thoroughly explored neither for South Asian languages (particularly
Urdu and Roman Urdu) nor for SMS genre. Therefore, this competition focuses
on multilingual (English and Roman Urdu) SMS-based author pro ling.
      </p>
      <p>The aim of MAPonSMS-Fire'18 (Multilingual Author Pro ling on SMS)
shared task is to identify the author's gender and age for a given multilingual
(Roman Urdu and English) SMS messages pro le, where each pro le consists of
an aggregation of SMS messages from a single author.</p>
      <p>The rest of the paper is organized as follows. Section 2 discusses the existing
work that has been done on SMS corpora and author pro ling. Section 3 gives
the details of train and test datasets, evaluation measure used to evaluate the
performance of submitted systems, and system submission process. Section 4
describes the overview of submitted systems. Section 5 presents the results and
analysis of submitted approaches. Finally, Section 6 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Although, collecting SMS messages for creating a standard evaluation resource
is a very challenging task, however, few efforts have been made in developing
datasets by using SMS messages for various tasks including SMS text
normalization [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], linguistic[
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], machine translation systems [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] and spam detection
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Among the existing SMS-based corpora, NUS SMS corpus1 is the largest
and most widely used SMS based dataset, which was initially developed to
improve the predictive text in mobile devices [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Its rst version was released in
2004 having English SMS messages and the second version came out in 2010
with an increase in the size of corpus as compared to the rst release. The nal
corpus released in 2013, consisted of two sub-corpora for English and Chinese
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, not all the pro les were associated with demographic information
because many people shared only messages. The NUS SMS corpus has also been
used for forensic authorship analysis task [13{15], authorship detection [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and
author identi cation [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        To date, the PAN competitions provide a major contribution of benchmark
monolingual corpora for identifying different author traits, particularly age and
gender, in various languages and genres. In the 2013 PAN competition, English
and Spanish blog posts were collected for monolingual age and gender prediction
tasks [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. In the 2014 PAN competition, four genres (hotel reviews, tweets, social
media and blogs) in English and Spanish were considered for monolingual age
and gender prediction tasks [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. In the 2015 PAN competition, tweets were
collected in four different languages, including English, Spanish, Italian and Dutch
for monolingual personality trait detection, age and gender prediction [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. In
2016, PAN competition task shifted from same genre author pro ling to cross
genre author pro ling in monolingual setting. [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. The train and test datasets
of PAN 2014 were merged for this year competition. The training was carried
out on tweets, and the test dataset constituted of blogs, social media and hotel
reviews for monolingual age and gender prediction [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. In PAN 2017, the task
was gender and language identi cation for tweets considering four languages
(Arabic, English, Spanish and Portuguese) [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. In PAN competitions from 2014
to 2017, it can be noted that one out of four sub-corpora consisted of tweets.
      </p>
      <p>
        Apart from PAN competitions, some research studies also explored tweet
based datasets for the authorship analysis task, such as author identi cation
[
        <xref ref-type="bibr" rid="ref17 ref20">20, 17</xref>
        ], gender identi cation [
        <xref ref-type="bibr" rid="ref2 ref37">2, 37</xref>
        ]. Few researchers carried out experiments on
the combined datasets of SMS messages and tweets for sentiment analysis [18,
1 http://www.comp.nus.edu.sg/ rpnlpir/downloads/corpora/sms/ Last visited:
22-092018
1]. Although, the construction of a tweet based corpus is quite easy due to
having less privacy concerns and its readily availability, but tweets cannot be an
alternative of SMS genre. It is because, SMS messages are purposely built for
private conversations while tweets are meant for public conversations [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>To summarize, the above mentioned corpora are predominantly monolingual
(for English and other European languages) and are not suitable for South Asian
languages such as Roman Urdu. Moreover, existing SMS and tweet based
corpora are not suitable for the multilingual author pro ling task. Therefore, this
competition addressed the problem by providing a dataset of multilingual SMS
based author pro les for gender and age prediction. We believe that this
competition will foster research on multilingual text (in general) and Roman Urdu
(an under-resourced language) more speci cally.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation Framework</title>
      <p>This section describes the characteristics of the train and test datasets,
performance measure used to evaluate the performance of submitted systems, baseline
approach and the procedure of submissions by the participants.
3.1</p>
      <sec id="sec-3-1">
        <title>Corpus</title>
        <p>
          A subset of SMS-AP-18 corpus [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is used for the rst shared task on Multilingual
Author Pro ling on SMS. The original SMS-AP-18 corpus consists of 810 author
pro les. For the MAPonSMS-FIRE'18 shared task, a subset of 500 author pro les
was selected from the SMS-AP-18 corpus. The reason for selecting a subset of
original corpus is to have a balanced train/test dataset.
        </p>
        <p>Train Dataset The train dataset consists of 350 multilingual (Roman Urdu
and English) SMS based author pro les (see table 1 for detailed statistics). For
gender, a multilingual author pro le may belong to either Male or Female class.
With regard to age, a multilingual author pro le may fall into one of the three
categories: 15{19, 20{24, 25{xx.</p>
        <p>The gender and age information associated with each multilingual author
pro le were stored in a separate truth le which was provided with the train
dataset. All author pro les in the train dataset were stored in the \.txt " format.
Test Dataset The dataset consists of 150 multilingual (Roman Urdu and
English) SMS based author pro les (see table 1 for detailed statistics). The
associated information (age and gender) was unknown for participants. All author
pro les in the test dataset were also stored in the \.txt " format.
The performance of submitted author pro ling systems was computed using
Accuracy measure. Accuracy is de ned as the proportion of correctly classi ed
author pro les.</p>
        <p>Accuracy =</p>
        <p>N o: of Correctly P redicted Author P rof iles</p>
        <p>T otal N o: of Author P rof iles</p>
        <p>We computed Accuracy in two ways: (1) Individual Accuracy of gender and
age traits, and (2) Joint Accuracy of gender and age traits. The submitted
systems were ranked based on Joint Accuracy score.</p>
        <p>Baseline Approach For baseline approach, we used MCC (Majority Common
Category) which is computed by assigning the most common category to all the
instances in the dataset. The MCC of test dataset for: (1) Gender = 0:60, (2)
Age = 0:51 and (3) Joint = 0:32.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Submission</title>
        <p>The participants were asked to submit: (1) Executable multilingual author pro
ling system, (2) Output of the system (predictions) for the test dataset in \.csv "
format for age and gender.</p>
        <p>For multilingual author pro ling system2, some guidelines were provided: (i)
It should be executable generically by commands for both age and gender so that
it can be re-trained on demand for maximizing the sustainability. (ii) It should
predict for each case found in the test corpus and write the output in .CSV le(s)
for both age and gender. The results of multiple runs were not allowed for the
submission.
2 The participants retain the full copyrights of their submitted systems.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Overview of Submitted Systems</title>
      <p>For the rst MAPonSMS competition, a total of 9 submission were received,
however, one of the participating teams did not submit the notebook paper. We
now present the detailed analysis of the 8 approaches we received.
4.1</p>
      <sec id="sec-4-1">
        <title>Preprocessing</title>
        <p>
          Four of the total eight participants that submitted their systems used shallow
text preprocessing before applying methods to extract features from the
multilingual corpus. The authors in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] cleaned the text by removing multiple space
characters, tabs and garbage characters. In [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] only punctuation marks were
removed while in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] only case conversion (lowercasing) was applied during text
preprocessing. The authors of [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] discarded stop words, punctuation marks and
then lowercased the text in the preprocessing step. Four participating systems
[
          <xref ref-type="bibr" rid="ref21 ref23 ref31 ref33">21, 33, 31, 23</xref>
          ] did not use any text preprocessing method.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Feature Extraction</title>
        <p>
          In terms of methods used to extract features from the multilingual corpus,
majority of the submitting systems [
          <xref ref-type="bibr" rid="ref21 ref23 ref31 ref35 ref6 ref9">6, 23, 9, 31, 21, 35</xref>
          ] opted for content based
methods using BoW (Bag of Words) and Tf-Idf (Term frequency - Inverse document
frequency). One of the participating team [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] used language dependent and
independent style based methods. Another team [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] utilized style, vocabulary
and emoticon based methods for feature extraction.
        </p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] used both word and character-based Tf-Idf whereas [
          <xref ref-type="bibr" rid="ref21 ref35 ref9">21,
35, 9</xref>
          ] used only word based Tf-Idf. Moreover, before applying Tf-Idf, [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] rst
normalised the text using a dictionary to translate Roman Urdu words to
English. Another participant, [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] used Tf only and did not consider words with
less than 5 occurrences. Furthermore, the authors in [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] applied a statistical
approach to select the best features out of a large set of generated features.
        </p>
        <p>
          Different style based (e.g. punctuation marks and other symbols, count of
distinct words, words per line, number of lines etc.), vocabulary based (e.g.
abbreviations, academic terms, contractions and slang words) and emoticons
based (i.e. happy, sad, cry, unsure, squint, kiss and wink) features were extracted
by [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. The authors also experimented with different combinations of these three
set of features. A set of stylistic features which are language independent (i.e. avg.
word and sentence length, number long short words and sentences, number of
different punctuation marks) and language dependent (POS-based e.g. number
of adjectives, interjections, nouns etc.) were used by [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ] during the feature
extraction step.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Classi cation</title>
        <p>
          All the participating systems employed supervised ML to identify age and
gender from the multilingual text. Most of them used multiple ML classi ers and
reported the results using the best one(s). In some cases, age was reported with
one classi er while gender with a different one. All the systems submitted for
the task used Support Vector Machines as one of the classi er, however, the
authors in [
          <xref ref-type="bibr" rid="ref23 ref6">6, 23</xref>
          ] used only Support Vector Machines. Apart from [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], all the other
approaches used Random Forest too. Logistic Regression was another favorite
used by 3 [
          <xref ref-type="bibr" rid="ref21 ref33 ref9">9, 33, 21</xref>
          ] submitted systems.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ], the authors experimented with 11 different classi ers i.e., Multinomial
Nave Bayes, Gaussian Nave Bayes, Decision Tree, Random Forest, Extra Trees,
Ada Boost, Gradient Boosting, Support Vector Machines, Stochastic Gradient
Descent, Multi Layer Perceptron and Multinomial Nave Bayes. They reported
best results using Multi Layer Perceptron and Multinomial Nave Bayes. In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ],
Random Forest, Support Vector Machines, Logistic Regression and Nave Bayes
were used, Nave Bayes outperformed others. The authors of [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], tried 3
different classi ers i.e. Random Forest, Nave Bayes and Support Vector Machines.
They showed that Random Forest for gender and Support Vector Machines for
Age performed best. In [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ], the authors went for Random Forest and Meta
Bagging by Decision Tree as its component classi er. In [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ], Nave Bayes, J48,
Random Forest and Logistic Regression were used for the classi cation task. The
authors reported that Random Forest performed best for gender and Logistic
Regression for age. In [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], Logistic Regression, Nave Base, Multi-layer
Perceptron and Gradient Boosting were used. Furthermore, the authors ensemble all
four classi ers to report the best result.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation of the Submitted Systems</title>
      <p>In this section, we discuss the results of the 9 teams that submitted their systems
for the MAPonSMS task. Table 2 shows the age, gender and joint accuracies
obtained by the submitting systems. As can be seen, majority of the systems
performed better than the baseline accuracies. The highest reported accuracies
are with sharmila-18, as they performed best in age, gender as well as joint
accuracy. abdul-18 secured the lowest results and is the only approach that is
below the baseline. Expectedly, as the gender prediction was binary classi cation
task whereas age was multi classi cation, all the team have performed better in
the former. Moreover, the low scores obtained on age classi cation has effected
the joint accuracies as well.</p>
      <p>
        The approach used by sharmila-18 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] outperformed others and achieved
the highest accuracy for the MAPonSMS task. Their team used both word and
character based Tf-Idf features which resulted in its overall best performance.
On the other hand, thenmozhi-18 [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] and ali-18 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] achieved results very close
to the sharmila-18, and they are among the top three. It can be observed that the
top 3 ranked teams have used Tf-Idf for feature extraction from the multilingual
corpus. Contrarily, the teams that utilized stylistic features are ranked last and
3rd last.
      </p>
      <p>The top two teams have used shallow text preprocessing methods before
feature engineering which indicates that text preprocessing have shown positive
impact on the results of the task.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        In this paper, we present [
        <xref ref-type="bibr" rid="ref12 ref21 ref23 ref31 ref33 ref35 ref6 ref9">6, 35, 21, 31, 33, 9, 23, 12</xref>
        ] the results of the 1st
International shared task of Multilingual Author Pro ling on SMS (MAPonSMS) at
FIRE'18. Given a reasonable and realistic collection of SMS messages for age
and gender identi cation with multilingual setting was a challenging task and 9
teams participated in the competition.
      </p>
      <p>
        Participants used several different methods for solving the task such as BoW
(Bag of Words) based Tf-Idf, stylistic, vocabulary and emoticon based features.
Majority of the participating teams performed better than the baseline
accuracies. The highest age, gender, and joint accuracy (0.87, 0.65, and 0.57) was
achieved by sharmila-18 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] by using word and character based Tf-Idf method
and Support Vector Machines.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aboluwarin</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andriotis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takasu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tryfonas</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Optimizing short message text sentiment analysis for mobile device forensics</article-title>
          .
          <source>In: 12th IFIP WG 11.9 International Conference on Advances in Digital Forensics XII</source>
          . pp.
          <volume>69</volume>
          {
          <fpage>87</fpage>
          . Springer, New Delhi, India (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Alowibdi</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buy</surname>
            ,
            <given-names>U.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Language independent gender classi cation on Twitter</article-title>
          .
          <source>In: ASONAM '13: Proceedings of the 2013 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining</source>
          . pp.
          <volume>739</volume>
          {
          <fpage>743</fpage>
          . ACM, Niagara, Ontario, Canada (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Anstead</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>O'Loughlin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Social Media Analysis and Public Opinion: The 2010 UK General Election</article-title>
          .
          <source>Journal of Computer-Mediated Communication</source>
          <volume>20</volume>
          (
          <issue>2</issue>
          ),
          <volume>204</volume>
          {
          <fpage>220</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henderson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarrella</surname>
          </string-name>
          , G.:
          <article-title>Discriminating Gender on Twitter</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <volume>1301</volume>
          {
          <fpage>1309</fpage>
          . Association for Computational Linguistics, Edinburgh, United
          <string-name>
            <surname>Kingdom</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Creating a live, public short message service corpus: the NUS SMS corpus</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>47</volume>
          (
          <issue>2</issue>
          ),
          <volume>299</volume>
          {
          <fpage>335</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Devi</surname>
            <given-names>V</given-names>
          </string-name>
          , S.,
          <string-name>
            <surname>Kannimuthu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Safeeq</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
            <given-names>M</given-names>
          </string-name>
          , A.:
          <article-title>KCE DAlab@MAPonSMSFIRE2018: Effective Word and Character-based Features for Multilingual Author Pro ling</article-title>
          .
          <source>In: Working Notes for MAPonSMS at FIRE'18 - Workshop Proceedings of the 10th International Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fatima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anwar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naveed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arshad</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
            ,
            <given-names>R.M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iqbal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masood</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Multilingual SMS-based author pro ling: Data and methods</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>24</volume>
          (
          <issue>5</issue>
          ),
          <volume>695</volume>
          {
          <fpage>724</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fatima</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anwar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
            ,
            <given-names>R.M.A.</given-names>
          </string-name>
          :
          <article-title>Multilingual author pro ling on Facebook</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>53</volume>
          (
          <issue>4</issue>
          ),
          <volume>886</volume>
          {
          <fpage>904</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gaur</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ayyar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kumar</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.R.</surname>
          </string-name>
          :
          <article-title>Multilingual Author Pro ling from SMS</article-title>
          .
          <source>In: Working Notes for MAPonSMS at FIRE'18 - Workshop Proceedings of the 10th International Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Giannella</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Winder</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>(Un/Semi-)supervised SMS text message SPAM detection</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>21</volume>
          (
          <issue>4</issue>
          ),
          <volume>553</volume>
          {
          <fpage>567</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Glance</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hurst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nigam</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stockton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tomokiyo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Deriving Marketing Intelligence from Online Discussion</article-title>
          .
          <source>In: KDD '05: Proceedings of the Eleventh ACM SIGKDD International Conference on Knowledge Discovery in Data Mining</source>
          . pp.
          <volume>419</volume>
          {
          <fpage>428</fpage>
          . Chicago, Illinois, USA (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Imran</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iqbal</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>MAPonSMS'18: Multilingual Author Pro ling using Combination of Features</article-title>
          . In: Working Notes for MAPonSMS at FIRE'
          <fpage>18</fpage>
          - Workshop Proceedings of the 10th International
          <article-title>Forum for Information Retrieval Evaluation (FIRE'18)</article-title>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ishihara</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A forensic authorship classi cation in sms messages: A likelihood ratio based approach using n-gram</article-title>
          .
          <source>In: Proceedings of the Australasian Language Technology Association Workshop 2011</source>
          . pp.
          <volume>47</volume>
          {
          <fpage>56</fpage>
          .
          <string-name>
            <surname>Canberra</surname>
          </string-name>
          ,
          <string-name>
            <surname>Australia</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ishihara</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A Forensic Text Comparison in SMS Messages: A Likelihood Ratio Approach with Lexical Features</article-title>
          .
          <source>In: WDFIA 2012 : Seventh International Workshop on Digital Forensics &amp; Incident Analysis</source>
          . pp.
          <volume>55</volume>
          {
          <fpage>65</fpage>
          .
          <string-name>
            <surname>Crete</surname>
          </string-name>
          ,
          <string-name>
            <surname>Greece</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Ishihara</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A likelihood ratio-based evaluation of strength of authorship attribution evidence in SMS messages using N-grams</article-title>
          .
          <source>International Journal of Speech, Language &amp; the Law</source>
          <volume>21</volume>
          (
          <issue>1</issue>
          ) (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Juola</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Industrial Uses for Authorship Analysis</article-title>
          .
          <source>In: Mathematics and Computers in Sciences and Industry</source>
          , pp.
          <volume>21</volume>
          {
          <fpage>25</fpage>
          .
          <string-name>
            <surname>INASE</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Kebede</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tefrie</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sohn</surname>
            ,
            <given-names>K.A.</given-names>
          </string-name>
          :
          <article-title>Anonymous Author Similarity Identication</article-title>
          .
          <source>In: 2015 5th International Conference on IT Convergence and Security (ICITCS)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          .
          <string-name>
            <given-names>Kuala</given-names>
            <surname>Lumpur</surname>
          </string-name>
          ,
          <string-name>
            <surname>Malaysia</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kiritchenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>Sentiment analysis of short informal texts</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          <volume>50</volume>
          ,
          <volume>723</volume>
          {
          <fpage>762</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kretchmar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Text Message Authorship Classi cation Using Kernel Support Vector Machines</article-title>
          .
          <source>In: CSCI 2014: International Conference on Computational Science and Computational Intelligence</source>
          . vol.
          <volume>2</volume>
          , pp.
          <volume>215</volume>
          {
          <fpage>218</fpage>
          . IEEE,
          <string-name>
            <surname>Las</surname>
            <given-names>Vegas</given-names>
          </string-name>
          , Nevada, USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Layton</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watters</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dazeley</surname>
          </string-name>
          , R.:
          <article-title>Authorship attribution for twitter in 140 characters or less</article-title>
          .
          <source>In: Cybercrime and Trustworthy Computing Workshop (CTC)</source>
          ,
          <year>2010</year>
          Second. pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Nemati</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gender and Age Prediction Multilingual Author Pro les Based on Comments</article-title>
          .
          <source>In: Working Notes for MAPonSMS at FIRE'18 - Workshop Proceedings of the 10th International Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Oliva</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serrano</surname>
            ,
            <given-names>J.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Castillo</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Igesias</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A SMS normalization system integrating multiple grammatical resources</article-title>
          .
          <source>Natural Language Engineering</source>
          <volume>19</volume>
          (
          <issue>01</issue>
          ),
          <volume>121</volume>
          {
          <fpage>141</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Orts</surname>
            ,
            <given-names>O.G.i.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A statistical approach to gender and age range classication in multilingual corpus</article-title>
          .
          <source>In: Working Notes for MAPonSMS at FIRE'18 - Workshop Proceedings of the 10th International Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Peersman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
          </string-name>
          , W.,
          <string-name>
            <surname>Van Vaerenbergh</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Predicting Age and Gender in Online Social Networks</article-title>
          .
          <source>In: Proceedings of the 3rd International Workshop on Search and Mining User-generated Contents (SMUC'11)</source>
          . pp.
          <volume>37</volume>
          {
          <fpage>44</fpage>
          . ACM, Glasgow, Scotland, UK (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Ragel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herath</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Senanayake</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Authorship detection of SMS messages using unigrams</article-title>
          .
          <source>In: 2013 IEEE 8th International Conference on Industrial and Information Systems (ICIIS</source>
          <year>2013</year>
          ). pp.
          <volume>387</volume>
          {
          <fpage>392</fpage>
          . IEEE,
          <string-name>
            <surname>Sri Lanka</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moshe</surname>
            <given-names>Koppel</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Inches</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          :
          <article-title>Overview of the Author Pro ling Task at PAN 2013</article-title>
          .
          <article-title>In: CLEF 2013 Evaluation Labs</article-title>
          and Workshop { Working Notes Papers. Valencia,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 5th Author Proling Task at PAN 2017: Gender and Language Variety Identi cation in Twitter</article-title>
          .
          <source>In: Working Notes Papers of the CLEF 2017 Evaluation Labs. CEUR Workshop Proceedings, CLEF and CEUR-WS.org</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Overview of the 3rd Author Pro ling Task at PAN 2015</article-title>
          .
          <article-title>In: CLEF 2015 Evaluation Labs</article-title>
          and Workshop { Working Notes Papers.
          <article-title>CEUR-WS</article-title>
          .org, Toulouse, France (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trenkmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daeleman</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Overview of the 2nd Author Pro ling Task at PAN 2014</article-title>
          .
          <article-title>In: CLEF 2014 Evaluation Labs</article-title>
          and Workshop { Working Notes Papers.
          <article-title>CEUR-WS</article-title>
          .org, Sheffield, UK (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 4th author pro ling task at PAN 2016: cross-genre evaluations</article-title>
          . In: Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          . pp.
          <volume>750</volume>
          {
          <fpage>784</fpage>
          . vora - Portugal (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Safdar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akhter</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inayat</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khalid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Using Bag-of-Words and PsychoLinguistic Features For MAPonSMS</article-title>
          . In: Working Notes for MAPonSMS at FIRE'
          <fpage>18</fpage>
          - Workshop Proceedings of the 10th International
          <article-title>Forum for Information Retrieval Evaluation (FIRE'18)</article-title>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Shrestha</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rey-Villamizar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sadeque</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bethard</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solorio</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Age and gender prediction on health forum data</article-title>
          .
          <source>In: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ).
          <source>European Language Resources Association (ELRA)</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Sittar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ameer</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>: Multi-lingual Author Pro ling Using Stylistic Features</article-title>
          . In: Working Notes for MAPonSMS at FIRE'
          <fpage>18</fpage>
          - Workshop Proceedings of the 10th International
          <article-title>Forum for Information Retrieval Evaluation (FIRE'18)</article-title>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strassel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garland</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fore</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gainor</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabe</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Thomas,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Callahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Sawyer</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Collecting Natural SMS and Chat Conversations in Multiple Languages: The BOLT Phase 2 Corpus</article-title>
          . In
          <source>: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Reykjavik, Iceland (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Thenmozhi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalaivani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrabose</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>: Multi-lingual Author Pro ling on SMS Messages using Machine Learning Approach with Statistical Feature Selection</article-title>
          .
          <source>In: Working Notes for MAPonSMS at FIRE'18 - Workshop Proceedings of the 10th International Forum for Information Retrieval Evaluation (FIRE'18)</source>
          .
          <article-title>CEUR-WS.org</article-title>
          , CEUR,
          <string-name>
            <surname>DAIICT</surname>
          </string-name>
          , Gujarat, India (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Treurniet</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Clercq</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , Van Den Heuvel, H.,
          <string-name>
            <surname>Oostdijk</surname>
          </string-name>
          , N.:
          <article-title>Collecting a corpus of Dutch SMS</article-title>
          .
          <source>In: 8th International conference on Language Resources and Evaluation Conference (LREC</source>
          <year>2012</year>
          ). pp.
          <volume>2268</volume>
          {
          <fpage>2273</fpage>
          .
          <string-name>
            <surname>European Language Resources Association</surname>
          </string-name>
          (ELRA), Istanbul, Turkey (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Vicente</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batista</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carvalho</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>Improving Twitter Gender Classi cation using Multiple Classi ers</article-title>
          .
          <source>In: ESCIM 2016 : 8th European Symposium on Computational Intelligence and Mathematics</source>
          <year>2016</year>
          . pp.
          <volume>121</volume>
          {
          <fpage>127</fpage>
          .
          <string-name>
            <surname>So</surname>
            <given-names>a</given-names>
          </string-name>
          ,
          <source>Bulgaria</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Wanner</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Multiple Language Gender Identi cation for Blog Posts</article-title>
          .
          <source>In: Proceedings of the 37th Annual Meeting of the Cognitive Science Society</source>
          . pp.
          <volume>2248</volume>
          {
          <issue>2251</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>