<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ensemble Learning Approach f or Author Profiling</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Gilad Gressel</institution>
          ,
          <addr-line>Hrudya P, Surendran K, Thara S</addr-line>
          ,
          <institution>Aravind A, Prabaharan Poornachandran Amrita Center for Cyber Security, Amrita University</institution>
          ,
          <addr-line>Kollam</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <fpage>1148</fpage>
      <lpage>1156</lpage>
      <abstract>
        <p>With the evolution of internet, author profiling has become a topic of great interest in the field of forensics, security, marketing, plagiarism detection etc. However the task of identifying the characteristics of the author just based on a text document has its own limitations and challenges. This paper reports on the design, techniques and learning models we adopted for the PAN-2014 Author Profiling challenge. To identify the age and gender of an author from a document we employed ensemble learning approach by training a Random Forest classifier with the training data provided by PAN organizers for English language only. Our work indicate that readability metrics, function words and structural features play a vital role in identifying the age and gender of an author.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        With more than 2 billion users, internet has provided a solid platform for people to
share, communicate and express their ideas globally. Though online social media
have brought people together, they are vulnerable to crimes like identity thefts, false
information, identity masking etc. A lot of people fake their original identity either to
remain anonymous or to perform different cybercrimes. Zheng et al. in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has
showed that anonymity is a significant characteristic in online communities. The
process of identifying the traits of an author like age, gender, country, religion etc
from a document has become one of the hot topics for researches in the fields of
security, forensics, marketing, etc.
      </p>
      <p>In this paper, we present the working of our system which performs the
Author Profiling task by PAN-2014. This task aims at identifying the age and gender
of an author from four different datasets which are twitter tweets, blog data, social
media posts and hotel reviews. The training data for this analysis is also provided by
the PAN organizers. We employed different Natural Language Processing techniques
to extract features from a text document and using the Random Forest classifier we
determine the age and gender of the author. This paper presents the working model of
our system along with our architecture diagram, the machine learning algorithms and
techniques we used to complete the task.</p>
      <p>The rest of the paper is organized in the following order. Section 2 discusses the
related work by other researches we went through. In Section 3 we discuss our system
details and in section 4 we explain in detail our methodology and implementation
details. Finally in section 5 we draw conclusions and present future plans for
extending and improvising our author profiling task.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Literature Survey</title>
      <p>
        A considerable amount of research has already been done to identify the
traits of an author from a document using different machine learning algorithms and
statistical models. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] Peng et al. mention that each author has his/her own unique
stylometry of writing and refer to this feature as author profile. Pennebraker and
Stone in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] used the LIWC dataset and showed a relationship between language used
and the age of the author. The study performed by Vimala Balakrishnan and Paul H.P
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] prove that age and gender of a mobile phone users do influence their texting style.
John D. Burger, John Henderson and co-writers presented a language independent
classifier for gender prediction from the twitter micro-blogging site [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] Dong
Nyugen et al. presented a linear regression model to predict the age of an author from
a given text document. Claudia Peersman performed an exploratory study for
predicting age and gender form chat texts using Netlog corpus data [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. R
Chandramouli in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] compared the working of SVM, AdaBoost and Logistic
regression for gender identification. Shlomo Argamon shows the differences in the
male and female writing using the British National Corpus in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] the authors propose a tool TAT to profiling the authors of Arabic emails.
Calix et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] used 55 different features to analyze the stylometry of authors for
email author identification and authentication. All previous years PAN Author
Profiling research papers could be found from [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 System Architecture</title>
      <p>In this notebook we propose our solution for author profiling from a given
set of text documents. We have used a combination of Semantic, Syntactic and
Natural Language Processing (NLP) analysis for finding the same. The output of each
analysis is fed into a trained ensemble classifier which determines the age and gender
of the author. Figure 1 shows the detailed architectural diagram of our work.</p>
      <p>For performing the Author Profiling task we used the data corpus provided by
PAN-2014. The corpus consisted of various xml documents which had to be handled
in an offline (Hotel reviews and Social media) and online (Twitter and Blog) mode.
The Twitter training data had to be downloaded due to Twitter terms of service. The
cleaned blog corpus from RSS feeds was provided by PAN organizers, but it only
contained partial data. To obtain full data we decided to crawl it online and then use it
for analysis. The data crawled from both these modes is then cleaned for the removal
of xml contents, urls, user mentions (twitter), etc. The cleaned data is then pushed into
a database.</p>
      <p>After crawling and data cleaning we had the following data entries in the database
for training. Table 1 and 2 shows the number of posts we extracted overall.
The data set is then retrieved from the database and sent for text analysis which
involves Natural Language Processing, Readability Analysis, Syntactic analysis and
Structural Analysis. We extract a total of 22 features from each data set and send it to
Random Forest Classifier for training purposes. This trained ensemble classifier is
then used later for training purposes. In the next module we will explain our
implementation process in detail.</p>
    </sec>
    <sec id="sec-4">
      <title>4 Implementation Details</title>
      <p>
        As shown in Figure 1 our model consists of four main modules:
a. Crawlers: The training corpus provided by PAN-2014 contains documents in the
XML format. However for some data sets like Twitter and Blogs, data had to be taken
from HTML links in the XML file. Hence we had two modes of crawlers; one for
offline data sets (Social Media and Hotel Review) and another for online data sets
(Twitter and Blogs).
b. Data Cleaning: The raw text obtained from the crawlers has to be cleaned to
remove noisy data like „\ufffd‟, XML tags, urls, twitter user mentions, hashtags etc.
The presence of this noisy data could affect and reduce the accuracy of the entire
analysis. The cleaned data is then pushed into a database.
c. Text Analysis: From the database the cleaned data is retrieved and we employ
Natural Language Processing (NLP) techniques on the text data for its analysis. We
used the NLTK [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] platform for performing NLP techniques like Stemming,
Tokenizing, Parts of Speech (POS) tagging etc. Using these techniques we extract
features which is further divided into 3 subsets listed below:
i. Readability Metrics: Though readability metrics were created to find the
factors that help in making the text easy to read, it also plays a vital role in
identifying the characteristics of the author [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The readability score is a
statistical technique that computes readability based on the structure and
semantics of the sentence [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We use the following readability metrics for
our analysis:
a. ARI:
b. Flesch Reading Ease:
c. Flesch Kincaid Grade Level:
d. SMOG Index:
e. Gunning Fog Index:
f. Coleman Liau Index:
g. LIX:
h. RIX:
ii. Function Words: For extracting the function words features we used the
POS tagger of NLTK toolkit. We extract a total of 7 features from the text
which include the number of nouns, adjectives, verbs, pronouns, determiners,
adverbs and foreign words. Foreign words are those words which are mostly
slangs used in internet like “Helloooo”, “Whaaaat”, “yipeee”, “ROFL” etc.
iii. Syntactic Features: We extract the syntactic features from the text using the
NLTK Tokenizer. We extract a total of 7 features which include number of
single quotes, commas, periods, colons, semi-colons, question marks,
exclamation marks etc.
d. Classifier: All of the 22 features collected are then fed into an ensemble classifier.
For our analysis we used Random Forest classifier due to speed and accuracy. The
classifier is trained with the whole data corpus and used later for testing purposes. The
working of a cleaned test corpus is shown in Figure 2 given below.
In this section we present the results we obtained from our PAN Author Profiling
Task. We evaluated our analysis on the test corpuses provided by PAN organizers
using TIRA software. Table 3 and 4 show the age and gender predictions for English
corpus 1 and 2 respectively.
      </p>
    </sec>
    <sec id="sec-5">
      <title>6 Conclusion</title>
      <p>AGE
In this work, we present our system which identifies the age and gender of an author
from a given document. We employed supervised Random Forest ensemble classifier
for the Author Profiling task. We have performed our analysis on the 343013 training
data for the English language corpus provided by the PAN-2014 organizers.</p>
      <p>In our future work, we would like to perform a deeper analysis on the different
features and traits and techniques that would help to improve the efficiency of our
current system.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <surname>Rong</surname>
          </string-name>
          , et al.
          <article-title>"A framework for authorship identification of online messages: Writing‐style features and classification techniques</article-title>
          .
          <source>" Journal of the American Society for Information Science and Technology 57.3</source>
          (
          <year>2006</year>
          ):
          <fpage>378</fpage>
          -
          <lpage>393</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <surname>Fuchun</surname>
          </string-name>
          , et al.
          <article-title>"Language independent authorship attribution using character level language models</article-title>
          .
          <source>" Proceedings of the tenth conference on European chapter of the Association for Computational Linguistics-Volume 1. Association for Computational Linguistics</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Pennebaker</surname>
          </string-name>
          , James W., and
          <string-name>
            <surname>Lori</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Stone</surname>
          </string-name>
          .
          <article-title>"Words of wisdom: language use over the life span</article-title>
          .
          <source>" Journal of personality and social psychology 85.2</source>
          (
          <year>2003</year>
          ):
          <fpage>291</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Balakrishnan</surname>
          </string-name>
          , Vimala, and Paul HP Yeow.
          <article-title>"Texting satisfaction: Does age and gender make a difference."</article-title>
          <source>International Journal of Computer Science and Security 1</source>
          .1 (
          <year>2007</year>
          ):
          <fpage>85</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>John D.</given-names>
          </string-name>
          , et al.
          <source>"Discriminating gender on Twitter." Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Nguyen</surname>
            , Dong,
            <given-names>Noah A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>and Carolyn P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosé</surname>
          </string-name>
          .
          <article-title>"Author age prediction from text using linear regression."</article-title>
          <source>Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities. Association for Computational Linguistics</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Peersman</surname>
            , Claudia,
            <given-names>Walter</given-names>
          </string-name>
          <string-name>
            <surname>Daelemans</surname>
          </string-name>
          , and Leona Van Vaerenbergh.
          <article-title>"Predicting age and gender in online social networks." Proceedings of the 3rd international workshop on Search and mining user-generated contents</article-title>
          .
          <source>ACM</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Cheng, Na,
          <string-name>
            <given-names>Rajarathnam</given-names>
            <surname>Chandramouli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. P.</given-names>
            <surname>Subbalakshmi</surname>
          </string-name>
          .
          <article-title>"Author gender identification from text</article-title>
          .
          <source>" Digital Investigation 8.1</source>
          (
          <year>2011</year>
          ):
          <fpage>78</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Argamon</surname>
          </string-name>
          ,
          <string-name>
            <surname>Shlomo</surname>
          </string-name>
          , et al.
          <article-title>"Gender, genre, and writing style in formal written texts."TEXT-THE HAGUE THEN</article-title>
          AMSTERDAM THEN BERLIN-
          <volume>23</volume>
          .3 (
          <year>2003</year>
          ):
          <fpage>321</fpage>
          -
          <lpage>346</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Estival</surname>
          </string-name>
          ,
          <string-name>
            <surname>Dominique</surname>
          </string-name>
          , et al.
          <article-title>"TAT: an author profiling tool with application to Arabic emails</article-title>
          .
          <source>" Proceedings of the Australasian Language Technology Workshop</source>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Calix</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , et al.
          <article-title>"Stylometry for e-mail author identification and authentication</article-title>
          .
          <source>"Proceedings of CSIS Research Day</source>
          , Pace University (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Bird</surname>
          </string-name>
          , Steven.
          <article-title>"NLTK: the natural language toolkit." Proceedings of the COLING/ACL on Interactive presentation sessions</article-title>
          .
          <source>Association for Computational Linguistics</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Pitler</surname>
            , Emily, and
            <given-names>Ani</given-names>
          </string-name>
          <string-name>
            <surname>Nenkova</surname>
          </string-name>
          .
          <article-title>"Revisiting readability: A unified framework for predicting text quality</article-title>
          .
          <source>" Proceedings of the Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Luyckx</surname>
            , Kim, and
            <given-names>Walter</given-names>
          </string-name>
          <string-name>
            <surname>Daelemans</surname>
          </string-name>
          .
          <article-title>"Shallow text analysis and machine learning for authorship attribution</article-title>
          .
          <source>"</source>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Tim</surname>
            <given-names>Gollub</given-names>
          </string-name>
          , Martin Potthast, Anna Beyer, Matthias Busse, Francisco, Rangel, Paolo Rosso, Efstathios Stamatatos, and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>Recent Trends in Digital Text Forensics and its Evaluation</article-title>
          . In Pamela Forner, Henning Müller, Roberto Paredes, Paolo Rosso, and Benno Stein, editors,
          <source>Information Access Evaluation meets Multilinguality, Multimodality, and Visualization. 4th International Conference of the CLEF Initiative (CLEF 13)</source>
          ,
          <year>September 2013</year>
          . Springer.
          <source>ISBN 978-3-642-40801-4.</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Francisco Rangel, Paolo Rosso, Moshe Koppel, Efstathios Stamatatos and
          <string-name>
            <given-names>Giacomo</given-names>
            <surname>Inches</surname>
          </string-name>
          .
          <article-title>Overview of the Author Profiling Task at PAN 2013</article-title>
          . In Pamela Forner, Roberto Navigli, and Dan Tufis, editors,
          <source>Working Notes Papers of the CLEF 2013 Evaluation Labs</source>
          ,
          <year>September 2013</year>
          .
          <source>ISBN 978-88- 904810-3-1.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rangel</surname>
          </string-name>
          , Francisco, and Paolo Rosso.
          <article-title>"Use of Language and Author Profiling: Identification of Gender and Age."</article-title>
          <source>Natural Language Processing and Cognitive Science</source>
          (
          <year>2013</year>
          ):
          <fpage>177</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>