<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Author's Traits Prediction on Twitter Data using Content Based Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fahad Najib</string-name>
          <email>choudharyfahad@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Waqas Arshad Cheema</string-name>
          <email>waqascheema06@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rao Muhammad Adeel Nawab</string-name>
          <email>adeelnawab@ciitlahore.edu.pk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, COMSATS Institute of Information Technology</institution>
          ,
          <addr-line>Lahore</addr-line>
          ,
          <country country="PK">Pakistan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>This paper describes the methods we have employed to solve the author profiling task at PAN-2015. The proposed system is based on simple content based features to identify an author's age, gender and other personality traits. The problem of author profiling was treated as a supervised machine learning task. First content based features were extracted from the text and then different machine learning algorithms were applied to train the models. Results showed that content based features approach can be very useful in predicting the author's traits from his/her text.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Authorship attribution concerns with the classification of documents into the classes to
be predicted based on the writing style of their authors. In the case of author verification
and author identification tasks, the style of individual authors is examined. Whereas
author profiling mean to distinguish between classes of authors studying their sociolect
aspect, that is, how language is shared between people. This helps in predicting
profiling aspects such as age, gender or personality type. Author profiling is a problem
of increasing importance in several applications like forensics, security and marketing.
E.g., from a forensic linguistics prospect, the linguistic profile of the sender of a
harassing SMS message can be identified. Similarly, from a marketing perceptive, companies
would like to know the demographics of the people that like or dislike their products on
the basis of the text analysis of online product reviews and blogs.</p>
      <p>
        In recent years, automatic detection of an author’s profile from his/her text has
become an emerging and popular research area
        <xref ref-type="bibr" rid="ref13">(Rangel et al., 2013)</xref>
        . Automatically
predicting the identity of authors from their texts has a lot of future applications. for e.g.,
forensics analysis
        <xref ref-type="bibr" rid="ref1 ref5">(Corney et al., 2002; Abbasi and Chen, 2005)</xref>
        , marketing intelligence
        <xref ref-type="bibr" rid="ref8">(Glance et al., 2005)</xref>
        and classification and sentiment analysis
        <xref ref-type="bibr" rid="ref12 ref9">(Oberlander and Nowson,
2006)</xref>
        .
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        A significant amount of research in automatic classification of texts into the classes to be
predicted, has already been done by different researchers and linguists using several
different machine learning techniques
        <xref ref-type="bibr" rid="ref15">(Sebastiani, 2002)</xref>
        . Over the past few years, a large
variety of techniques have been devised for predicting the text based on its author’s
traits
        <xref ref-type="bibr" rid="ref1 ref11 ref12 ref14 ref7 ref9">(Abbasi and Chen, 2005; Houvardas and Stamatatos, 2006; Schler et al., 2006;
Argamon et al., 2009; Estival et al., 2008; Koppel et al., 2009)</xref>
        . Previously different
machine learning classifiers tried to include variety of techniques such as Lazy learners
(IBk)
        <xref ref-type="bibr" rid="ref6 ref7">(Estival et al., 2008, 2007)</xref>
        , Support Vector Machine (SVM)
        <xref ref-type="bibr" rid="ref11 ref6">(Koppel et al., 2009;
Estival et al., 2007)</xref>
        , LibSVM
        <xref ref-type="bibr" rid="ref7">(Estival et al., 2008)</xref>
        , RandForest
        <xref ref-type="bibr" rid="ref7">(Estival et al., 2008)</xref>
        ,
Information Gain
        <xref ref-type="bibr" rid="ref12 ref9">(Houvardas and Stamatatos, 2006)</xref>
        , Baysian Regression
        <xref ref-type="bibr" rid="ref11">(Koppel et al.,
2009)</xref>
        , Exponential Gradient
        <xref ref-type="bibr" rid="ref10">(Koppel et al., 2002)</xref>
        etc.
      </p>
      <p>
        Several approaches have been implemented and experiments conducted for
selecting the best possible features set for the most accurate classification. Houvardas and
Stamatatos (2006) showed the usefulness of n-gram, whereas Koppel et al. (2009) shown
the effect of gender and age in blogging sites by considering different word classes and
showing the relation of the word classes with the author’s age and gender. Koppel et al.
(2009); Estival et al. (2007) have identified that the Part-of-speech is also an
commendable linguistic feature and
        <xref ref-type="bibr" rid="ref4">(Calix et al., 2008)</xref>
        achieved the accuracy of 76.72% using
55 different features.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experimental setup</title>
      <p>The data used in our experiments is the training dataset of PAN-2015 1. The corpus
consists tweets on different topics, grouped by author and labeled with his/her language,
gender, age group and 5 personality traits ( extroverted (Ex), stable (St), agreeable (Ag),
conscientious (Co) and open (Op)). The documents are categorized as in languages
(English, Dutch, Italian and Spanish), two genders (male and female), and four groups
( 18-24, 25-34, 35-49 and 50-XX). With regard to personality traits, for each trait the
scores lies between -0.5 and 0.5. Documents in the corpus consist of a collection of
posts made by a single user.</p>
      <p>The corpus was balanced gender wise within each age group but imbalanced in
terms of age representation and the five personality traits scores distribution (-0.5 to
0.5). The proportion of languages, gender and each age group in the corpus within the
training dataset is presented in Table 1, whereas the personality traits distribution in
table 2.</p>
      <p>Prior to any model training or testing, we apply some pre processing steps to all
documents. We eliminated all the data contents that were not determined to be the text
written from the user like XML tags, as our primary source of features is the text written</p>
      <sec id="sec-3-1">
        <title>1 http://pan.webis.de/</title>
        <p>Class</p>
        <p>Agreeable
Conscientious
Extroverted</p>
        <p>Open
Stable
by an author. Now as all the user posts lie within the unparsed data tags of the source’s
xml file, we disregard any text not within these tags and HTML tags also.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Feature selection</title>
      <p>
        As male and females like to write about different topics, they use different words
accordingly. This leads to the fact that content based features can be an important tool
to distinguish between texts of males and females
        <xref ref-type="bibr" rid="ref14">(Schler et al., 2006)</xref>
        . For example, a
tweet concerned to sports will be more likely to be written by a male author rather by
a female. That tweet may contain words like goal, score, world cup etc. So the
occurrence of words like these will increase the chances of it being written by a male author.
Similarly occurrence of words or phrases like my husband, shopping, nailpolish etc will
increase the chances of it being written by a female author. In a similar fashion, people
in their teen age like to write more about their school life, and friends. Whereas people
in their 20’s like to write more about their college life and people of 30’s write more
about jobs, marriage and politics. So the content based features can be an important tool
to distinguish between texts written by people belonging to different profiles.
      </p>
      <p>We calculated the frequencies of different unigrams in the texts written by a
particular profile. Then, for every unigram, we calculated the ratio of its frequencies in
the tweets of different classes like male and female, different age groups and different
personality traits. Finally we selected the features on the basis of two combined factors.
First the unigrams with the highest frequencies in the corpus, and second the
difference of frequencies in different classes which are to be classify from one another. The
frequencies of some of the most frequently used unigrams for English language in the
corpus and their frequency comparison along gender and age groups is given in table 3.
Similarly this routine has carried out for all the four languages individually and four
sets of content based features selected for each language each. Then for each language,
two different sets of features has been used, one for gender and age group prediction
and one for the personality traits classification.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Models and Evaluation</title>
      <p>For each of the four languages, we trained different models. Then for each language,
we trained two models, one for age and gender and one for the personality traits. So
there were total training models build all on the content based features. We ran the
experiments on four machine learning classifiers: J48, Random Forest, Support Vector
Machines (SMO), and Naive Bayes. The evaluation measures used as instructed by PAN
2015 2, accuracy for age and gender and root mean squared error for the five personality
traits.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Results and Analysis</title>
      <p>Language Gender Age</p>
      <p>English 0.914 0.967
Italian 1.000 NA
Spanish 0.940 0.990
Dutch 1 NA</p>
      <p>Both Ex St Ag
0.894 0.076 0.093 0.088
NA 0.028 0.051 0.048
0.930 0.085 0.110 0.084
NA 0 0 0
Table 4. Results on training data.</p>
      <p>Based on the performance of the four classifiers (see section 5) on the training
data, we choose only one single classifier (SVM) for all the classes of age, gender and
personality in our final software that we have submitted in the competition. We are
repoting the results that we have achieved on the training data in table 4 and finally
on the testing data in table 5. Here the combined accuracy for age and gender (open),
the combined root mean squared error for all five personality traits (RMSE) have been
reported as well. Again here the age group results for Italian and Dutch are missing in
both training and testing data suggesting there is only 1 class.</p>
      <p>In comparison of performance in different languages, our system performed better
for English and Dutch as compared to Spanish and Italian. Traits wise the performance</p>
      <sec id="sec-6-1">
        <title>2 http://pan.webis.de/</title>
        <p>is reasonably good for Spanish gender, English age and Dutch personality. Overall, our
system couldn’t perform as well on the testing data as it performed on training data.
The possible reasons for this are the variation of language in training and testing data
and the nature of the content based approach.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>In this paper, we have presented a content based technique for the automatic
classification of the author’s gender, age and personality from their writing. This work has a
number of potential applications like marketing, forensics, and security. We have
performed our experiments on the training data provided by the PAN-2015 organizers. We
applied some simple content based techniques and the results we achieved are highly
motivational showing the usefulness of content based features in predicting the author’s
profile from text. In future work, the results can be further improved by incorporating
and finding more suitable features.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Ahmed</given-names>
            <surname>Abbasi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hsinchun</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Applying authorship analysis to arabic web content</article-title>
          .
          <source>In Intelligence and Security Informatics</source>
          , pages
          <fpage>183</fpage>
          -
          <lpage>197</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          , Moshe Koppel, James W Pennebaker, and Jonathan Schler.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>Automatically profiling the author of an anonymous text</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>52</volume>
          (
          <issue>2</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          3.
          <string-name>
            <given-names>K</given-names>
            <surname>Calix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Connors</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H</given-names>
            <surname>Manzar</surname>
          </string-name>
          ,
          <string-name>
            <surname>G MCabe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S</given-names>
            <surname>Westcott</surname>
          </string-name>
          .
          <article-title>Stylometry for e-mail author identification and a uthentication</article-title>
          .
          <source>Proceedings of CSIS Research Day</source>
          , Pace University,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Malcolm</given-names>
            <surname>Corney</surname>
          </string-name>
          , Olivier de Vel,
          <string-name>
            <surname>Alison Anderson</surname>
            , and
            <given-names>George</given-names>
          </string-name>
          <string-name>
            <surname>Mohay</surname>
          </string-name>
          .
          <article-title>Genderpreferential text mining of e-mail discourse</article-title>
          .
          <source>In Computer Security Applications Conference</source>
          ,
          <year>2002</year>
          .
          <source>Proceedings. 18th Annual</source>
          , pages
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          . IEEE,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Estival</surname>
          </string-name>
          , Tanja Gaustad, Son Bao Pham, Will Radford, and
          <string-name>
            <given-names>Ben</given-names>
            <surname>Hutchinson</surname>
          </string-name>
          .
          <article-title>Author profiling for english emails</article-title>
          .
          <source>In Proceedings of the 10th Conference of the Pacific Association for Computational Linguistics (PACLINGâA˘Z´07)</source>
          , pages
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          . PACLING,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Estival</surname>
          </string-name>
          , Tanja Gaustad, Ben Hutchinson, Son Bao Pham, and
          <string-name>
            <given-names>Will</given-names>
            <surname>Radford</surname>
          </string-name>
          .
          <article-title>Author profiling for english and arabic emails</article-title>
          .
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Natalie</given-names>
            <surname>Glance</surname>
          </string-name>
          , Matthew Hurst, Kamal Nigam, Matthew Siegler, Robert Stockton, and
          <string-name>
            <given-names>Takashi</given-names>
            <surname>Tomokiyo</surname>
          </string-name>
          .
          <article-title>Deriving marketing intelligence from online discussion</article-title>
          .
          <source>In Proceedings of the eleventh ACM SIGKDD international conference on Knowledge discovery in data mining</source>
          , pages
          <fpage>419</fpage>
          -
          <lpage>428</lpage>
          . Association for Computing Machinery,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          8.
          <string-name>
            <given-names>John</given-names>
            <surname>Houvardas</surname>
          </string-name>
          and
          <string-name>
            <given-names>Efstathios</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          .
          <article-title>N-gram feature selection for authorship identification</article-title>
          .
          <source>In Artificial Intelligence: Methodology, Systems, and Applications</source>
          , pages
          <fpage>77</fpage>
          -
          <lpage>86</lpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Moshe</given-names>
            <surname>Koppel</surname>
          </string-name>
          , Shlomo Argamon, and Anat Rachel Shimoni.
          <article-title>Automatically categorizing written texts by author gender</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <fpage>401</fpage>
          -
          <lpage>412</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          10.
          <string-name>
            <surname>Moshe</surname>
            <given-names>Koppel</given-names>
          </string-name>
          , Jonathan Schler, and
          <string-name>
            <given-names>Shlomo</given-names>
            <surname>Argamon</surname>
          </string-name>
          .
          <article-title>Computational methods in authorship attribution</article-title>
          .
          <source>Journal of the American Society for information Science and Technology</source>
          ,
          <volume>60</volume>
          (
          <issue>1</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>26</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Jon</given-names>
            <surname>Oberlander</surname>
          </string-name>
          and
          <string-name>
            <given-names>Scott</given-names>
            <surname>Nowson</surname>
          </string-name>
          .
          <article-title>Whose thumb is it anyway?: classifying author personality from weblog text</article-title>
          .
          <source>In Proceedings of the COLING/ACL on Main conference poster sessions</source>
          , pages
          <fpage>627</fpage>
          -
          <lpage>634</lpage>
          . Association for Computational Linguistics,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          12. Francisco Rangel, Paolo Rosso, Moshe Koppel, Efstathios Stamatatos, and
          <string-name>
            <given-names>Giacomo</given-names>
            <surname>Inches</surname>
          </string-name>
          .
          <article-title>Overview of the author profiling task at pan 2013</article-title>
          .
          <source>Notebook Papers of CLEF</source>
          , pages
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          13.
          <string-name>
            <surname>Jonathan</surname>
            <given-names>Schler</given-names>
          </string-name>
          , Moshe Koppel, Shlomo Argamon, and James W Pennebaker.
          <article-title>Effects of age and gender on blogging</article-title>
          .
          <source>In AAAI Spring Symposium: Computational Approaches</source>
          to Analyzing Weblogs, volume
          <volume>6</volume>
          , pages
          <fpage>199</fpage>
          -
          <lpage>205</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          14.
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          .
          <source>Machine learning in automated text categorization. ACM computing surveys (CSUR)</source>
          ,
          <volume>34</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>