<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identification of Author Personality Traits using Stylistic Features Notebook for PAN at CLEF 2015</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ifrah Pervaz</string-name>
          <email>1ifrahpervaz23@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iqra Ameer</string-name>
          <email>2iqraameer133@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abdul Sittar</string-name>
          <email>3abdulsittar72@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rao Muhammad Adeel Nawab</string-name>
          <email>4adeelnawab@ciitlahore.edu.pk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>COMSATS Instiute of Inforamtion Technology</institution>
          ,
          <addr-line>Lahore</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Author profiling is the task of determining the age, gender or type of the author's personality by studying their sociolect aspect, that is, how the language is shared by people. This paper presents the COMSATS Institute of Information Technology, Lahore entry for the PAN 2015 competition on Author Profiling task. Our proposed system is based on stylometry features. We implemented 29 different stylistic features, many of which are language independent. Since the training data was available in multiple languages, one of our main objectives was to explore which language independent features are most effective. The problem of author profiling was casted as a supervised document classification task. Results showed that features (Percentage of Question Sentences, Average Sentence Length, Percentage of Punctuations, Percentage of Comma and Percentage of Full stops) were most effective multilingual features.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Personality in Encyclopedia of Psychology is defined as “individual differences in
characteristic patterns of thinking, feeling and behaving.” [1] These differences can be
reflected through one’s speech, writing, images etc. Authorship analysis deals to find
out the methods to identify these differences and use them to know profile of author
such as gender, age, native language, education, profession or personality type. So
author profiling can simply be defined as: given the set of texts, you need to identify
age, gender, profession, education, native language and similar personality traits.
Authorship analysis has attracted much attention in recent years due to the rapid
increase of electronic text and the need for expert systems able to handle this
information. Like from a marketing viewpoint, companies may be interested in
knowing about the trends regarding their products, on the basis of the analysis of blogs
and online product reviews, what types of people like or dislike their products, what is
required by the customer and which customer category/ class they can attract more.
Similarly from a forensic viewpoint, determining the linguistic profile of a person i.e.
who wrote a "suspicious text"’ may provide valuable background information.
In this paper we explore how different stylistic features help in conveying about the
personality of writer. Moreover, we also tried to find out how the use of different
stylistic features affect multilingual results. We used the datasets provided by PAN
organizers, and applied different machine learning practices for predicting writer’s
traits.</p>
      <p>The rest of this paper is organized as follows: Section 2 describes related work. Section
3 presents the approach used to identify personality traits of author. Section 4 describes
the results obtained on training data. Section 5 presents the results obtained on test data
and section 6 concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>With the advancement in web technology, there is an abundant increase in use of
electronic media. Along with useful information there is also heap of anonymous data
available so the need to identify “who is who” has become important.
A lot of establishments have been done in this field so far, Pennebaker et al. [2] joined
dialect utilization with identity qualities, examining how the variety of semantic
attributes in a content can give data in regards to the gender and age of its author.
Argamon et al. [3] analyzed formal written texts extracted from the British National
Corpus [4]1 combining function words with part-of-speech features. Koppel studied the
problem of automatically determining an author’s gender by proposing combinations
of simple lexical and syntactic features.</p>
      <p>Holmes and Meyerhoff [7], Burger and Henderson [6] have also investigated obtaining
age and gender information from formal texts.</p>
      <p>Seifeddine and Maher [8], focus is on author’s discussions, they combine content based
approach and statistic approach. They show a hemi- strategy that merges the analysis
of information in writings with a machine learning technique. To get a higher
organization of these statistics, they depended on the utilization of the “Decision table
calculation".</p>
    </sec>
    <sec id="sec-3">
      <title>3 Our Approach</title>
    </sec>
    <sec id="sec-4">
      <title>3.1 Stylistic Features</title>
      <p>
        The approaches commonly used for the automatic identification of an author’s
personality traits from text can be categorized into three broad categories:
1 www.natcorp.ox.ac.uk
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Stylometry based approaches (which aim to identify an author’s traits from his
writing style), (2) Content based approaches (which identify author traits using features
extracted from the content of the document) and (
        <xref ref-type="bibr" rid="ref2">3</xref>
        ) Topic based approaches (which try
to predict an author’s profile based on the topics used in the document).
12. PCeorncjeunntcatgieons
13. Percentage of Comma
14. Percentage of Articles
15. Percentage of Words with One Syllable
16. SPyerllcaebnlteasge of Words with Three Plus
17. Average Syllables per Word
18. Percentage of Adjectives
19. Percentage of Adverbs
20. Percentage of Capitals
21. Percentage of Colons
22. Percentage of Determiners
23. Percentage of Digits.
24. Percentage of Full stop
25. Percentage of Interjections
26. Percentage of Modals
27. Percentage of Nouns
28. Percentage of Personal Pronouns
29. Percentage of Verbs
      </p>
      <p>Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes</p>
      <sec id="sec-4-1">
        <title>English</title>
        <p>Yes
Yes
Yes
Yes
Yes
Yes</p>
      </sec>
      <sec id="sec-4-2">
        <title>Languages</title>
      </sec>
      <sec id="sec-4-3">
        <title>Dutch Spanish</title>
        <p>Yes Yes
Yes Yes
Yes Yes
Yes Yes
Yes Yes
Yes Yes
Yes
Yes
Yes
---------Yes
------------------Yes
Yes
---Yes
Yes
---------------Yes
Yes
Yes
---------Yes
------------------Yes
Yes
---Yes
Yes
---------------</p>
      </sec>
      <sec id="sec-4-4">
        <title>Italian</title>
        <p>
          Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
Yes
---------Yes
------------------Yes
Yes
---Yes
Yes
---------------The system presented in this PAN Author Profiling Competition is based on stylometry.
Table 3.1 shows the list of 29 stylistic features used for the development of our proposed
author profiling detection system. As can be noted that these features aim to extract
different stylistic information from a single document, which can be helpful in
identifying the age, gender, personality type i.e. stable, open, extroverted, agreeable and
conscientious .In this year’s training data, the tweets are available for four different
languages: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) English, (2) Dutch, (
          <xref ref-type="bibr" rid="ref2">3</xref>
          ) Spanish and (4) Italian. This study aims to
identify some language independent stylistic features which are probable to perform on
multiple languages as well. As can be noted from Table 3.1 that features numbered 1,
2, 3, 4, 5, 6, 7, 8, 9, 13, 20, 21, 23 and 24 are language independent and can be used for
extracting stylistic information from document in any of the four languages i.e. English,
Dutch, Spanish and Italian, while the remaining ones can only be used for English
language.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3.2 PAN 2015 Author Profiling Training Datasets</title>
      <p>
        Our proposed system was trained using the PAN 2015 Author Profiling training data,
which consists of Twitter tweets in four different languages: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) English, (2) Dutch, (
        <xref ref-type="bibr" rid="ref2">3</xref>
        )
Spanish and (4) Italian. In the training dataset there are 152, 34, 100 and 38 author
profiles for English, Dutch, Spanish and Italian languages respectively.
In the PAN 2015 Author Profiling training dataset, there are two classes for gender
(male and female), four classes for age (18-24, 25-34, 35-49 and 50-xx) for the
remaining personality traits open, stable, agreeable, extroverted and conscientious there
are two classes: (yes or no).
      </p>
    </sec>
    <sec id="sec-6">
      <title>3.3 Evaluation Methodology</title>
      <p>The task of identifying an author’s profile from text was treated as a supervised machine
learning problem. N-fold-cross-validation was used. Due to difference in the sizes of
training data for different languages, we used 5-fold cross validation for English
language corpus, 4-fold cross validation for Dutch and 3-fold cross validation for the
Italian and Spanish corpora respectively.</p>
      <p>We applied a range of machine learning algorithms on the training data including
Navies Bayes, Support Vector Machine, Random Forest, J48 and Logistics. The
WEKA’s2 [9] implementation of these algorithms was used. The scores generated using
the stylistic features (see Table 3.1) were used as “input features” to the machine
learning algorithms.</p>
      <p>We used WEKA’s “attribute selection” approach for selecting “best features” from the
complete feature set. For that purpose we explored CFsSubSetEval, Filtered Attribute,
2 mailto:https://weka.wikispaces.com/
SVM Attribute and Classifier Subset as evaluators, and Best first and ranker as search
methods.</p>
    </sec>
    <sec id="sec-7">
      <title>3.4 Evaluation Measure</title>
      <p>As recommended by PAN 2015 organizers, the performance of the proposed system for
age and gender personality traits was measured using accuracy, whereas for other
personality traits (stable, open, extroverted, agreeable and conscientious) average Root
Mean Squared Error was used.
4</p>
    </sec>
    <sec id="sec-8">
      <title>Results on Training Data</title>
      <p>
        We carried out three sets of experiments: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) performance on individual features, (2)
performance on combined features (mean all 29 features are used as input) and (
        <xref ref-type="bibr" rid="ref2">3</xref>
        )
performance on “selected features”. The best performance was obtained using “selected
features”, therefore, we are only reporting the results for them.
In this paper we have discussed our participation in the Author Profiling task. We have
covered the role of stylistic features in identification of author personality traits. For
that purpose we figured out 29 features and perform different experiments on these, like
comparing accuracy by using all features, then checking accuracy for single feature and
finally using subsets of them and come to conclusion that best results are achieved by
feature selection techniques.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alan</surname>
            <given-names>E</given-names>
          </string-name>
          . Kazdin,
          <source>PhD, Editor-in-Chief Encyclopedia of Psychology: 8 Volume Set 2</source>
          .
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehl</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Niederhoffer</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          ,
          <article-title>Psychological aspects of natural language use: Our words, our selves</article-title>
          .
          <source>Annual review of psychology</source>
          ,
          <volume>54</volume>
          (
          <issue>1</issue>
          ):
          <fpage>547</fpage>
          -
          <lpage>577</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          3.
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fine</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shimoni</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          , Gender, genre, and
          <article-title>writing style in formal written texts</article-title>
          . TEXT,
          <volume>23</volume>
          :
          <fpage>321</fpage>
          -
          <lpage>346</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          5.
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <article-title>Automatically profiling the author of an anonymous text</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>52</volume>
          (
          <issue>2</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          ,
          <year>February 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          6.
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henderson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zarrella</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>Discriminating gender on twitter</article-title>
          .
          <source>In Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP '11</source>
          , pages
          <fpage>1301</fpage>
          -
          <lpage>1309</lpage>
          , Stroudsburg, PA, USA,
          <year>2011</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          7.
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Meyerhoff</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <article-title>The Handbook of Language and Gender</article-title>
          . Blackwell Handbooks in Linguistics. Wiley,
          <year>2003</year>
          . ISBN 9780631225027.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mechti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaoua</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belguith</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>Faiz1</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ,
          <article-title>Machine learning for classifying authors of anonymous tweets, blogs, reviews and social media Notebook for PAN at CLEF</article-title>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jason</surname>
          </string-name>
          <article-title>Brownlee Feature Selection to Improve Accuracy and Decrease Training Time</article-title>
          ,
          <source>March</source>
          <volume>12</volume>
          , 2014 in Uncategorized
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ian H. Witten</surname>
            , Eibe Frank and
            <given-names>Mark A.</given-names>
          </string-name>
          <string-name>
            <surname>Hall</surname>
          </string-name>
          ,
          <source>Data Mining Practical Machine Learning Tools and Techniques</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Inche</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tufis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , (ed),
          <source>Overview of the Author Profiling Task at PAN 2013 Working Notes Papers of the CLEF's. 2013 Evaluation Labs</source>
          ,
          <year>September 2013</year>
          .
          <source>ISBN 978-88-904810- 3-1.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>