<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Statistical Learning Methods for Profiling Analysis</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Deparment - University of Neuchâtel</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Author profiling is the task to infer some information about an author by analyzing her/his writing style. It's application in forensics, business intelligence and psychology makes this topic interesting for researching. In this notebook, we present our baseline approach using SVM and Linear Discriminant Analysis (LDA) classifiers. We analyze features obtained from LIWC dictionaries, these are frequencies of use words by categories, which gives a general view about how the author writes and what he/she is talking about. According the experimental results, those are significant features to differentiate gender, age-group and personality. Although they are relatively few (not more than 100), they allow to discriminate with an acceptable accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Studies have demonstrated evidence of differences in the writing style according the
gender and age of the authors. These differences are detected with the use of
functionwords and content-words. The function-words define how the person use the grammar
and build sentences. On the other hand, content-words indicate what the person is
talking about. For example, Pennebaker [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] found that women tend to use more personal
pronouns and words referring to emotions. By the contrary, men tend to use more nouns,
prepositions and big words (defined as words with more than 6 characters).
      </p>
      <p>
        In the case of age, Pennebaker [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] found that younger writers use more personal
pronouns in first person and past tense verbs; while older writers tend to use more
articles, nouns prepositions and future tense verbs. Another example of this differences
is in Schler et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], where these authors found that men’s writing is more related to
money, job and TV, while women’s writing is more related to family, sex and eating. In
the case of age, younger people use to write more about sports, friends and emotions;
while older people write more about money, job, and family.
      </p>
      <p>
        More recently studies expanded this analysis to determinate weather the personality
influence the writing style too. For example Yarkoni [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] presented a detailed work were
he found that extroverted people are more likely to speak about leisure activities, family
and other persons than non extroverted. People open to new experiences talk more about
friends, time and positive emotions than other people; and some other similar relations
for all different personalities.
      </p>
      <p>Based on this evidence, we believe that the word categories are important features
to determine the profile of an author of a text, and we will try to measure how much they
can tell us about the authors of tweets. Therefore, in this notebook, we present a
baseline author profile classifier based on statistical learning over word category features.
First, we present the feature pre-processing, extraction and selection. Then, the
classification models to identify gender, age group, and personality. And finally, we present
the description of the experiments, the results and the conclusions.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Features</title>
      <sec id="sec-2-1">
        <title>LIWC</title>
        <p>
          The studies mentioned previously used the Linguistic Inquired and Word Count (LIWC)
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. This tool propose a list of word categories, each category is formed by agreement
between at least three judges. Then, given a text, it counts the number of words that the
text have per each category. The idea is to know how frequent a person use each word
category and with that data estimate some information about him or her.
        </p>
        <p>The LIWC categories are grouped in linguistic dimensions (e.g., part-of-speech);
content dimensions(e.g., emotions and activities); or spoken dimensions (e.g. fillers and
no fluencies markers). LIWC has dictionaries in a variety of languages, for this
experiment English, Italian, Spanish and Dutch dictionaries are used. Each dictionary
contains the list of predefined categories and the words associated with them, for example
for the category positive emotions some associated words are “fun", “nice", “succes",
etc. They were created over reiterative process of human judgment and were tested in
several times in different studies to assure their validity.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Additional Categories</title>
        <p>Additionally, we include other two groups of categories: punctuation marks and tweet.
We include seven categories in punctuation marks: question mark, exclamation, period,
comma, colon, semi-colon and all punctuation. The last one groups any punctuation
mark including the mentioned before. The tweet categories are added for the nature of
the corpus data and because they are frequently employed by tweet users: emoticons,
hyper-link, hashtag and references to other users. In Table 1 we can see a summary of
the total categories analyzed in this experiment.
The tweets were given in XML files. Each XML file was processed to extract only the
text information given as scale. First, the text was tokenized using white space
(including tab, change of line among others) and punctuation marks as separation characters.
We consider the apostrophe as part of the word, for example “she’s" is consider a single
word. This choose is related to the LIWC dictionaries that consider them in this way.
Once the tokens are obtained, they were counted in the corresponding categories. As
mentioned before, in the dictionary each category has a set of words, so, if the token is
part of the set of word of a category then it sums one to that category. One token can
appear in many categories (e.g. “I" as pronoun and function word).</p>
        <p>The granularity of the models is set by user and not by tweet, so, there is a vector
of categories per each user. Once we complete the counting per user, we divide each
count by the total number of words by user in other to obtain the relative frequencies.
Finally, we keep the frequencies in a matrix format, where the columns are the word
categories and the rows are the users. Thus, each row represent the distribution over
LIWC categories for the given user.</p>
        <p>One additional step is performed before the feature selection, the relative word
frequencies x are scaled by calculating the z score respect to each category according
to the following formula:
z =
x
:
(1)
where, and are the mean and standard deviation of the frequencies in each category.</p>
        <p>The frequency of use of words in not uniform in a language. Some of them are
highly used (e.g., function words) and some others have low frequency of use (e.g.,
topic related words), the relative frequencies are scaled because we need to compare
them obtaining their use related to each particular category and not to the general use
of language.
2.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Feature Selection</title>
        <p>
          Even we have a reduce set of features, we need to ignore noisy and irrelevant features
before applying a classification scheme. Additionally, it derives in an easier linguistic
explanation about the key features to discriminate among the different classes.
Gender and Age Group Fourth feature selection methods were evaluated to
determine the more suitable for the data: Manual selection, Information Gain, Odd ratio
and Support Vector Machine Recursive Feature Elimination (SVM RFE). The manual
selection was based on [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], where it is explained which are the more general
categories to differentiate an author according she/he’s age and gender. The
Information Gain and Odd Ratio were based on the study of Sebastiani [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] where he compares
different methods for feature reduction in text categorization. The SVM RFE proposed
by Guyon [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] is a backward feature elimination using SVM, it eliminates one feature
at time given a ranking criteria. The three last methods were implemented with Weka
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. For their evaluation, three different classifiers were tested and the results were
compared according the accuracy of classification. The best subset was obtained with SVM
RFE.
        </p>
        <p>Personality In the case of personality, the number of classes to discriminate is larger
than the previous models and it is more difficult to associate specific categories to
each class. Consequently, the methods mentioned before did not have significant
improvement in comparison of using the full set of features. So for this case, we applied
Forward-Backward Feature Selection, trying to improve the Root Mean Squared Error
(RMSE).
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Classification</title>
      <sec id="sec-3-1">
        <title>Gender and Age Group</title>
        <p>
          These classes were defined as categorical. We have two classes for gender: “Male"
and “Female", and fourth classes for age: “18-24", “25-34", “35-49", and “50-xx". The
classification was made with -SVM [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], which is a variant of the original SVM but
with an easier interpretation for the cost parameter called . In the experiments, was
set to 0.01 and we used radial kernel. The implementation was made in R with the
library “e1071".
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Personality</title>
        <p>In the case of personality, we define one model per each personality. Our first approach
was to define the classes as categorical without taking into account the order of the
score of the personality, so we have 11 classes from “-0.5" to “0.5" with one decimal of
difference. We choose two classifiers: -SVM with set to 0.01 and with radial kernel,
and Linear Discriminant Analysis (LDA). The implementation was done in R with the
libraries “e1071" and “MASS".
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <sec id="sec-4-1">
        <title>Training</title>
        <p>
          The corpus to develop the models is a training set of tweets in English, Spanish, Italian
and Dutch given by PAN 2015 Author Profiling task [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The validation was made
measuring the accuracy for gender and age-group, and RMSE for personality (according
the specification of the task). We used the full training set with leave-one-out validation.
The results are showed in Tables 3 and 4.
        </p>
        <p>In the feature selection, the experiment shows that there are some categories to
discriminate between gender which are independent of language while other are different
for each language, and the same patter for the other models. Tables 5, 6, and 7 contain
the common categories that were found in one or more languages.</p>
        <p>
          The training and testing are implemented separately. The outputs are the models
and the vectors of means and standard deviation calculated with the training data. These
vectors are used to calculate the z score of the testing data. This step corresponds to
Software 1 of TIRA [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
Linguistic Prepositions, word count, you, pronouns
Content Family, affect, space, swear, feel, emotions, body, home, work, TV, money,
future, motion, school, inclusion (and, we, both), exclusion (or, either, but)
Spoken None
Punctuations Question mark, exclamation mark, colon
Tweet Emoticon, reference to other users, hyper-links
Extroverted Word count, big words, pronouns, I, we, us, others, article, social, family,
emoticons, reference to other users
Stable Pronouns, oneself, we, us, others, article, affection, positive emotions, optimist,
anxiety, sadness, anger, emoticons, reference to other users, hyper-links
Agreeable Pronouns, I , others, prepositions, inhibition, sadness, certain, see, listen,
discrepancy, causation, cognitive process, emoticons, reference to other users,
hyper-links
Conscientious Pronouns, I , us, others, time, present, past, work, motion, home, optimist,
positive emotions, number, reference to other users, hyper-links
Open Pronouns, I , us, others, negation, preposition, number, affection, optimist,
certain, discrepancy, cause, tentative, see, insight, emoticons, reference to other
users, hyper-links
are the input files and the models. This step corresponds to Software 2 on TIRA [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
better than the average with less runtime than the majority. The best global results were
in Dutch and English, and the worse with Italian. In the case of gender and age-group,
the results were good comparing with the state of the art using similar features,
Argamon et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] reported 72% accuracy to distinguish gender and 67% for age-group
(having 3 groups). Specially in the case of Spanish, where we obtained 92% of
accuracy in gender. Nevertheless, the accuracy of classification of age-group was close to
average. In the case of personality, the results are also good taking into account the
difficulty of the data: bigger number of classes to discriminate many of them with very
few or none samples to train, and the few quantity of features used (less than 100). The
selected runs for testing personality where using SVM classifier.
        </p>
        <p>The global ranking of our solution for English was 7th over 22, Dutch 5th over 20,
Italian 9th over 19, and Spanish 8th over 21.
are measured by accuracy in %,“BOTH" is the accuracy when gender and age were both well
classified. The personality traits were measure with RMSE rounded to two decimal points, “RMSE"
is the average of all personalities traits.</p>
        <p>English
Performance
Best
Our solution
Mean
Worse
Dutch
Performance
Best
Our solution
Mean
Worse
Italian
Performance
Best
Our solution
Mean
Worse</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The present approach using LIWC Categories has demonstrated being a good solution
regarding the limitation of having a few quantity of features compared with other
solutions. According to the testing results, it had better performance than the average state
of the art. Moreover, it is simple and efficient. But the most important point is that we
can linguistically justify the classification decision because we can know which are the
key features for the decision process. Deeper analysis is needed to extract the
explanation of correct and incorrect assignment of classes; and to compare the differences
in the results using SVM or LDA classifier. We think that it can be improved in future
with a finer analysis of features and selection methods, and with a more appropriate
definition of the classes and modeling for age-group and personality traits.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schler</surname>
          </string-name>
          , J.:
          <article-title>Automatically profiling the author of an anonymous text</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>52</volume>
          (
          <issue>2</issue>
          ),
          <fpage>119</fpage>
          -
          <lpage>123</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Gollub</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burrows</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Ousting ivory tower research: towards a web framework for providing experiments as a service</article-title>
          .
          <source>In: Proceedings of the 35th international ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <fpage>1125</fpage>
          -
          <lpage>1126</lpage>
          . ACM (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Guyon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barnhill</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vapnik</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Gene selection for cancer classification using support vector machines</article-title>
          .
          <source>Machine learning 46(1-3)</source>
          ,
          <fpage>389</fpage>
          -
          <lpage>422</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The weka data mining software: an update</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Groom</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handelman</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          :
          <article-title>Gender differences in language use: An analysis of 14,000 text samples</article-title>
          .
          <source>Discourse Processes</source>
          <volume>45</volume>
          (
          <issue>3</issue>
          ),
          <fpage>211</fpage>
          -
          <lpage>236</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Booth</surname>
          </string-name>
          , R.J.:
          <article-title>Linguistic inquiry and word count: LIWC 2001</article-title>
          . Mahway: Lawrence Erlbaum Associates
          <volume>71</volume>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The Secret Life of Pronouns: What Our Words Say About Us</article-title>
          . Bloomsbury
          <string-name>
            <surname>Publishing</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Overview of the 3rd author profiling task at pan 2015</article-title>
          . In: Cappellato L.,
          <string-name>
            <surname>Ferro</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gareth</surname>
            <given-names>J.</given-names>
          </string-name>
          and San Juan E. (Eds). (Eds.)
          <article-title>CLEF 2015 Labs and Workshops, Notebook Papers</article-title>
          .
          <article-title>CEUR-WS.org (</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Schler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J.W.:</given-names>
          </string-name>
          <article-title>Effects of age and gender on blogging</article-title>
          . In: AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs. vol.
          <volume>6</volume>
          , pp.
          <fpage>199</fpage>
          -
          <lpage>205</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Schölkopf</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smola</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williamson</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bartlett</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          :
          <article-title>New support vector algorithms</article-title>
          .
          <source>Neural Computation</source>
          <volume>12</volume>
          (
          <issue>5</issue>
          ),
          <fpage>1207</fpage>
          -
          <lpage>1245</lpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Sebastiani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>: Machine learning in automated text categorization</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 34(1)</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Yarkoni</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Personality in 100,000 words: A large-scale analysis of personality and word use among bloggers</article-title>
          .
          <source>Journal of Research in Personality</source>
          <volume>44</volume>
          (
          <issue>3</issue>
          ),
          <fpage>363</fpage>
          -
          <lpage>373</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>