<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Analysis of Big Five Personality Traits by Processing of Social Media Users Activity Features</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Proceedings of the XX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2018)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>RUDN University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Maxim Stankevich © Ivan Smirnov Institute for Systems Analysis, Federal Research Center “Computer Science and Control” of the Russian Academy of Sciences</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Nikolay Ignatiev RUDN University, Moscow, Russia © Oleg Grigoriev Federal Research Center “Computer Science and Control” of the Russian Academy of Sciences, Moscow, Russia © Natalia Kiselnikova Psychological Institute of Russian Academy of Education</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>162</fpage>
      <lpage>166</lpage>
      <abstract>
        <p>The study focused on the analysis of relation between Big Five personality traits of a user and his activity in popular Russian social media Vkontakte. In order to receive Big Five personality trait scores, we asked Vkontakte users to complete a psychological survey and then analyzed data from their personal public social media pages. The purpose of the study was to investigate the relation between social media activity features and users' level of neuroticism, conscientiousness, extraversion, openness to experience and agreeableness. To perform the task, we used machine learning classification algorithms.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The Big Five personality traits model is a popular
psychological tool, which is commonly used for
describing the human personality through the following
measurements: neuroticism, conscientiousness,
extraversion, openness to experience and agreeableness
[1]. Personality traits scores are usually calculated with
the help of questionnaires. Widespread use of social
media makes it possible to receive information about
social media users by analyzing data retrieved from their
public pages. However, there are only a few studies
related to the analysis of users’ Big Five personality traits
by using social media activity information from
Russianspeaking social networks. A number of researchers are
involved in Big Five personality traits prediction and
analysis for English-speaking social networks [2,3], but
there are no in-depth studies for Russian. The proposed
approach and the dataset thus collected are new for the
Russian social network analysis. The purpose of the
study was to investigate the relation between social
media activity features and user’s level of neuroticism,
conscientiousness, extraversion, openness to experience
and agreeableness by using machine learning algorithms.</p>
      <p>In order to form the dataset, we asked volunteers to
complete NEO-FFI questionnaire [4] and then to provide
access to their public pages information under privacy
constraints. Thus, we received data of 165 users from popular
Russian social network Vkontakte. We presented five
personality traits scores on the following scale: low level,
medium, and high. The idea is to present the problem as
multiclass classification. Classification features are based on
a 1-year period of user activity represented as posts on their
public pages and general information about users’ profiles
such as gender and a total number of friends and followers.
To evaluate methods, we ran two sets of experiments with
different classifiers: support vector machine and random
forest.</p>
      <p>The main issue that we faced with was a lack of
training examples. Though because of insufficient data
we couldn’t significantly improve classification
performance, we came to the conclusion that feature
format should be redesigned and text analysis-based
features should be added. We continue data collection
and look forward to improve our results in the nearest
future.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Related works</title>
      <p>There is a lot of studies that investigate social
mediabased data usage for classification in different
psychology related tasks.</p>
      <p>Besides Big Five personality analysis, detection of
depression, post-traumatic stress disorder and anxiety, is
also a very important problem. For example, CLPsych
2015 Shared Task organizers built the dataset consisting
of messages collections of depressed and non-depressed
users and asked contributors to share the performance of
their depression detection models [5]. This shared task as
well as the other similar studies, such as [6] and [7] used
textual data natural language processing methods to form
features for a predictive model. Authors of related works
[8] and [9] used a social media activity features to
improve classification performance. One should take
into account that depression detection task is
timedependent – it is necessary to consider time constraints
while dataset preparation, at the same time Big Five
personality traits are more consistent in time [10].</p>
      <p>One of the most significant studies, related to
social media language and Big Five personality
traits, presented in [3]. The authors performed
analysis of 700 million words, phrases, and topic
instances collected from Facebook messages of
75,000 volunteers, who took a standard personality
test. The work demonstrates some important
dependencies between language use and users’
personality attributes. For each personality trait,
they formed the list of related words which showed
valuable correlations with a neuroticism,
conscientiousness, extraversion, openness to
experience and agreeableness levels.</p>
      <p>The work presented in [11] describes the Big
Five personality traits prediction models for Twitter
users. This research dataset contains most recent
2000 tweets of 279 volunteers. To perform the task
authors decided to present personality traits scores
as values on a normalized 0-1 scale. The features
used were based on text analysis. The authors utilize
Linguistic Inquiry and Word count tool [12] to
produce statistics on 81 different features. The
MRC Psycholinguistic Database [13] was used to
retrieve features from users’ vocabulary. The
authors also performed social media activity
features. The results of correlation analysis revealed
that some of them had correlations with five-factor
personality model. The proposed models showed
about 15% of mean absolute error on a normalized
scale for each personality trait as a measurement of
model prediction accuracy.</p>
      <p>Another work related to the task of Big Five
traits prediction described in [14]. The data for the
research contains information about likes of 58,466
volunteers from the Facebook social media. The
authors used decomposed User-Likes matrix with
logistic and linear regression classificators to
predict users Big Five personality traits and other
6
mypersonality.org
personal attributes. The model achieved high
prediction performance for personal attributes such
as gender, age, and nationality (~80% Area Under
Curve) and about 35% of accuracy score for users’
Big Five personality traits.</p>
      <p>The research presented in [15] has a similar with
our work idea. Authors collected users’ data from
Vkontakte and performed correlation analysis using
social media activity indicators. The main interest
was on photos published on the users’ public pages.
According to the results, most significant
correlations were found between extraversion and
such activity indicators as a number of friends and
followers, total numbers of posts and some photo
information-based indicators. Neuroticism score
also showed valuable positive correlation with a
users’ total number of posts.</p>
      <p>We analyzed related works and came up to the
following conclusion. The background studies
propose valuable methodologies for Big Five
personality traits analysis and prediction which are
mainly related to language use of English-speaking
social media users. For Russian-speaking social
networks this problem is not well studied. For
example, in 2007 myPersonality project 6 started to
gather social media data and results of psychology
questionnaires from Facebook users. The huge
volume of this project was successfully used for
different academic studies. However, there are no
available and appropriate datasets based on
Russian-speaking social media. This is the main
reason why we had to form our original for the task
of Big Five personality traits analysis of Vkontakte
users.</p>
    </sec>
    <sec id="sec-3">
      <title>3 Dataset</title>
      <p>To build the dataset we asked volunteers from
Vkontakte to take part in a psychological survey and
complete NEO-FFI questionnaire. After this part,
we requested access to their public pages under
privacy constraints. Finally, for those who provided
their acceptance and completed questionnaire we
collected all available information from their public
profile pages. Overall, data from 165 profiles was
assembled. Personal information that can reveal the
identity of a persons was removed from the data.</p>
      <p>We divided collected data into two categories:
general information about users and information
about user messages posted during the time period
from January 2017. The first part contains such
features as - number of friends, number of
followers, gender, number of followed groups and
communities, etc. The second part contains the text
of the users’ messages, timestamps, and numbers of
likes, commentaries, and reposts (analog of a
retweet on Twitter).</p>
      <p>It is worth mentioning, that we continue to expand
our dataset with new examples. This study is based on
the current amount of available data, but we consider this
number only as an intermediate stage.</p>
    </sec>
    <sec id="sec-4">
      <title>4 Methods</title>
      <sec id="sec-4-1">
        <title>4.1 Big five personality traits</title>
        <p>Here we describe the methodology for the Big Five
personality traits prediction. We also describe the
personality scores representation and features that we
extract from available data.</p>
        <p>
          As a first step, we divide the initial NEO-FFI score
scale (
          <xref ref-type="bibr" rid="ref1 ref10 ref11 ref12 ref13 ref14 ref15 ref16 ref17 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-48</xref>
          ) of each personality trait as following: low
level (
          <xref ref-type="bibr" rid="ref1 ref10 ref11 ref12 ref13 ref14 ref15 ref16 ref17 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-20</xref>
          ), medium level (21-32) and high level (33-48)
[16]. As a result, one of these three classes were assigned
to each of user’s scores of neuroticisms,
conscientiousness, extraversion, openness to experience
and agreeableness. Thus, the initial task is transformed to
the task of multiclass classification. It should be noted
that such approach imposes some restrictions on
evaluation method. Figure 1 represents the class
distribution among users’ level of extraversion.
        </p>
        <p>Despite the fact that the medium level covers the
shortest score interval, Figure 1 illustrates that the
majority of users fall into this class. The same situation
is observed with other personality traits. The statistics for
each of Big Five personality trait presented in Table 1.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.1 Features</title>
        <p>The format of Vkontakte personal page provides a wide
range of user information. We used gender, number of
friends, number of followers, number of followed
groups, number of photo, and number of audio tracks to
form a users’ feature set. While filling out a Vkontakte
personal page, users can provide their opinion on
predefined question such as, how they relate to smoking
or what is the most important in people and life. Our data
include all the answers, but it is hard to present such
information as a feature. Since these questions are not
mandatory for Vkontakte users, we decided to assign</p>
        <sec id="sec-4-2-1">
          <title>Conscientiousness, %</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>Medium:</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Extraversion, %</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>Medium:</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Neuroticism, %</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>Openness to experience, %</title>
        </sec>
        <sec id="sec-4-2-7">
          <title>Agreeableness, % Low:</title>
        </sec>
        <sec id="sec-4-2-8">
          <title>Medium:</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>High: Low:</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>High: Low:</title>
        </sec>
        <sec id="sec-4-2-11">
          <title>High:</title>
          <p>Low:</p>
        </sec>
        <sec id="sec-4-2-12">
          <title>Medium:</title>
        </sec>
        <sec id="sec-4-2-13">
          <title>High:</title>
          <p>Low:</p>
        </sec>
        <sec id="sec-4-2-14">
          <title>Medium:</title>
          <p>High:
them binary values that represent if a user provided this
information or not. We assume that these answers can
indicate users’ general readiness to share their opinion
with other people and that might be valuable for future
analysis.</p>
          <p>As it was mentioned, we collected users’ messages
from their public pages. We used information about
likes, commentaries and reposts related to these
messages to calculate their averaged values on a single
post. The fact that for every user we collected messages
posted during an equal time period allows as to use total
number of assembled posts as a feature. The messages
timestamps were used to calculate the proportion of
users’ messages posted during night time (12 P.M – 6
A.M.).</p>
          <p>Figure 2 Number of words in users’ messages.</p>
          <p>However, Vkontakte profiles in personal pages
provide much less text data than Facebook and Tweeter.
The most popular format of Vkontakte users’ activity is
reposting. A large amount of communities provides
different kind of content and users usually only repost
this content on their personal pages without giving any
commentaries or opinions. Overall, we collected 13152
posts, but majority of them were empty reposts. Only
2637 of them contain texts written by users themselves.
The total number of used words for each user is presented
on Figure 2.</p>
          <p>As we can see on the Figure 2, current data contains
a very limited amount of information about Vkontakte
language. Considering this, we decided to perform
classification without language analysis. It is necessary
to collect much more data before applying text analysis
and compiling text-based features. In this work, we
perform classification task using mostly social media
activity features.</p>
          <p>Despite this fact that we ignored lexical features in
this research, we processed messages data to form
several additional features. For example, the average
number of sentences and words. We also computed the
proportion of uppercase words as well as the number of
ellipses in the users’ writings. We assume that described
features could reveal some specifics of people’s behavior
in social media.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Results of experiments</title>
      <p>The following chapter represents the results of our
experiments. To perform the evaluations, we used
scikitlearn implementation of random forest and multiclass
SVM algorithms [17]. The parameters for the
classification were set up by grid-search with 4-fold
cross-validation.</p>
      <p>We calculated the macro variation of recall,
precision, and f1-score to present classification
performance. To evaluate the accuracy of our models we
compiled 10 runs of 4-fold cross-validation on the data.
The results of our experiments presented as an averaged
value of these runs for each metric. The multiclass
classification results with a 4-fold cross-validation
presented in Table 2. The best values for each metric
highlighted in bold.</p>
      <p>The best performance was shown for the
agreeableness and neuroticism with a 49% and 53% of
f1-score respectively. The slightly worse results were
received for extraversion and openness to experience
with a 45% and 46% of f1-score. Random forest
classification algorithm was used to get these results. The
conscientiousness personal trait performance was the
lowest in our experiments with only 36% of f1-score
received by SVM. It is worth to note that in the most
cases SMV achieved more precision than RF, but recall
score was significantly less.</p>
      <p>In general, we can’t define considered performance
as good. However, limited information about language
use of Vkontakte users prevented the possibility to
compile lexical features and perform text analysis.
According to the results of studies based on
Englishspeaking social media, text features might serve as an
effective revealing tool for users Big Five personality
traits. Thus, in this study, we mostly tested social media
activity features, which we can describe as being useful
for the considered task.</p>
    </sec>
    <sec id="sec-6">
      <title>6 Conclusion</title>
      <p>In this work, we performed the prediction of Big Five
personality traits of social media users. We collected
results of NEO-FFI questionnaire taken by 165
volunteers and compiled dataset using social media
activity information from their personal pages. The
personality traits scores were represented as low,
medium, and high levels to transform the task into
multiclass classification.</p>
      <p>We can define two limitations that we faced during
our work. The first one consists of the fact that Vkontakte
users’ messages provide a very small amount of text data.
We observed that collected messages, for the most part,
are empty reposts, which don’t provide any text written
by users personally. This limitation imposes some
restriction on our current study. The features for the
classification were compiled by processing of social
media activity information without any lexical features.
We assume that such features can greatly improve
classification results. The second limitation is a simple
lack of examples in our current dataset.</p>
      <p>Considering this limitation, we can admit that our
most important task now is to add much more new
examples to the dataset. With a greater size of data, we
can utilize text analysis approaches and investigate the
relation between Big Five personality traits and
Russianspeaking social media language, which is currently an
unresearched field of study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Gosling</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rentfrow</surname>
            ,
            <given-names>P. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Swann</surname>
            Jr,
            <given-names>W. B.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>A very brief measure of the Big-Five personality domains</article-title>
          .
          <source>Journal of Research</source>
          in personality,
          <volume>37</volume>
          (
          <issue>6</issue>
          ),
          <fpage>504</fpage>
          -
          <lpage>528</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Ortigosa</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carro</surname>
            ,
            <given-names>R. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Quiroga</surname>
            ,
            <given-names>J. I.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Predicting user personality by mining social interactions in Facebook</article-title>
          .
          <source>Journal of computer and System Sciences</source>
          ,
          <volume>80</volume>
          (
          <issue>1</issue>
          ),
          <fpage>57</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>H. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichstaedt</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dziurzynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramones</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Ungar</surname>
            ,
            <given-names>L. H.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Personality, gender, and age in the language of social media: The open-vocabulary approach</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>8</volume>
          (
          <issue>9</issue>
          ),
          <year>e73791</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>P. T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          (
          <year>1989</year>
          ).
          <article-title>NEO five-factor inventory (NEO-FFI)</article-title>
          . Odessa, FL: Psychological Assessment Resources.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Coppersmith</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hollingshead</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp; Mitchell,
          <string-name>
            <surname>M.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>CLPsych 2015 shared task: Depression and PTSD on Twitter</article-title>
          .
          <source>In Proceedings of the 2nd Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality</source>
          (pp.
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Yazdavar</surname>
            ,
            <given-names>A. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Olimat</surname>
            ,
            <given-names>H. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ebrahimi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bajaj</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirunarayan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Sheth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2017</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Semi-Supervised Approach to Monitoring Clinical Depressive Symptoms in Social Media</article-title>
          .
          <source>In Proceedings of the 2017 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining</source>
          <year>2017</year>
          (pp.
          <fpage>1191</fpage>
          -
          <lpage>1198</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Jamil</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Monitoring Tweets for Depression to Detect At-risk Users (Doctoral dissertation</article-title>
          , Université d'Ottawa/University of Ottawa).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>De</given-names>
            <surname>Choudhury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Counts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            , &amp;
            <surname>Horvitz</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          (
          <year>2013</year>
          , May).
          <article-title>Social media as a measurement tool of depression in populations</article-title>
          .
          <source>In Proceedings of the 5th Annual ACM Web Science Conference</source>
          (pp.
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2013</year>
          , April).
          <article-title>A depression detection model based on sentiment analysis in microblog social network</article-title>
          .
          <source>In Pacific-Asia Conference on Knowledge Discovery and Data Mining</source>
          (pp.
          <fpage>201</fpage>
          -
          <lpage>213</lpage>
          ). Springer, Berlin, Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Cobb-Clark</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Schurer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>The stability of big-five personality traits</article-title>
          .
          <source>Economics Letters</source>
          ,
          <volume>115</volume>
          (
          <issue>1</issue>
          ),
          <fpage>11</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Golbeck</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robles</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edmondson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Turner</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2011</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Predicting personality from twitter</article-title>
          .
          <source>In Privacy, Security, Risk and Trust (PASSAT)</source>
          and
          <source>2011 IEEE Third Inernational Conference on Social Computing (SocialCom)</source>
          ,
          <year>2011</year>
          IEEE Third International Conference on (pp.
          <fpage>149</fpage>
          -
          <lpage>156</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francis</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Booth</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Linguistic inquiry and word count: LIWC 2001</article-title>
          . Mahway: Lawrence Erlbaum Associates,
          <volume>71</volume>
          (
          <year>2001</year>
          ),
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Coltheart</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>1981</year>
          ).
          <article-title>The MRC psycholinguistic database</article-title>
          .
          <source>The Quarterly Journal of Experimental Psychology Section A</source>
          ,
          <volume>33</volume>
          (
          <issue>4</issue>
          ),
          <fpage>497</fpage>
          -
          <lpage>505</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kosinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stillwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Graepel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Private traits and attributes are predictable from digital records of human behavior</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>110</volume>
          (
          <issue>15</issue>
          ),
          <fpage>5802</fpage>
          -
          <lpage>5805</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Shchebetenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Big Five and usage of the VK online social network</article-title>
          . Bulletin of South Ural State University, Series “Psychology” (pp.
          <fpage>73</fpage>
          -
          <lpage>83</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>P. T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>R. R.</given-names>
          </string-name>
          (
          <year>1992</year>
          ).
          <article-title>Normal personality assessment in clinical practice: The NEO Personality Inventory</article-title>
          .
          <source>Psychological assessment</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , ... &amp;
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of machine learning research</source>
          ,
          <volume>12</volume>
          (Oct),
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>