<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting Psychology Attributes of a Social Network User</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>8</lpage>
      <abstract>
        <p>Nowadays, the number of people using social network site increases every day. The social networking sites, such as Facebook or Twitter, are sources of human interaction, where users are allowed to create and share their activities, thoughts and place di erent information about themselves. However, most of this information remains unnoticed. In this work, we propose a machine learning approach to predict Big-Five personality using information from users accounts from the social network. The predictions can be used in di erent areas such as psychology, business, marketing.</p>
      </abstract>
      <kwd-group>
        <kwd>Social Networks</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Psychology</kwd>
        <kwd>Big Five Personality</kwd>
        <kwd>Shwartz Human Values</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Today we cannot imagine our life without social media resources. We
communicate with our friends, share information with them or just spend our free time
in social networks, for example on Facebook. Eventually, this process results in
the accumulation of a huge amount of information, which can tell almost
everything about the person, since people tend to write a lot about themselves. In
our work, we will use information from Vkontakte accounts (a social network,
which is popular in Russia, the analogue of Facebook) to predict information
about their owners. These predictions can be used in various areas of our lives,
such as psychology, business or marketing.</p>
      <p>
        In the present work, we will use the psychological model of the Big Five [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
and the psychological model of Schwartz [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to predict the portrait of a
person. This work is relevant because it allows to solve a lot of problems and has
a clear practical application. It is no coincidence that similar works have been
conducted for a long time and have become especially popular in recent years.
In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], IBM published the research report on predicting the socio-psychological
portrait on the scale of the Big Five based on the analysis of the owner's Twitter
account logs. In [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the authors published a paper on the construction of
correlations between the signs of a person in a social network and his psychological
characteristics on the "Big Five" scale. However, in the most works character
recognition is based on analysing the semantics of the text at the moment (for
example, posts on Twitter or statuses in social networks), and the accuracy of
the predictions was extremely low. In the present work, we want to use the
various characteristics of a person, which he or she points out in his or her own
account in the social network to predict the portrait of the person. The
disadvantage of this approach is that very often people do not write full information
about themselves or knowingly indicate false information about themselves. In
case of a large number of noisy data, the construction of a good predictive model
is a complex task. We use machine learning methods to implement predictions.
In particular, we have used such methods as Random forest, Gradient Boosting,
SVM.
      </p>
      <p>Chapter 2 presents the results achieved to date in this eld. In Chapter 3 we
give the main de nitions and describe predictions of the psychological scales of
Big-Five and the psychological scale of Schwartz. In Chapter 4 we give the main
conclusions on the work done.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        There are a lot of articles related to the study of the Big-Five scale. One of
the earliest ones is the work written by S.D. Gosling et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] examines the
psychological scale of Big-Five and the correlation with measurements on it for
di erent psychological tests. Their result noticed the fact that di erent surveys
do not have a high correlation with respect to the values of the characteristics
among themselves. Such a result can lead to understanding why person's
psychological portrait is di cult to predict. There were also many works on the
direct prediction of the psychological portrait, for example, in F. Mairesse et al.
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the prediction of a psycho-portrait was based on the linguistic analysis of
the Twitter posts. The prediction used a decision tree as a model while dividing
into several classes of the original Big-Five scale. We also used similar division
into several classes to improve quality of prediction model.
      </p>
      <p>
        Later on, in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], a model for predicting the characteristics of a Twitter account
has been constructed. Their work was divided into 2 parts. In the rst one they
tried to use the data from the account (some statistical information) and in the
second one they tried to predict the signs for certain linguistic variables received
during the processing of texts posts. Such work was conducted by IBM and
described in research report [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Similar research was conducted in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in which S. Bai et al. used information
from the popular social network in China to predict the psychological portrait
based on the Big- ve scale. They used a less accurate version of this scale
consisting of 2 components (0, 1) with decision trees prediction models only. Correlation
tables for some of the signs from users' accounts on the Facebook network and
the current scale of signs were built in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        More advanced machine learning were applied in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], in which the authors
tried to predict the characteristics of a person on the Big-Five scale using neural
networks, but also for reduced binary scale version from [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It should also be
noted that, in the work, the generated signs for the analysis of Twitter posts
were used as signs for the neural network.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Model Description</title>
      <p>In our work, we use the Big-Five scales and the Schwartz scale. Big-Five scale
contains 5 dimensions:
{ Extraversion vs. Introversion
{ Agreeableness vs. Antagonism
{ Conscientiousness vs. Lack of direction
{ Neuroticism vs. Sensible stability
{ Openness vs. Closeness to experience</p>
      <p>Schwartz human values scale contains 10 dimensions:
{ Self-Direction
{ Stimulation
{ Hedonism
{ Achievement
{ Power
{ Security
{ Conformity
{ Tradition
{ Benevolence
{ Universalism</p>
      <p>First of all, we consider the correlation between the initial signs and the
resulting values of psychological traits. Correlation tables were constructed for
the data, the most signi cant results are shown in Table 1.</p>
      <p>city sibling grandparent parent child sex
age</p>
      <p>
        Further, the scale data were considered from the point of view of regression
and classi cation for constructing the predictive model. We used algorithms of
machine learning, such as Gradient boosting, Random forest, SVM, Linear /
Logistic Regression from Scikit-learn [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The results were obtained by 5- and
10- fold cross-validation.
      </p>
      <p>For regression algorithms, we compare the results presented for R2 and MSE,
and accuracy, f1-score for classi cation algorithms. For each of the problems, we
consider some of the best features that could be predicted (three in each case).
In the case of the classi cation problem, we split our initial scale into 5 classes
and try to predict the hit of an element in 1 of the classes. The results obtained
are presented in Tables 2 and 3.</p>
      <p>It can be noted that even the best results for regression are extremely poorly
predicted by the signi cance of the features, which is indicated by Negative value
of R2. In the case of the classi cation problem, the results are encouraging,
but still the feature selection based on only social pro le is not su cient. It
is also worth noting that the results obtained for the remaining characteristics
do not provide su cient grounds for stating the conclusions about successful
classi cation.</p>
      <p>As a result of machine learning, the best methods were SVR, Random Forest
- in the case of a regression problem, and Gradient Boosting, Logistic regression
- in the case of a classi cation problem. Also, attempts were made to improve
the result using SVM (the use of di erent kernels, but this transformation did
not yield tangible results).</p>
      <p>mean R2 std R2 mean MSE std MSE
Self-Direction</p>
      <p>mean f1-score std f1-score mean accuracy std accuracy
Logistic Regression</p>
      <p>Random Forest
Gradient Boosting</p>
      <p>SVM</p>
      <p>Power
Logistic Regression</p>
      <p>Random Forest
Gradient Boosting</p>
      <p>SVM</p>
      <p>Conformity
Logistic Regression</p>
      <p>Random Forest
Gradient Boosting</p>
      <p>SVM
0.23 0.11 0.30 0.09
0.24 0.05 0.27 0.06
0.20 0.05 0.23 0.05
0.17 0.06 0.29 0.07
mean f1-score std f1-score mean accuracy std accuracy
0.24 0.03 0.66 0.01
0.23 0.05 0.63 0.02
0.22 0.04 0.61 0.03
0.22 0.05 0.69 0.03
mean f1-score std f1-score mean accuracy std accuracy
As a training dataset we used a small database of users, which were manually
marked by experts. Additional veri cation was obtained via psychological tests
processed by users who allowed access to their social network pro les when
passing the tests. Information from accounts in the social network we will get
by using VKontakte API. The information we will get contains the following
characteristics:
{ 'status', 'about', 'wall comments', 'relation',
{ 'has mobile', 'has photo', 'inspired by',
{ 'people main', 'life main', 'political',
{ 'tv', 'movies', 'games', 'music',
{ 'smoking', 'alcohol', 'religion',
{ 'home town', 'city', 'languages',
{ 'universities','education',
{ 'activities', 'interests',
{ 'occupation', 'career',
{ 'followers count',
{ 'sex', 'age'.</p>
      <p>We separate \complex" characteristics, such as
{ 'people main', 'life main', 'inspired by',
{ 'alcohol', 'smoking', 'political', 'religion',
{ 'languages', 'city', 'followers count'
For the items of this group we considered the value that forms persons attitude
to the given attribute (each characteristic is associated with a scale). Other
attributes we associated with 0,1-values, such that 1 corresponds to the indicated
attribute in the pro le, while 0 means that characteristic is not speci ed. We
also added signs of sex and age in addition to numerical data. After that we
form data based on the characteristics.</p>
      <p>The distribution of result features is shown in Fig.1
First of all, a correlation table was constructed for the signs, which revealed
signi cant values for some psychological signs and background data from social
networks. However, for the regression problem we have received unsatisfactory
results, regardless of the machine learning model. For the classi cation problem
our results can be considered satisfactory. We have received predictions on
certain grounds (in particular, Self-Direction, Power, Conformity). It is also worth
noting that such results can be a consequence of a small sample, so we cannot
unequivocally answer the question of the signi cance of the results obtained. As
possible explanations for this result, there are three main reasons. Firstly, this
can mean that in the future it is necessary to consider a more carefully selected
set of characteristics. Secondly, it should be noted that such a poor result can be
a consequence of a subjective scale of psychological signs, because tests cannot
accurately determine the psychological portrait of a person. Thirdly, which is
also important, the presence of unveri ed information in social accounts makes
it di cult to build a working model.
6</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>During the work, a set of characteristics for the prediction of an account from
a social network was generated, 2 models of prediction of features (regression
model, classi cation model) were constructed. Predictive models were SVM,
SVR, Logistic Regression, Random Forest, Gradient Boosting. As a result,
significant correlations were found in some of the psychological traits with signs from
social networks, and the model of prediction with classi cation was successfully
reviewed. Adding information from semantics of posted text information should
bene t our research in the future work.</p>
      <p>Acknowledgements. The article was prepared within the framework of the
Basic Research Program at the National Research University Higher School of
Economics (HSE) and supported within the framework of a subsidy by the
Russian Academic Excellence Project `5-100'.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          , T., Cheng, L.:
          <article-title>Big- ve personality prediction based on user behaviors at social network sites</article-title>
          .
          <source>arXiv preprint arXiv:1204.4809</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>L.R.:</given-names>
          </string-name>
          <article-title>An alternative" description of personality": the big- ve factor structure</article-title>
          .
          <source>Journal of personality and social psychology 59(6)</source>
          ,
          <volume>1216</volume>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Gosling</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rentfrow</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swann</surname>
          </string-name>
          , W.B.:
          <article-title>A very brief measure of the big- ve personality domains</article-title>
          .
          <source>Journal of Research in personality 37(6)</source>
          ,
          <volume>504</volume>
          {
          <fpage>528</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kalghatgi</surname>
            ,
            <given-names>M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramannavar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sidnal</surname>
            ,
            <given-names>N.S.:</given-names>
          </string-name>
          <article-title>A neural network approach to personality prediction based on the big- ve model</article-title>
          .
          <source>International Journal of Innovative Research in Advanced Engineering (IJIRAE) 2</source>
          (
          <issue>8</issue>
          ),
          <volume>56</volume>
          {
          <fpage>63</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mairesse</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Words mark the nerds: Computational models of personality recognition through language</article-title>
          .
          <source>In: Proceedings of the Cognitive Science Society</source>
          . vol.
          <volume>28</volume>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>T.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lefringhausen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferenczi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>The big ve, self-esteem, and narcissism as predictors of the topics people write about in facebook status updates</article-title>
          .
          <source>Personality and Individual Di erences 85</source>
          ,
          <volume>35</volume>
          {
          <fpage>40</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pedregosa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gramfort</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grisel</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prettenhofer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubourg</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanderplas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Passos</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cournapeau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brucher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perrot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duchesnay</surname>
          </string-name>
          , E.:
          <article-title>Scikit-learn: Machine learning in python</article-title>
          .
          <source>J. Mach. Learn. Res</source>
          .
          <volume>12</volume>
          ,
          <issue>2825</issue>
          {2830 (Nov
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bilsky</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          : (
          <year>1987</year>
          ).
          <article-title>toward a universal psychological structure of human values</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>53</volume>
          (
          <issue>3</issue>
          ),
          <volume>550</volume>
          {
          <fpage>562</fpage>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Sumner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Byers</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boochever</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Park, G.J.:
          <article-title>Predicting dark triad personality traits from twitter usage and a linguistic analysis of tweets</article-title>
          .
          <source>In: Machine learning and applications (icmla)</source>
          ,
          <year>2012</year>
          11th international conference on.
          <source>vol. 2</source>
          , pp.
          <volume>386</volume>
          {
          <fpage>393</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Identifying user needs from social media</article-title>
          .
          <source>IBM Research Division</source>
          , San Jose p.
          <volume>11</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>