<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Estimating the Political Orientation of Twitter Users in Homophilic Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Morteza Shahrezaye</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Hegelich</string-name>
          <email>simon.hegelich@hfp.tum.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bavarian School of Public Policy at Technical University of Munich Richard-Wagner street 1 80333 Munich</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>There have been many efforts to estimate the political orientation of citizens and political actors. With the burst of online social media use in the last two decades, this topic has undergone major changes. Many researchers and political campaigns have attempted to measure and estimate the political orientation of online social media users. In this paper, we use a combination of metric learning algorithms and label propagation methods to estimate the political orientation of Twitter users. We argue that the metric learning algorithm dramatically increases the accuracy of our model by accentuating the effect of homophilic networks. Homophilic networks are user clusters formed due to cognitive motivational processes linked with cognitive biases. We apply our method to a sample of Twitter users in Germany's six-party political sphere. Our method obtains a significant accuracy of 62% using only 40 observations of training data for each political party.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Measuring and estimating the political orientation of normal
citizens and political actors has always been a relevant
question. The answer to this question is essential for electoral
campaigns
        <xref ref-type="bibr" rid="ref21 ref7 ref9">(Gayo Avello, Metaxas, and Mustafaraj 2011;
Dokoohaki et al. 2015; Papakyriakopoulos et al. 2018)</xref>
        ,
agenda setting, policy making
        <xref ref-type="bibr" rid="ref20">(McCombs 2014)</xref>
        , and
research purposes
        <xref ref-type="bibr" rid="ref1 ref13 ref9">(Golbeck and Hansen 2011; Barbera´ 2014;
Hegelich and Shahrezaye 2015)</xref>
        . The methodological efforts
to answer this crucial question possess three qualities.
      </p>
      <p>
        The first quality is related to the number and type of
inputs in the algorithm: What type of features are considered
while estimating the latent political orientation of the users?
The second quality is if the method is designed to estimate
the political orientation of a specific group of political
actors
        <xref ref-type="bibr" rid="ref10 ref27">(Wong et al. 2013; Groseclose and Milyo 2005)</xref>
        or a
more general group of citizens
        <xref ref-type="bibr" rid="ref1">(Barbera´ 2014)</xref>
        . If a method
is designed based on a specific group of political actors or
citizens, it cannot be generalized to estimate the political
orientation of other groups of political actors or citizens. Cohen
and Ruths have presented that methods that have accuracy
greater than 90% in estimating if a Twitter user is a
Democrat or Republican, would have accuracy level of less than
65% when applied on general Twitter users. The last quality
is if the method measures the political orientation on a one
dimensional or a multidimensional latent space. Most of the
literature has been designed based on the two-party
political system of the United States. Thus, they are inherently
designed to estimate a one-dimensional latent variable.
      </p>
      <p>
        In this work, we use a combination of metric learning
algorithms and label propagation methods to estimate the
political orientation of Twitter users. Our method has three
distinguishing features. First, the method requires a minimal
number of features as training data because it exploits the
homophilic structure of social networks
        <xref ref-type="bibr" rid="ref18 ref8">(Geschke, Lorenz,
and Holtz 2018; Madsen, Bailey, and Pilditch 2018)</xref>
        .
Second, the proposed method estimates on a multidimensional
latent space; therefore, the proposed method can be used to
estimate the political orientation of users in a multiparty
political system. The third feature is that our method is
extendable to multiple groups or cluster of users. Our method can
estimate the political orientation of users even if the target
users have zero political activity on the platform.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>We use a combination of metric learning algorithms with
label propagation methods to estimate the political orientation
of Twitter users. The goal of label propagation algorithms is
to estimate the labels of a large set of unlabeled observations
from the small set of labeled observations.</p>
      <p>Suppose there are l labeled observations
(x1; y1); : : : ; (xl; yl) and u unlabeled observations
such that l &lt; u, and n = l + u. Consider a connected
graph G = (V; E) with nodes L = f1; : : : ; lg and
U = fl + 1; : : : ; l + ug corresponding, respectively, to
the labeled or training observations and unlabeled or test
observations. A label propagation algorithm propagates
the labels for the set U , based on the distances between
its observations to the observations in L. Within the label
propagation algorithm, the labels of the vertices in set
L would be fixed, but the labels of the set U would be
estimated based on a function of their distance to set L.</p>
      <p>Let n be the total number of Twitter users we have
including l users for whom we already know their political
orientation and u users for whom we want to estimate their
political orientation. We use only the structure of the friends’
network to estimate the political orientations. Let F be the
set of friends of all n users with size m. Therefore, we can
create the binary matrix A with dimension n m, which
would represent the friends of each of the n users. Before
constructing graph G from matrix A, we transform matrix
A by using a proper metric learning algorithm.</p>
      <p>
        The reason for transforming matrix A is that we believe
there are hidden information within the network structure,
which we could use to increase the estimation accuracy. By
contrast with the rational choice theory, the human
judgment is influenced by various cognitive biases, prior
judgments, environmental features, and stimulus-feedback loops
        <xref ref-type="bibr" rid="ref13 ref16 ref5">(Kenrick et al. 2010; Donkin, Heathcote, and Brown 2015)</xref>
        .
Cognitive biases reproduce human judgments that could be
systematically different from rational reasoning
        <xref ref-type="bibr" rid="ref12 ref13 ref15">(Kahneman
and Tversky 1973; Haselton, Nettle, and Murray 2015)</xref>
        . The
cognitive biases make the human brain process the
information in a distorted manner compared with an objective
reality
        <xref ref-type="bibr" rid="ref24 ref9">(Sharot, Korn, and Dolan 2011)</xref>
        . Although there is
a list of cognitive biases that affect the online activity of
the users, we are specifically interested in cognitive biases
related to self-categorization. Self-categorization describes
the motivations and circumstances under which
communities with shared identities form. The self-categorization
theory articulates that the spectrum of human behavior can be
analyzed from a pure interpersonal or individualistic and a
pure intergroup or collectivist perspective. Humans have the
desire for a positive and secure self-concept; therefore, they
connect with individuals that confirm their pre-existing
attitudes, verify their self-views, and increase their social
identity. The aforementioned behaviour is called confirmation
bias
        <xref ref-type="bibr" rid="ref8">(Geschke, Lorenz, and Holtz 2018)</xref>
        . In addition, “If we
are to accept that people are motivated to have a positive
self-concept, it flows naturally that people should be
motivated to think of their groups as good groups”
        <xref ref-type="bibr" rid="ref14">(Hornsey
2008)</xref>
        . Striving for a positive and secure self-concept,
humans’ collectivist behaviors contribute to the formation of
online and offline communities with shared social identities
        <xref ref-type="bibr" rid="ref23">(Ridings and Gefen 2004)</xref>
        . Consequently, users with similar
labels, that is, similar political preferences, are expected to
be relatively closer to each other. Therefore, if we were to
supposedly apply a k-nearest neighbors learning method, it
makes sense to use a distance function that interprets
similar users closer to each other. Instead of using an
off-theshelf distance function such as Euclidean distance, we use
an alternative distance function that guarantees higher
accuracy for the labeled or training observations after running the
learning method.
      </p>
      <p>A brief description of the steps of our method is as
follows. First, we acquire matrix A, which includes the labeled
observations and the unlabeled observations as rows.
Second, we learn the optimized distance or metric function that
guarantees higher accuracy within the labeled observations
by exhausting the special structure of homophilic networks.
We transform matrix A by using the learned metric to
construct graph G. Finally, we apply the learning method or the
label propagation algorithm.</p>
      <sec id="sec-2-1">
        <title>Metric Learning for Large Margin Nearest Neighbor</title>
      </sec>
      <sec id="sec-2-2">
        <title>Classification (LMNN)</title>
        <p>
          The accuracy of each learning algorithm is a function of the
distance function or the metric used to compute the distance
between the observations. The metric learning algorithm we
use is based on the following: a precise k-nearest neighbors
classification will correctly classify a labeled observation if
its k-nearest neighbors share the same label. The algorithm
then attempts to increase the number of labeled observations
with this property by learning a linear transformation of the
input space that proceeds the final learning method. The
linear transformation of LMNN is derived by maximizing a loss
function with two terms. The first term minimizes the large
distances between observations within class, and the second
term maximizes the distances between the observation
between the classes
          <xref ref-type="bibr" rid="ref25">(Weinberger and Saul 2009)</xref>
          .
        </p>
        <p>
          In general, metric learning algorithms estimate the
positive semidefinite transformation matrix M such that the
distance between two observations, xi and xj , is derived by the
Mahalanobis distance,
dM(xi; xj ) =
q
(xi
xj )T M(xi
xj )
which follows certain features. If we replace M with
the identity matrix, the resulting metric would be Euclidean
metric. LMNN learns a linear transformation matrix M,
such that the training or labeled observation satisfies the
following items
          <xref ref-type="bibr" rid="ref25">(Weinberger and Saul 2009)</xref>
          :
        </p>
        <p>Each labeled observation should share the same label as
its k- nearest neighbors. This is achieved by introducing
a loss function that penalizes large distances between
observations belonging to the same class,</p>
        <p>X
j i
pull(L) =
jjL(xi
xj )jj2
where j i indicates that j is an observation that we
desire to be close to i, and L is the function representing
the transformation by matrix M.</p>
        <p>The labeled observations with different labels should be
significantly separated. This separation is achieved by
introducing a loss function that penalizes small distances
between observations belonging to different classes,
push(L) =</p>
        <p>X</p>
        <p>X[1+jjL(xi xj )jj2
jjL(xi xl)jj2]
i;j i l
where the inner sum iterates over all the observations with
a different class to i, and l invades the perimeter of i and j
plus unit margin. In other words, the observation l satisfies
jjL(xi
xl)jj2
jjL(xi
xj )jj2 + 1
The final loss function is a weighted combination of the two
defined components,
(L) = (1
) pull(L) +
push(L)
Although the general loss function above is not convex, by
limiting the solution space to positive semidefinite matrices,
the loss function will be a convex function.</p>
        <p>The solution to the minimization of the loss function,
given the labeled subset of A, is the desirable matrix M.
We transform matrix A to obtain matrix AM by</p>
        <p>We construct graph G using the AM of size n m by
using the nearest neighbor graph method. In other words,
using n rows of AM, we define n vertices of G and then
define edges between each vertex and its kG nearest neighbors
by using the Euclidean distance function.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Label Propagation Using Gaussian Fields and Harmonic</title>
      </sec>
      <sec id="sec-2-4">
        <title>Functions</title>
        <p>
          The goal of applying a label propagation algorithm to a
graph is to estimate the labels of unlabeled vertices by using
their connections to the few labeled vertices. This problem
is usually formulated as an iterative process within which
the labels are gradually diffused over the matrix, such that
the state of the graph would converge to a stationary state.
This iterative process might have an analytical solution that
would be more efficient than applying the algorithm
iteratively
          <xref ref-type="bibr" rid="ref2 ref29">(Barrett et al. 1994; Zhu and Ghahramani 2002)</xref>
          . The
most crucial implication of a label propagation algorithm
for our question regarding estimating political orientation of
Twitter users is that the only requirement for estimating the
political requirement of a user is that the user should be
connected to graph G. Hence, the user should not necessarily
have politicians or other political actors as friends.
        </p>
        <p>The algorithm we use for label propagation is based on
Zhu, Ghahramani, and Lafferty. Let the simple graph G =
(V; E) and the set of the labeled and unlabeled vertices, L
and U , be as defined. The goal is to compute the real-valued
function f : V ! R on the simple graph G. f must assign
the same given labels for the set L or fl(i) yi for i 2 l. To
estimate the function f they defined the energy function
E(f ) = 1 X wi;j (f (i)
2
i;j
f (j))2
and the Gaussian field
p (f ) =
e E(f)
Z
where is an inverse temperature function and Z =
Rf exp( E(f )) which normalizes over all functions
constrained to the constraint fl(i) yi on the labeled vertices.
Then, they demonstrate the result of the minimization
f = arg min E(f )</p>
        <p>f
which is a harmonic function that satisfies the constraint
fl(i) yi on the labeled vertices. The harmonic property
implies that the value of f at each unlabeled vertex is the
average of f at neighboring vertices. Therefore, the estimated
labels would be a function of the similarity of all
neighboring vertices.</p>
        <p>The estimated f has an interpretation within the
framework of random walks. The estimated f (i) for an unlabeled
vertex i 2 U would be a vector of size equal to number of
possible classes. The jth element of f (i) would be the
probability that a particle that started at vertex i would first hit
a vertex with class j. Therefore, the resulting algorithm can
be used to estimate the political orientation of a user in a
multidimensional latent space.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data and Results</title>
      <sec id="sec-3-1">
        <title>Data Preparation</title>
        <p>We require two sets of data for training and testing. We
acquire both sets from the public Twitter API. In the first step,
we obtained the list of all the members of the main and
local German parliaments who are available on Twitter. This
list contains 623 Twitter users from one of the six parties
CDU/CSU, SPD, Gru¨ne, Linke, FDP and AfD.</p>
        <p>From a database of German political Tweets, we obtained
a list of 400,000 random Twitter users. We downloaded the
list of all their friends and their last 4,000 Tweets by
using the public API. We counted how many times each user
retweeted the Tweets of members of each of the political
parties we acquired in the first step. If a user has retweeted
a minimum of five Tweets from members of party j but no
retweets from other parties, we tag this user as a user with
a political orientation to party j. From the 400,000 initial
users, we could label 8,146 based on the mentioned
heuristic.</p>
        <p>To reduce the complexity of the computations, we
reduced the sample size to 50,000 from 400,000. Thus, we
created matrix A using 50,000 random users including all
of the 8,146 labeled users. Matrix A has at this step 50,000
rows as users, which we want to use for our training and test
set, and 7,194,153 columns as the friends. To further reduce
the complexity of the computations, we removed the friends
who are friends of less than 0.01% of the users. The final
matrix A has the dimension 50,000 552,136.</p>
        <p>We confirm that our test data has a minor bias in the sense
that we already know our test data includes users who have
engaged in some type of political activity. This assumption
is because these users are randomly chosen from a database
of German political Tweets. On the other side, this bias is
mildly mitigated in two steps. First, matrix A is created by
a list of friends of all 50,000 random users and not only the
friends of the labeled 8,146 users. Thus, the feature sets are
from a bigger set of observations. Second, we added some
randomness by removing some columns of matrix A in the
final step.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Metric Learning and Label Propagation</title>
        <p>We resampled 60 users per political party out of the 8,146
labeled users of A. We learned matrix M based on the 240
users. Next, we transformed the whole matrix A using M
by applying</p>
        <p>M
Using the transformed AM, we made a 10-nearest neighbors
graph using a Euclidean distance function to make graph G.
Finally, we applied the label propagation algorithm on G
that has 50,000 vertices, out of which, the labels of 240 are
introduced to the algorithm. The labels of the other 49,760
are estimated using the label propagation algorithm.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Results</title>
        <p>We performed the resampling and the computations 10 times
to make sure the results are robust. For each trial, we
applied a random forest classifier on the 240 training data as a
random forest
label propagation
random forest
label propagation
AM (transformed)
benchmark result. We also applied the random forest
classifier and the label propagation method on A directly to
improve our understanding regarding how much the LM N N
metric learning method contributes to the accuracy of the
results. Table 1 shows the average accuracy of the
estimations on the remaining 8,146-240=7,906 labeled users with
a known political orientation.</p>
        <p>Referring to Table 1, we observe that the transformation
increases the accuracy of the random forest classifier and
the label propagation algorithm. We also observe that the
combination of the metric learning algorithm and the label
propagation method results to a much higher accuracy of
estimation.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>In this paper, we proposed a new method to estimate the
political orientation of Twitter users. Our method has many
distinguishing features: The method requires few training
observations, requires few learning features, is based on a
multidimensional latent space, and is easily expendable to
new users even if they have zero political activity on
Twitter.</p>
      <p>Based on Table 1, the high accuracy of the model is due
to the transformation of the initial matrix using the function
learned by the LMNN algorithm. The cost function of the
LMNN algorithm has two parts. One part pulls the
observations of the same class closer to each other, and the other
part pushes the observations of different classes far apart.
Additionally, since the LMNN algorithm is based on
optimizing a k-nearest neighbor model on the training
observations, the trained matrix M transforms the observations
based on their relation to other observations in their vicinity
and not the whole dataset. These characteristics have crucial
implications reagarding the accuracy of our estimation.</p>
      <p>As aforementioned, the initial matrix, A, has a special
structural feature because it represents a homophilic social
network, which means that users with similar political
identity are assumed to demonstrate similar behavior on
Twitter. Therefore, we expected that users with similar political
identity would follow similar politicians, similar celebrities,
similar sportsmen, and so forth.</p>
      <p>
        When we apply the LMNN algorithm to this
homophilic network, we accentuate the extant distinctive
features formed due to the existing cognitive biases in
selfcategorization and group formation
        <xref ref-type="bibr" rid="ref18 ref8">(Geschke, Lorenz, and
Holtz 2018; Madsen, Bailey, and Pilditch 2018)</xref>
        .
      </p>
      <p>The matrix M learns different combinations of features
that help distinguish normal Twitter users based on their
political orientation. The matrix M also allows different
combination of features for each class because it is based on a
k-nearest neighbor algorithm that considers a bounded
proximity of the users. Our model detects the political orientation
of users with high accuracy, and by far outperforms other
algorithms that have been applied to this task.</p>
      <p>Due to the use of label propagation algorithm, this model
can be later applied on any new user e to estimate her or his
political orientation, as long as e is connected to the graph G.
More generally, to predict the political orientation of user e,
we must find a new set of users including e, forming a small
graph g connected to the initial graph G.</p>
      <p>This study provides valuable insights into the study of
user behavior on online social networks. This study
illustrates, that using mathematical algorithms that exhaust
properties of social theories, we can improve the performance of
models explaining human behavior. Furthermore, this study
contradicts the general claim that a huge amount of data is
required to make accurate predictions on social and
political behavior. Finally, our method provides a novel technique
to assign political partisanship, by having as input only the
network of interpersonal connections.</p>
      <p>Zhu, X.; Ghahramani, Z.; and Lafferty, J. D. 2003.
Semisupervised learning using gaussian fields and harmonic
functions. In Proceedings of the 20th International
conference on Machine learning (ICML-03), 912–919.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Barbera</surname>
            ´,
            <given-names>P.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Birds of the same feather tweet together: Bayesian ideal point estimation using twitter data</article-title>
          .
          <source>Political Analysis</source>
          <volume>23</volume>
          (
          <issue>1</issue>
          ):
          <fpage>76</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Barrett</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Berry,
          <string-name>
            <surname>M. W.</surname>
          </string-name>
          ; Chan,
          <string-name>
            <given-names>T. F.</given-names>
            ;
            <surname>Demmel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Donato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Dongarra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Eijkhout</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ;
            <surname>Pozo</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Romine,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Van der Vorst</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>1994</year>
          .
          <article-title>Templates for the solution of linear systems: building blocks for iterative methods</article-title>
          , volume
          <volume>43</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ruths</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Political orientation inference on twitter: It's not easy</article-title>
          .
          <source>Proc. of ICWSM.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          2015.
          <article-title>Predicting swedish elections with twitter: A case for stochastic link structure analysis</article-title>
          .
          <source>In Advances in Social Networks Analysis and Mining (ASONAM)</source>
          ,
          <year>2015</year>
          IEEE/ACM International Conference on,
          <fpage>1269</fpage>
          -
          <lpage>1276</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Donkin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Heathcote</surname>
            ,
            <given-names>B. R. A.</given-names>
          </string-name>
          ; and Brown,
          <string-name>
            <surname>S. D.</surname>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>The Oxford Handbook of Computational and Mathematical Psychology</source>
          <volume>121</volume>
          -141.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Gayo</given-names>
            <surname>Avello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ; Metaxas, P. T.; and
            <surname>Mustafaraj</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Geschke</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lorenz</surname>
            , J.; and Holtz,
            <given-names>P.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>The triplefilter bubble: Using agent-based modelling to test a metatheoretical framework for the emergence of filter bubbles and echo chambers</article-title>
          .
          <source>British Journal of Social Psychology.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Golbeck</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Computing political preference among twitter followers</article-title>
          .
          <source>In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>
          ,
          <volume>1105</volume>
          -
          <fpage>1108</fpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Groseclose</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Milyo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>A measure of media bias</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>The Quarterly Journal of Economics</source>
          <volume>120</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1191</fpage>
          -
          <lpage>1237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Haselton</surname>
            ,
            <given-names>M. G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nettle</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Murray</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>The evolution of cognitive bias</article-title>
          .
          <source>The handbook of evolutionary psychology 1-20.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Hegelich</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Shahrezaye</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>The communication behavior of german mps on twitter: Preaching to the converted and attacking opponents</article-title>
          .
          <source>European Policy Analysis</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>155</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hornsey</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <year>2008</year>
          .
          <article-title>Social identity theory and selfcategorization theory: A historical review</article-title>
          .
          <source>Social and Personality Psychology Compass</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>204</fpage>
          -
          <lpage>222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Kahneman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Tversky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>1973</year>
          .
          <article-title>On the psychology of prediction</article-title>
          .
          <source>Psychological review 80</source>
          <volume>(4)</volume>
          :
          <fpage>237</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Kenrick</surname>
          </string-name>
          , D. T.;
          <string-name>
            <surname>Neuberg</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Griskevicius</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>D. V.</given-names>
          </string-name>
          ; and Schaller,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Goal-driven cognition and functional behavior: The fundamental-motives framework</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>Current Directions in Psychological Science</source>
          <volume>19</volume>
          (
          <issue>1</issue>
          ):
          <fpage>63</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Madsen</surname>
            ,
            <given-names>J. K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bailey</surname>
            ,
            <given-names>R. M.</given-names>
          </string-name>
          ; and Pilditch,
          <string-name>
            <surname>T. D.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Large networks of rational agents form persistent echo chambers</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>Scientific reports 8</source>
          (
          <issue>1</issue>
          ):
          <fpage>12391</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>McCombs</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Setting the agenda: Mass media and public opinion</article-title>
          . John Wiley &amp; Sons.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Papakyriakopoulos</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hegelich</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Shahrezaye,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Serrano</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. C. M.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Social media and microtargeting: Political data processing and the consequences for germany.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>Big Data &amp; Society</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Ridings</surname>
            ,
            <given-names>C. M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gefen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Virtual community attraction: Why people hang out online</article-title>
          .
          <source>Journal of Computermediated communication</source>
          <volume>10</volume>
          (
          <issue>1</issue>
          ):
          <fpage>JCMC10110</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Sharot</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Korn</surname>
            ,
            <given-names>C. W.</given-names>
          </string-name>
          ; and Dolan,
          <string-name>
            <surname>R. J.</surname>
          </string-name>
          <year>2011</year>
          .
          <article-title>How unrealistic optimism is maintained in the face of reality</article-title>
          .
          <source>Nature</source>
          neuroscience
          <volume>14</volume>
          (
          <issue>11</issue>
          ):
          <fpage>1475</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K. Q.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Saul</surname>
            ,
            <given-names>L. K.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Distance metric learning for large margin nearest neighbor classification</article-title>
          .
          <source>J.</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Mach</surname>
          </string-name>
          .
          <source>Learn. Res</source>
          .
          <volume>10</volume>
          :
          <fpage>207</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Wong</surname>
            ,
            <given-names>F. M. F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>C. W.</given-names>
          </string-name>
          ; Sen,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; and Chiang,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>ICWSM</source>
          <volume>13</volume>
          :
          <fpage>640</fpage>
          -
          <lpage>649</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ghahramani</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <year>2002</year>
          .
          <article-title>Learning from labeled and unlabeled data with label propagation.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>