<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.1145/1235</article-id>
      <title-group>
        <article-title>Birds of a Feather Tweet Together: Computational Techniques to Understand User Communities in Social Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Burth Kurka</string-name>
          <email>d.kurka@ic.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan Godoy</string-name>
          <email>godoy@dca.fee.unicamp.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fernando J. Von Zuben</string-name>
          <email>vonzuben@dca.fee.unicamp.br</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>(1) University of Campinas</institution>
          ,
          <addr-line>(2) CPqD Foundation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>(1) University of Campinas, (2) Imperial College London</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Campinas</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>1691</volume>
      <fpage>21</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>The study of social systems shows that there is a relationship of mutual influence between social connections and individual behavior, known as homophily. In this work, we developed a methodology to allow the analysis of interests of groups of users in Twitter network, based on automatic community detection and tweets ranking. The techniques presented reveal evidences that the presence of communities is related to topic specialization, and allow the characterization of elaborate profiles of groups of users based only on their location on the network.</p>
      </abstract>
      <kwd-group>
        <kwd>Online social networks</kwd>
        <kwd>social network analysis</kwd>
        <kwd>community detection</kwd>
        <kwd>homophily</kwd>
        <kwd>complex systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The experience of the last decade has shown that online
social networks (OSNs) are not only useful for the
amusement of Internet users, but can also be a valuable source of
data for the study of social systems. The millions to billions
of users who access OSNs services everyday are providing
researchers an unprecedented possibility to gather
information from human activity and social behavior, enabling the
investigation of complex issues [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        One of the interesting topics that can be explored in this
context is the creation and diffusion of content. OSNs users
are constantly involved in complex dynamics of message
sharing, which may result in the emergence of new trends,
collective mobilization and opinion formation.
Understanding the essential mechanisms of this process can be useful to
areas such as politics [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and marketing campaigns [
        <xref ref-type="bibr" rid="ref10 ref20">10, 20</xref>
        ].
      </p>
      <p>This paper explores the interplay between social
connections and content shared by individuals in an online social
network, which is one aspect of social communication. We
investigated here how it is possible to understand
behavior and interests of a group of individuals based on their
social connections, paying special attention to the role of
homophily – the tendency for someone to establish
connections to similar peers. Analytical tools were developed –
and applied to real data extracted from OSNs – to pinpoint
content relevant to the understanding of interests of
communities and users composing them. The techniques used are
based solely on the knowledge of the users that shared each
message, not resorting to their contents. We believe that, in
addition to indicating the relevance of sharing information
for the classification of messages, such tools can be useful to
the practical study of social systems and to the development
of new applications, such as recommendation systems.</p>
      <p>The paper is structured as follows. In Section 2 we
introduce relevant literature on OSN and homophily. In
Section 3, we present the methods used to acquire and analyze
data. Then, in Section 4, we report and analyze the results
obtained using our methodology. Finally, in Section 5, we
present final remarks, discussing implications of this work
and possible future directions.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>BACKGROUND</title>
      <p>
        An increasing number of studies have been conducted over
the last years focusing on online social networks [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Dynamics of information diffusion [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], public opinion prediction [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ],
collective sentiment analysis [
        <xref ref-type="bibr" rid="ref18 ref9">9, 18</xref>
        ] and the formation of
social structures [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] are among the subjects explored.
      </p>
      <p>
        An important aspect investigated is the presence of
selforganized processes. Despite not having central controllers
that rule on how content is disseminated or connections
between users are created, OSNs display many organized
behaviors. Common examples of such process are the collective
curation of contents [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and information diffusion cascades
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Communities and Homophily</title>
      <p>
        Self-organizing processes also take part on how
connections are formed in an OSN. Usually, social connections are
not created uniformly between all individuals, but are
concentrated in few hubs [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Thus, when analysing topologies
of real networks, communities of individuals which are more
likely to be linked between each other, with fewer
connections between individuals from different communities, are
found. Social systems exhibit communities in many
different levels: people inside families, sharing specific common
interests, or living at the same city or nation have many more
connections to other in-groups than to out-groups [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Considering the flow of information that travels over the
network, groups with highly connected individuals may
imply a redundancy of communication channels, possibly
enhancing information that pass through or are generated by
these groups. In the case of online social networks, it can
imply that a content may be more easily proliferated and
reinforced inside a community after it is shared by a member
of such group. While communities are expected to influence
on how their members receive and process information, it
is also believed that community formation is influenced by
pre-existing affinities between members [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        Researchers identified in social networks outside the
virtual world a tendency (not only within communities) of
individuals with common interests to be usually connected to
each other [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Such phenomenon is called homophily and
is observed, also, in OSNs.
      </p>
      <p>
        Kwak et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], in an early study analyzing Twitter’s data,
showed evidences of homophily among users with same
localization and same number of friends (popularity). Also
on Twitter, Wu et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], found a strong tendency where
users belonging to a same category (e.g artists,
organizations, bloggers) would communication among themselves.
Romero et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] studied the relationship between the
(explicit) network of friendship and the (implicit) network of
topical affiliations – i.e., the communities formed by users
interested in a common topic. They showed that both
networks have considerable intersection and that users tend to
connect to other users with common interests. This
correlation allows the prediction of friendship connections from
hashtag diffusions and also the forecast of the future
popularity of a hashtag from the friends network.
      </p>
      <p>
        Bollen et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] verified that users’ emotions is strongly
correlated with social connections, showing that users
considered happy tend to be linked to each other. Salath´e et
al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] explored how a network with signs of homophily
interfere on the spread of sentiment towards a new vaccine,
showing how negative opinions can be reinforced in such
environments.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Communities Detection</title>
      <p>
        Some of the most common approach to community
detection are modularity-based algorithms [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which look to
partition a network in communities so that the summed weight
of all connections between two communities is minimized.
This approach, as most other techniques usually applied for
community detection, however, is insensitive to direction of
connections in the network. In systems where patterns of
flow among individuals are relevant, however, ignoring
connection direction may disregard information valuable to the
comprehension of collective behavior.
      </p>
      <p>
        In order to address this issue, Rosvall and Bergstrom
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposed a flow-based method, which defines a
community as a set of individuals “among which information
flows quickly and easily”. Their algorithm takes advantage
of both direction and weights of connections, using
information theoretical measures and a random walk as a proxy
of information flow. Using a greedy search, they look for a
partition of individuals that define a two-level description –
the first level indicating the partition and the second
indicating a individual inside the partition – that minimizes the
expected description length of a random walk in the graph.
This partitions are, thus, the communities of individuals in
the network.
3.
3.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>METHODS</title>
    </sec>
    <sec id="sec-6">
      <title>Data Acquisition</title>
      <p>As Twitter has plenty of public data available online, it is
a good source for creating a database of social events.
However, the rate limits imposed by its API1 hamper the
download of large volumes of data from events that took place in
the past. An alternative is to use Twitter’s Streaming APIs,
which allows the download of messages in real time, as they
are posted.</p>
      <p>Therefore, in order to collect a satisfactory amount of data
to be analyzed, we decided to track the interactions between
a popular user account producer of original content and the
users that share (i.e. retweet) these messages. As popular
Twitter accounts interact with many users daily, this
approach revealed to be an effective way for collecting message
diffusion processes and user interactions as they happen.</p>
      <p>We chose the Brazilian largest newspaper Twitter account,
Folha de Sa˜o Paulo2, and collected: original messages posted,
retweets of those messages posted by other users, account
details of those users and the relationships (followers and
followees) of all of them.
3.2</p>
    </sec>
    <sec id="sec-7">
      <title>Automatic Topic Classification</title>
      <p>As with many other newspapers, almost all messages
published by the chosen source are headlines, followed by a link
to the newspaper’s website with the news’ full content. As
the news articles on the website belong to thematic
categories (newspaper’s sections), it was possible to
automatically attribute a class to each tweet, based on these
categories. This procedure was carried out to all tweets and six
most common topics were verified, namely: “daily life news”,
“sports”, “world”, “politics”, “entertainment” and “market”.
3.3</p>
    </sec>
    <sec id="sec-8">
      <title>Detecting common interests in communities</title>
      <p>
        After collecting and classifying all data, we then
evaluated whether retweeting behaviors of communities’
members are coherent among themselves regarding the subjects
of the shared messages. In order to detect groups of tightly
connected users from the social connections observed, we
executed Rosvall and Bergstrom’s community detection
algorithm [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] in our social network, an algorithm focused on
locating groups of users among which information can flow
more quickly.
      </p>
      <p>
        Understanding the interests of a community is a hard task.
To address such issue, we propose here an adaptation of a
well-known statistics from text mining, the term
frequencyinverse document frequency (tf-idf ) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>The tf-idf method usually considers a corpus of
documents, each composed of different terms. By comparing the
frequency of specific terms inside and outside documents,
values are assigned to each term, pondering its importance
1https://dev.twitter.com/rest/public/rate-limiting
2https://twitter.com/folha
for each document. The tf-idf of a term t in a specific
document d is calculated as follows:</p>
      <p>tf idfd(t) = tfd(t) × idf (t),
where tfd(t) is a value that is higher the higher the frequency
of the term t in d while idf (t) is inversely proportional to
the frequency of t in all documents of a corpus.</p>
      <p>In the traditional use of the tf-idf algorithm in text-mining,
a term is a word and a set of terms is a textual document
(e.g. a book, or a webpage). This creates a document-term
matrix, where each row represents a document, each column
a word and each cell, indexed by (i, j), the frequency of the
term j in document i. For the case of our application, we
adapted this approach by considering each tweet as a term
and each community as a document. Therefore, we built a
community-tweet matrix where each cell (i, j) represents the
number of retweets of tweet j in community i.</p>
      <p>Thus, when the same operations of tf-idf are applied to a
community-tweet matrix, tweets that are both highly shared
by members of a community and more akin to such
community’s behavior are highlighted with higher scores than the
others.</p>
    </sec>
    <sec id="sec-9">
      <title>RESULTS</title>
    </sec>
    <sec id="sec-10">
      <title>Database Description</title>
      <p>From March 19, 2014 to September 21, 2014, all messages
(tweets) posted by the source account, as well as any share
of this content by other users (retweets) were collected.</p>
      <p>It was possible to track, from the data collected, a large
amount of information diffusion processes triggered by the
observed source. During the observed period, 13463 distinct
and original messages posted by the source account were
collected. From this set, a series of filters were sequentially
applied, forming a more appropriate dataset for the work,
as described below:
• Filter 1 - Only messages which had received at least
20 retweets were selected, resulting in a group of 4671
distinct messages;
• Filter 2 - Messages not belonging to one of the six main
categories (“everyday news”, “sports”, “world”,
“politics”, “entertainment” and “market”) were removed from
the above set, resulting in a group of 3185 messages;
• Filter 3 - Users who retweeted more than 60% of the
publications were considered automated scripts (bots)
and therefore removed from the database and from the
retweets count (just one user was identified as such);
• Filter 4 - As in 2014 Brazil held the Football World
Cup and presidential elections, there were a high
number of messages in categories “politics” and “sports”.
In order to balance the proportion of messages in each
topic, a maximum limit of 450 messages was set for
each category.</p>
      <p>The filtered data resulted in a collection of 2444 messages,
of which 44320 distinct users retweeted one or more
messages at some moment. From the users lists of followers and
followees, it was possible to characterize a network,
registering the social connections between them. The source user,
Folha de Sa˜o Paulo, was removed from the network, so that
the interactions among its followers became the focus of the
analysis. Table 1 presents a more complete characterization
of the collected data and the network formed between users.
It is worth pointing, also, that most users do not participate
actively on the diffusions, implying in high diversity of users
participating in the processes, but low recurrence: during
the observation period, each user retweeted in average only
two messages from the source.
4.2</p>
    </sec>
    <sec id="sec-11">
      <title>Collective Coherence in Communities</title>
      <p>After running the community detection algorithm, 4278
communities were detected, many with 2 elements (2
connected users, isolated from the rest of the network), but 26
larger groups with 200 users or more have also been
identified.</p>
      <p>A first experiment involving communities consisted of
checking how coherent communities behaviors were, by comparing
the frequency at which specific messages were shared inside
communities and in the complete network. For each
community with a relevant number of users (200 or more), it
was computed the number of retweets for each Folha de Sa˜o
Paulo’s original tweet. A high coherence was verified in the
behaviors of individuals belonging to the same community.
Table 2 compares the sharing rates inside and outside the
four largest communities, showing the five most distinctive
cases where the community has singular sharing behavior,
differentiating from the global behavior.</p>
      <p>The cases shown on Table 2 indicate that there is a
certain level of coordination in the selection of shared messages
by each community. It is interesting to see how some
messages are much more emphasized inside a community, than
it was in the general network. Community 2, for example,
present higher sharing rates for its messages, compared to
the rest of the network. Equally interesting are cases where
the community seems to suppress the spread of a message,
as seen in community 8, where highly retweeted messages
are less emphasized by the community’s members.</p>
      <p>For comparison purposes, an attempt to break the
relationship between user connections and postings was made,
by randomly swapping the retweet pattern of different users
on the database, while keeping their social connections. Thus,
Community 1, members: 2621, total tweets: 12292
1st 2nd 3rd 4th 5th
Inside 2.5% 1.9% 2.1% 1.7% 1.4%
Outside 0.6% 0.3% 0.6% 0.2% 0.0%
under this new experiment, the communities remain defined
as they were originally, but their retweets have patterns from
users of different communities, in order to eliminate any
homophily related to retweeting behavior but keeping intact
other characteristics of our data (as global retweet counts
and correlations between sharings by individual users). The
same analysis made before is now performed on the
randomized dataset, resulting on Table 3. It becomes evident
that when homophily is suppressed, communities lose their
particular behavior and present sharing rates closer to the
whole network rates, indicating that differences in
retweeting patterns are not only artifact of communities’ finite sizes.
dex3 for their number of retweets within each community in
the real data with the same index for the data in the
randomized dataset. Our results indicate that the real communities
have stronger preferences for specific messages, with
satisfactory statistical significance when considering all
communities (Wilcoxon test, p = 1.23 ∗ 10−15, Gini index difference
of 1.5 ∗ 10−4) and even more pronounced when restricting
the comparison only to larger communities, with at least
50 individuals (Wilcoxon test, p = 2.13 ∗ 10−4, Gini index
difference of 8.26 ∗ 10−3).
4.3</p>
    </sec>
    <sec id="sec-12">
      <title>Topic Specialization</title>
      <p>A second experiment consisted in the analysis of how
topics are distributed among communities. For this, the six
standard categories were considered and the number of
messages of each category shared by each community was
computed. First, Figure 1a shows the general distribution of
messages per topic for the whole network. Then, Figures
1b1h show the distribution for seven distinctive communities
chosen among those with more than 200 members. The
figures demonstrate how communities’ topic distribution may
have different profiles compared to the rest of the network.</p>
      <p>From the graphs presented, communities 1 and 16
(Figures 1b and 1f) seems to have a stronger interest in political
issues, with community 1 showing an interest in daily life
news slightly above the average. Community 12 (Figure 1e)
also shows a higher interest in politics, but divide it with
a focus on sports news. Community 21 graph (Figure 1g)
shows a preference for daily life news, but not in a very
distinctively way. Community 10 (Figure 1d) does not have
a distribution very different from the whole network
(Figure 1a), being a representative of the average preferences.
Community 25 (Figure 1h), in turn, is remarkably different,
with over 50% of its retweets being about entertainment and
practically all the rest regarding sports news, almost
ignoring the other topics.
4.4</p>
    </sec>
    <sec id="sec-13">
      <title>Detecting Relevant Messages</title>
      <p>By applying the tf-idf normalization on the retweet counts
of the communities we identified the most characteristic tweets
for each community. Table 4 presents the top five messages
for the largest communities in the network. The messages
content were translated from Portuguese to English, with
some translation notes (in brackets), where necessary.</p>
      <p>It was possible to deepen the analysis of each
community profile beyond what would be possible by simply
looking to the distribution of general topics in each community.
Analysing each group of messages individually, it is possible
to notice specific and subjective categories. For example,
although communities 1 and 16 both have an emphasis in
politics, community 1 seems to be supportive to the
Brazilian government – focusing on good results of politics made
by the government and scandals of the opposition – while
16 appears more involved in topics related to (at the time)
election’s opposition candidates.</p>
      <p>A very interesting conclusion comes from the analysis of
the relevant messages from community 12, as we discover
that the tweets are not connected by the newspaper sections,
3The Gini index is a measure of how unevenly a value is
distributed among elements of a group – in our case, if the
Gini index is close to 1 then most of the retweets seen in a
community were associate with few messages, if its value is
near 0, then the distribution of retweets is closer to uniform.
(a) Whole network
(b) Community 1,
members: 1022,
retweets: 3160
(c) Community 9,
members: 272,
retweets: 636
(d) Community 10,
members: 283,
retweets: 861
(e) Community 12,
members: 316,
retweets: 820
(f) Community 16,
members: 215,
retweets: 527
(g) Community 21,
members: 268,
retweets: 745
(h) Community 25,
members: 217,
retweets: 254
but by subjects regarding the Brazilian state of Pernambuco
and its capital, Recife. All the messages presented were
related to events taking place in the state, involving both
politics matters (police strike) and sports (2014 FIFA World
Cup events). This analysis shows both the limitations of the
standard categorization of topics (the six classes defined by
the newspaper) and the potential of the tf-idf technique on
revealing new subjective connections among messages.</p>
      <p>Another strong topic specialization is noticed in
community 9, where all the top five messages are related to the
football team Palmeiras, giving evidence that the
community consists mainly of the team’s supporters. Community
25 is specialized in entertainment topics, showing an
apparent tendency to emphasize messages related to international
pop culture. Interestingly, although the topic distribution
in this community was predominantly on entertainment and
sports, the tf-idf normalization reveals the relative relevance
of messages in other categories, such as world and market
(market is the least shared topic among the six categories).
Sports tweets were not present among the top five messages,
which can probably mean that the sports messages shared
by the community followed the general distribution, not
revealing a distinctive behavior of the community.</p>
      <p>This kind of qualitative analysis of communities behavior
could be made with most communities detected in the
network, but are not presented here, due to space constraints.
Other examples of topic specialization present in
communities include: regional news (from diverse Brazilian states),
international politics, economy, football discussions (in
general and regarding specific teams) and corruption.</p>
    </sec>
    <sec id="sec-14">
      <title>DISCUSSION</title>
      <p>
        This research presents a computational framework for
general investigations on collective behavior. When applied to
a large dataset, the method presents new evidences of
homophily in Twitter’s network. Despite the existence of
homophily in Twitter was already found in different studies [
        <xref ref-type="bibr" rid="ref2 ref22 ref8">2,
8, 22</xref>
        ], homophily may be based on many different criteria,
as ethnic background, social class, mood, etc. The results
we presented highlight the relation between shared interests
and Twitter’s structure.
      </p>
      <p>The use of tf-idf jointly with community detection was
able to group and order messages according to their
relevance to a social community, enabling the characterization
of complex behavior profiles inside communities. The
presented method was able to reveal more nuanced classes of
contents, such as political positions, regional matters, fan
clubs, that were not covered in the original six categories,
defined by the newspaper’s staff with the specific purposes of
organization and classification. It is relevant to notice that,
beyond Twitter, the same technique can be applied to the
analysis of different sets and databases from social networks,
enabling similar studies for different services and contexts.</p>
      <p>The empirical results of this study show new concrete
evidences of how individual behavior and social connections are
closely related, as expected in theory. In fact, so much
information is present in social connections that it was even
possible to find groups of similar messages without any
knowledge of their content, using techniques that only consider
network properties and sharing behaviour. This is reflected
on the list of the most representative messages for each
community (Table 4), which are usually related to few subjects.
Accordingly, by aggregating information about which
communities interacted with some object (a tweet, for instance)
to other techniques (e.g., natural language processing), it
may be possible to gather new knowledge about such object
that are not evident in it (as social or geographic contexts).
In an extended perspective of this result, one can wonder
if it might be possible to infer significant elements about
the nature of a process happening on a social network even
without access to the content traveling through such
network, if there is information available about its structure
and dynamics.</p>
      <p>Ethical implications of the power of such methods for
detecting users specific interests, without the direct access to
their personal information, should be considered. If, on one
hand, this knowledge can be used in order to improve the
performance of useful systems, as machine learning
algorithms, on the other hand it may also incur in risks to
privacy and security. This discussion is not conducted in details
here, but should not be forgotten.</p>
      <p>The communities observed had elements of cohesion on
their general behaviors, emphasizing or even repressing the
spread of certain types of content. An interesting conflict
between individual autonomy and collective behavior seems
to be part of information diffusion processes that take place
on OSNs self-organized in communities. The community
specialization in topics of interest is also evidenced. On a
context of proliferation of many different subjects, the
limitation of the scope of themes discussed within a community
can be an efficient strategy for individuals to deal with
information overload.</p>
      <p>Further steps of the research include a deeper analysis of
a database including more messages from multiple sources,
using text-mining to define messages subjects and
comparing such classification to results obtained by tf-idf. In future
studies, the homophily of sharing behavior in online social
networks can be subject of a deeper analysis, developing new
methods to try to determine how much of it is due to (1)
preference of individuals to establish new social connections
to similar peers; (2) social influence; (3) indirect homophily,
which occurs due to the existence of homophily of another
trait (e.g., if two individuals usually access Twitter in the
same hours of the day, despite of their connections they will
be more likely to read the same news and, thus, retweet it).
An even deeper analysis can be made about social influence,
in order to be able to divide it into its reactive part – an
individual exhibits a sharing behavior in favor of some subject
because his/her community publishes more about such
subject, not because of an inner preference – and its cognitive
part – by observing his/her peers’ behaviors, an individual
shapes his/her preferences according to those practiced by
his/her peers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baraba</surname>
          </string-name>
          <article-title>´si and</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Albert</surname>
          </string-name>
          .
          <article-title>Emergence of scaling in random networks</article-title>
          .
          <source>Science</source>
          ,
          <volume>286</volume>
          (
          <issue>5439</issue>
          ):
          <fpage>509</fpage>
          -
          <lpage>512</lpage>
          , Oct.
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bollen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          , G. Ruan, and
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          .
          <article-title>Happiness is assortative in online social networks</article-title>
          .
          <source>Artificial life</source>
          ,
          <volume>17</volume>
          (
          <issue>3</issue>
          ):
          <fpage>237</fpage>
          -
          <lpage>251</lpage>
          , Jan.
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Borge-Holthoefer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ban</surname>
          </string-name>
          <article-title>˜os, S. Gonza´lez-Bailo´n, and</article-title>
          <string-name>
            <given-names>Y.</given-names>
            <surname>Moreno</surname>
          </string-name>
          .
          <article-title>Cascading behaviour in complex socio-technical networks</article-title>
          .
          <source>Journal of Complex Networks</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dillon</surname>
          </string-name>
          .
          <article-title>Introduction to modern information retrieval</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>19</volume>
          (
          <issue>6</issue>
          ):
          <fpage>402</fpage>
          -
          <lpage>403</lpage>
          ,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Goel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Watts</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Goldstein</surname>
          </string-name>
          .
          <article-title>The structure of online diffusion networks</article-title>
          .
          <source>In Proceedings of the 13th ACM Conference on Electronic Commerce (EC '12)</source>
          , volume
          <volume>1</volume>
          ,
          <string-name>
            <surname>page</surname>
            <given-names>623</given-names>
          </string-name>
          , New York, New York, USA,
          <year>2012</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Halu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baronchelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Bianconi</surname>
          </string-name>
          .
          <article-title>Connect and win: The role of social networks in political elections</article-title>
          .
          <source>EPL (Europhysics Letters)</source>
          ,
          <volume>102</volume>
          (
          <issue>1</issue>
          ):
          <fpage>16002</fpage>
          ,
          <string-name>
            <surname>Apr</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kurka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Godoy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F. Von</given-names>
            <surname>Zuben</surname>
          </string-name>
          .
          <article-title>Online social network analysis: A survey of research applications in computer science</article-title>
          .
          <source>arXiv:0707.3168 [cs.SI]</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kwak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Moon</surname>
          </string-name>
          .
          <article-title>What is Twitter, a social network or a news media</article-title>
          ?
          <source>In Proceedings of the 19th International Conference on World Wide Web (WWW '10)</source>
          , page 591, New York, New York, USA,
          <year>2010</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lansdall-Welfare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lampos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Cristianini</surname>
          </string-name>
          .
          <article-title>Effects of the recession on public mood in the UK</article-title>
          .
          <source>In Proceedings of the 21st International Conference Companion on World Wide Web (WWW '12)</source>
          , pages
          <fpage>1221</fpage>
          -
          <lpage>1226</lpage>
          . ACM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Adamic</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Huberman</surname>
          </string-name>
          .
          <article-title>The dynamics of viral marketing</article-title>
          .
          <source>ACM Transactions on the Web</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):5,
          <string-name>
            <surname>May</surname>
          </string-name>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>M. McPherson</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Smith-Lovin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <article-title>Birds of a feather: Homophily in social networks</article-title>
          .
          <source>Annual Review of Sociology</source>
          ,
          <volume>27</volume>
          (
          <issue>1</issue>
          ):
          <fpage>415</fpage>
          -
          <lpage>444</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitchell. Complexity</surname>
          </string-name>
          :
          <string-name>
            <given-names>A Guided</given-names>
            <surname>Tour</surname>
          </string-name>
          . Oxford University Press,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Newman</surname>
          </string-name>
          .
          <article-title>Modularity and community structure in networks</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>103</volume>
          (
          <issue>23</issue>
          ):
          <fpage>8577</fpage>
          -
          <lpage>8582</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J.</given-names>
            <surname>Ugander</surname>
          </string-name>
          .
          <article-title>On the interplay between social and topical structure</article-title>
          .
          <source>arXiv:1112.1115 [cs.SI]</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosvall</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bergstrom</surname>
          </string-name>
          .
          <article-title>Maps of random walks on complex networks reveal community structure</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>105</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1118</fpage>
          -
          <lpage>1123</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Salath</surname>
          </string-name>
          ´e,
          <string-name>
            <given-names>D. Q.</given-names>
            <surname>Vu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Hunter</surname>
          </string-name>
          .
          <article-title>The dynamics of health behavior sentiments on a large online social network</article-title>
          .
          <source>EPJ Data Science</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):4,
          <string-name>
            <surname>Apr</surname>
          </string-name>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarcevic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Palen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Starbird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bagdouri</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Anderson</surname>
          </string-name>
          .
          <article-title>“beacons of hope” in decentralized coordination</article-title>
          .
          <source>In Proceedings of the ACM 2012 conference on Computer Supported Cooperative Work (CSCW '12)</source>
          , New York, New York, USA,
          <year>2012</year>
          . ACM, ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Stieglitz</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Dang-Xuan</surname>
          </string-name>
          .
          <article-title>Political communication and influence through microblogging - an empirical analysis of sentiment in Twitter messages and retweet behavior</article-title>
          .
          <source>In 2012 45th Hawaii International Conference on System Science (HICSS)</source>
          , pages
          <fpage>3500</fpage>
          -
          <lpage>3509</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Tumasjan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sprenger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sandner</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Welpe.</surname>
          </string-name>
          <article-title>Predicting elections with Twitter: What 140 characters reveal about political sentiment</article-title>
          .
          <source>In International AAAI Conference on Weblogs and Social Media</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Watts</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Dodds</surname>
          </string-name>
          .
          <article-title>Influentials, networks, and public opinion formation</article-title>
          .
          <source>Journal of Consumer Research</source>
          ,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <fpage>441</fpage>
          -
          <lpage>458</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. a.</given-names>
            <surname>Huberman</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>a. Adamic, and</article-title>
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Tyler</surname>
          </string-name>
          .
          <article-title>Information flow in social groups</article-title>
          .
          <source>Physica A: Statistical Mechanics and its Applications</source>
          ,
          <volume>337</volume>
          (
          <issue>1-2</issue>
          ):
          <fpage>327</fpage>
          -
          <lpage>335</lpage>
          ,
          <year>June 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Hofman</surname>
          </string-name>
          , W. a. Mason, and
          <string-name>
            <given-names>D. J.</given-names>
            <surname>Watts</surname>
          </string-name>
          .
          <article-title>Who says what to whom on twitter</article-title>
          .
          <source>In Proceedings of the 20th International Conference on World Wide Web (WWW '11)</source>
          , page 705, New York, New York, USA,
          <year>2011</year>
          . ACM Press.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>