<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recommending #-Tags in Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eva Zangerle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wolfgang Gassler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gunther Specht</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Databases and Information Systems Institute of Computer Science University of Innsbruck</institution>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Twitter, currently the most popular microblogging tool available, is used to publish more than 140,000,000 messages a day. Many users use hashtags to categorize their tweets. However, hashtags are not restricted in any way in terms of usage or syntax which leads to a very heterogeneous set of hashtags occurring in the Twitter universe and therefore, decreases the search capabilities. In this paper, we present an approach for the recommendation of highly appropriate hashtags to the user during the creation process. The recommendations aim at encouraging the user to (i) use hastags at all, (ii) use more appropriate hashtags and (iii) avoid the usage of synonymous hashtags. Therefore the vocabulary of hashtags becomes more homogenous regarding both syntax and semantics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        let the world know what they're up to or simply to share some information they
consider useful. The probably most important feature of Twitter is the retweet
functionality. It enables users to further broadcast tweets they consider worth
spreading within the Twitter network. Mostly, the retweeted message remains
unchanged. A retweeted message contains "RT: @originaluser\ followed by
the original message. This retweeting, which was also heavily analysed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
can spread an important message all over the world within minutes. Due to the
ever increasing amount of Twitter messages and the resulting chaos within the
Twittersphere, the microblogging community started to use so-called hashtags
as a means for the manual categorization of tweets. The categorization can be
used for either searching for certain topics based on the used hashtags or to be
able to follow certain conversations about a certain topic on Twitter. The only
requirement for hashtags is to start with a hash symbol #. Besides this fact,
hashtags do not have to conform to any rules or regulations and can be seen as
typcial tags as used in common Web 2.0 applications like e.g. Blogs. Hashtags
may appear at any arbitrary position within the message and may consist of
any arbitrary combination of characters. This makes them easy to use, but at
the same time leads to a signi cant lack of structure and uniformity. During
our research, we crawled a data set and analyzed it. In the process we found
that that users utilized very di erent popular hashtags for their tweets about
the same topic. For example, the Tour de France (a world-famous bicycle race
in France) was very popular. Tweets about this topic contain di erent hashtags,
such as #tdf, #tourdefrance, #cycling or #procycling. Twitter o ers its
users a search engine which is able to search for keywords, but also for hashtags.
Therefore, when searching for discussions about the Tour de France by using the
search hashtag #tourdefrance, the user might not be able to retrieve all tweets
containing information about the Tour de France. This is due to the fact that
other users used the hashtag #tdf, which the user did not specify in the search
query. Certainly, tweets containing the hashtag #tdf would also have been a
perfect match for the user's query. However, due to the heterogeneous hashtag
vocabulary used by the active Twitter community, many synonymous hashtags
are used for describing the same semantic information.
      </p>
      <p>In this paper we introduce an approach for the recommendation of hashtags. Our
approach computes recommendations based on an analysis of existing tweets by
other users and recommends suitable hashtags for the currently entered message
to the user. This recommendation mechanism aims at encouraging the user to
make use of hashtags and creating a more homogeneous hashtag vocabulary in
order to enhance the quality of search result. Additionally, we present general
statistics about the use of hashtags within Twitter and an evaluation of our
approach.</p>
      <p>The remainder of this paper is organized as follows. Section 2 describes the
basic concepts of Twitter and hashtags. Section 3 is concerned with the process of
hashtag recommendations. Section 4 contains the experiments and evaluations
of the presented approach. Subsequently, Section 5 describes important related
work and Section 6 concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>Hashtags</title>
      <p>Hashtagging is a simple and convenient way for users to categorize their own
tweets. Such a hashtag within a tweet can simply be speci ed by adding a hash
- '#' - followed by the tag itself. One tweet may also contain multiple hashtags,
like in the following example tweet: "Don't forget! Only 7 days till the
#SASWeb submission deadline #umap2011 http://bit.ly/dKgS82.
#recsys #um #adaptivity #web3.0 #ontologies\ which was posted by the
SASWeb workshop (@sasWeb2011).</p>
      <p>
        The most popular hashtags are either related to long-term popular topics or
to current events or topics, e.g. the hashtag #tdf was extensively used during the
crawling period as the Tour de France was taking place during this time. Typical
long-term topics are e.g. #Apple or #Obama which are featured in thousands of
messages a day [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
2.1
      </p>
      <sec id="sec-2-1">
        <title>Data Set and Hashtag Analysis</title>
        <p>In order to be able to analyse the hashtagging behaviour of Twitter users and to
build up a database which forms the basis for all recommendation computations,
we had to crawl tweets. Overall, we collected about 16,000,000 tweets from July
2010 until February 2011 via the Twitter Application Programming Interface.</p>
        <p>In order to retrieve a diverse and highly representative data set to base our
evaluations and analysis on, we decided to use Twitter's API2. The basis for our
search queries was an English dictionary containing more than 32,000 words.
We iterated over the words contained in the dictionary and used them as search
keywords for the Twitter Search API. All search results were stored whereby only
tweets containing hashtags were used for further analysis. Another approach was
to retrieve the public timeline, which basically consists of the ten latest tweets.
The timeline is displayed on the Twitter website and is also available via the
API. However, these tweets are only updated once a minute. Therefore, only 600
tweets could be retrieved per hour and considering the fact that only 20% of all
tweets contain hashtags, this approach was not feasible for crawling a su ciently
large dataset.</p>
        <p>After having crawled the data, we had to perform multiple preprocessing
steps. This included removing all non-english messages (based on Twitter's
language classi cation mentioned in the metadata of every tweet) and all messages
not containing hashtags at all. Furthermore, all messages were transformed to
lower-case. Table 1 contains an overview about the crawled data set and its
characteristics. Out of the crawled tweets, more than 3 million tweets contained
at least one hashtag, which marks 20% of all crawled tweets. The hashtags
ltered from all tweets were further analysed in regards to their usage and
popularity. Figure 1 displays the long tail distribution of hashtags and their
usage. The fact that stands out about this distribution is that 86% of all
hashtags within the data set were used within less than ve tweets. On the other
2 http://search.twitter.com/search
hand, the most popular hashtags within the data set (#jobs, #nowplaying,
#zodiacfacts, #news and #fb) were used in 8% of all messages containing
hashtags. Another interesting fact is the distribution of the number of
hashtags used per tweet which can be seen in Figure 2. We expected the number of
hashtags per message to be decreasing steadily. This is mostly the case for
messages contains less than 15 hashtags. However, the sudden amplitude at 17
hashtags per message is somewhat surprising. We therefore examined these messages
and discovered that these were spam tweets which only contained hashtags and
a URL, like e.g. "RT @Bhupesh tweet: #Quad #loop-http://bit.ly/ciHX2U
#retweet #India #Jobs #World #news #canada #ad #win #USA #tdf #oea
#hacking #icantstop #sdcc #game\. Such tweets typically also feature a high
retweet-rate by using a spam network consisting of many Twitter users created
for spam purposes.</p>
        <p>Characteristic</p>
        <sec id="sec-2-1-1">
          <title>Crawled messages total</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>Messages containg at least one hashtag</title>
        </sec>
        <sec id="sec-2-1-3">
          <title>Messages containing no hashtags</title>
        </sec>
        <sec id="sec-2-1-4">
          <title>Retweets</title>
        </sec>
        <sec id="sec-2-1-5">
          <title>Direct messages</title>
        </sec>
        <sec id="sec-2-1-6">
          <title>Hashtags usages total</title>
        </sec>
        <sec id="sec-2-1-7">
          <title>Hashtags distinct</title>
        </sec>
        <sec id="sec-2-1-8">
          <title>Average number of hashtags per message</title>
        </sec>
        <sec id="sec-2-1-9">
          <title>Maximum number of hashtags per message</title>
        </sec>
        <sec id="sec-2-1-10">
          <title>Hashtags occurring &lt; 5 times in total</title>
        </sec>
        <sec id="sec-2-1-11">
          <title>Hashtags occurring &lt; 3 times in total</title>
          <p>Hashtags occuring only once
16,034,195
3,209,281
12,824,914
2,556,617
3,073,948
5,097,545
510,170
1.5884</p>
          <p>23
437,266
328,348
384,187
Value Percentage
100%
20%
80%
16%
19%
{
{
{
{
{
{
{
The aim of the approach presented in this paper is to nd a set of hashtags
suitable for any tweet the user enters. These hashtags are then recommended
to the user during the creation process of the new tweet. Recommendations are
basically be computed by performing the following steps:
1. nding the most similar messages in the crawled data set for the tweet just
entered by the user
)e 10000
l
-scag
o
l
(s 1000
e
c
n
e
rrc
u
c 100
O
fr
o
e
ubm 10
N</p>
          <p>1
1e+07
1e+06
100000
e 10000
sg
sseaM 1000
100
10
1 0</p>
          <p>Hashtags ids (ordered by popularity)
2. retrieving the set of hashtags used within these most similar messages
3. ranking the computed set of hashtag recommendation candidates
These steps for the computation of hashtag recommendations are discussed in
the following sections.
3.1</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Similarity of Tweets</title>
        <p>In order to be able to determine similar tweets, a similarity measure for the
comparison of two at most 140 character long messages has to be introduced.
This metric is used to rank the results gathered from searching similar tweets
within the crawled data set. These similar messages are subsequently considered
to contain valuable hashtag recommendation candidates. A straightforward
solution is to use the term frequency - inverse document frequency measure for
the comparison of tweets. In order to be able to use tf/idf for the computation
of the similarity of tweets, the formula stated in Equation (1) is used.
tf idft;d = tft;d idft
(1)
idft = log</p>
        <p>jDj
jfd : t 2 dgj
(3)</p>
        <p>In the case of searching a set of tweets, the set D of documents which have
to be searched is the set of tweets in the system. The term frequency basically is
the number of occurrences of a term t within a given document d (tweet). The
inverse document frequency (idf) constitutes the importance of a term t within
the whole set of documents which are searched. This is computed by taking the
number of all documents (jDj) within the index and dividing it by the number
of documents which contain the searched term (jfd : t 2 dgj). The computation
of the tf/idf measure for a given search query (in our case the tweet inserted by
the user), is subsequently accomplished by computing the sum of all tf/idf of all
terms t occurring within the search query d: Pt in d tf idf (t). Futhermore, the
nal score is increased if more of the terms of the query are matched. The nal
set of similar tweets (those obtaining the highest tf/idf-based score ratings) is
restricted to a set of tweets having a score above a certain treshold corresponding
to the total number of results and the speci ed limit of total results.
3.2</p>
      </sec>
      <sec id="sec-2-3">
        <title>Ranking</title>
        <p>After having obtained the set of the most similar messages to the tweet the user
just entered, the hashtags are extracted from these tweets. These hashtags are
referred to as hashtag recommendation candidates throughout the remainder of
this paper. The ranking of these hashtag recommendation candidates is crucial
for the success of recommendations. This is due to the fact that both the
cognition of the user and the space available for displaying the recommendations is
limited. In most cases a set of 5-10 recommendations is most appropriate which
also correspond to the capacity of short-term memory (Miller, 1956). Therefore
the top-k recommendations are shown to the user, where k denotes the size of the
set of recommended hashtags presented to the user. This restricted set is based
on the set of all hashtags which were extracted from the most similar messages
to the newly created tweet. To present the most suitable top-k hashtags to the
user, the recommendation candidates have to be ranked. For our approach, we
evaluated three ranking methods, which can be summarized as follows:
{ OverallPopularityRank: This ranking approach is based on the popularity of
the hashtag recommendation candidates. It basically considers the number of
occurrences of the respective hashtag within our data set. The more popular
a hashtag is overall, the higher the resulting rank of the hashtag.
{ RecommendationPopularityRank: This ranking method basically counts the
occurrences of each hashtag within the set of recommendation candidates.
The higher the number of occurrences, the more (similar) messages contain
this hashtag. Therefore, it is likely that the hashtag is suitable for the tweet
the user just entered.
{ SimilarityRank: This ranking method is based on the similarity value
between the tweet entered by the user and the tweet which provides a hashtag
recommendation candidate. The more similar the messages are, the more
likely it is that the hashtags contained in this similar message are suitable
for the tweet entered by the user. In the case that multiple tweets contain
the hashtag which has to be ranked, the similarity of the most similar tweet
is used. As a metric for the similarity of tweets, we used tf/idf as described
in 3.1.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <p>A recommendation engine prototype implementing this approach has been
developed based on Apache's Lucene3 fulltext index. We used the fulltext index to
store the crawled tweets which enabled us to nd the most similar messages by
using Lucene's Search Index.
4.1</p>
      <sec id="sec-3-1">
        <title>Test Setup</title>
        <p>The evaluation was done on a CentOS release 5.1 machine with 8 GB of RAM.
The evaluation of the hashtag recommendation approaches was conducted by
performing a leave-one-out test. This test was based on the data set described in
Section 2.1. Based on the crawled data set, we built a fulltext index comprising all
3.2 mil. cleaned messages without hashtags of this data set. From this index, we
randomly chose 10,000 messages with less than six hashtags for each test run. For
each of these messages, the contained hashtags were removed from the message
and the resulting string was used as the input tweet for the recommendation
engine. Naturally, the currently used tweet was removed from the Lucene Index
and was not considered for the computation of recommendation candidates.
Additionally, no retweets were used as test input tweets as search for similar
messages would return an identical retweeted message which would obviously
distort the evaluation results.</p>
        <p>Based on the hashtag recommendations computed by the recommendation
engine, we evaluted the three ranking methods described in 3.2. This was done
by computing the precision and recall values of the top-k recommendations with
k = 1, k = 2,..., k = 10 as described in the next section.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Precision and Recall</title>
        <p>For the evaluation of the quality of the computed recommendations, we chose to
use the precision and recall values of the recommendations. These metrics are
de ned as follows:
precision (Hrec) = jHrec \ Horigj
jHrecj
(4)
3 http://lucene.apache.org/
where Horiginal is the set of original hashtags which were removed from
the original tweet and Hrecommended is the set of top-k recommendations. We
performed ten test runs for each ranking method with k = 1, k = 2,..., k = 10.
Each test run computed the respective average recall and precision value of
10,000 test tweets. Thus, the evaluation is based on the computation of 100,000
top-k recommendation sets for each ranking method.
(5)
4.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Results</title>
        <p>The experiments conducted showed that the approach is feasible of
recommending suitable hashtags. The recall values for the top-k recommended hashtags can
be seen in Figure 3. In this gure, the recall values for k (the number of
recommended hashtags) being between 1 and 10 has been evaluated for the three
considered ranking methods. This Figure shows that ranking based on the
overall popularity of the hashtag (OverallPopularityRank) and also based on the
popularity of the hashtag within the hashtag recommendation candidates
(RecommendationPopularityRank) do not perform well. In contrast, SimilarityRank
(ranking based on the similarity of the original tweet and the tweet
containing the recommendation candidate) is able to perform signi cantly better. This
ranking method leads to promising recall values which are well above the 40%
mark for k &gt; 2.</p>
        <p>The precision values for the computed recommendation sets decrease with
an increasing k. This is due the fact, that we only use test tweets with at most
5 hashtags per message. Therefore even a set of 10 recommendations featuring
a recall value of 100% only results in a precision of 50% as ve of the ten
recommended hashtags are not applicable as the original message only features
ve hashtags.</p>
        <p>Overall, the evaluations showed that our approach is suitable for the
recommendation of hashtags. Another fact which can be derived from the evaluations
is that our approach shows the best performance when restricting the set of
recommended hashtags to k = 5, as the recall value does not improve much with
additional recommendations and the precision value is still reasonable.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Related Work</title>
      <p>The recommendation of Twitter hashtags can bene t from various other elds of
research. These areas are (i) tagging of online resources, (ii) traditional
recommender systems, (iii) social network analysis and (iv) Twitter analysis. However,
to the best of our knowledge, there is no other approach aiming at recommending
hashtags to Twitter users.</p>
      <p>
        The recommendation of tags of online resources like images, bookmarks or
bibliographic entries is directly related to our approach. Such approaches can
be based on the co-occurrence of tags, like e.g. in [
        <xref ref-type="bibr" rid="ref14 ref20">14, 20</xref>
        ]). The notion of
cooccurrence of tags describes the fact that two tags are used to tag the same
photo. Therefore, only partly tagged photos can be subject to tag
recommendations. Based on these relatively simple approaches, the paper by Rae et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
proposes a method for Flickr tag recommendations which takes di erent
contexts into account. Rae distinguishes four di erent contexts for the computation
of recommendations: (i) the user's previously used tags, (ii) the tags of the user's
contacts, (iii) the tags of the users which are members of the same groups as
the user and (iv) the collectively most used tags by the whole community. A
similar approach has also been facilitated by Garg and Weber in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Another
example for recommendations of tags is based on the BibSonomy platform which
basically allows its users to tag bibliographic entries [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This approach extracts
tags which might be suitable for the entry from the title of the entry, the tags
previously used for the entry and tags previously used by the current user.
Based on these resources, the authors propose di erent approaches for merging
these sets of tags. The resulting set is subsequently recommended to the user.
Jaschke et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] propose a collaborative ltering approach for the
computation of tag recommendations. This computation is based on a graph consisting
of the users, their tags and the tagged resources. After having constructed this
graph, a PageRank-like ranking algorthm (called FolkRank) is applied.
Furthermore, [
        <xref ref-type="bibr" rid="ref15 ref2">2, 15</xref>
        ] are mainly concerned with the motivation of users to tag resources.
John Hannon et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] developed the Twittomender system which facilitates an
approach for the recommendation of followees. This is done by creating
proles of users and applying a collaborative ltering approach to these pro les.
The Twittomender system also provides search functionality (based on
arbitrary keywords) which returns pro le information about the found users like e.g.
the latest popular keywords used by the speci c user or his latest tweet.
Another approach directly connected to Twitter and recommendations is
described by Phelan et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In this approach, Twitter is used for the
recommendation of news articles. In particular, Twitter is used to rank the news stories
originating from various RSS feeds based on the user's tweets, the user's friends
tweets or the public most recent tweets. Also, Jilian Chen et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] focused on
recommendations based on tweets. In this case, interesting URLs are
recommended to the user. Romero et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] analyzed how hashtags spread within the
Twitter Universe. The hashtags were analyzed with regards to how a hashtag
might be used by a user who is exposed to this hashtag by his followers and
followees. The authors categorized the top-500 hashtags used within their data
set and found that the adoption of hashtags is dependent on the category of
the hashtags. E.g. multiple exposure to a hashtag for political or sports topics
lead to the adoption of the hashtag with a higher probability than in any other
hashtag category.
      </p>
      <p>
        Kwak et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] did a thorough analysis of the Twitter universe focusing on
information di usion within the network. Further analysis of Twitter messages
are also contained in [
        <xref ref-type="bibr" rid="ref11 ref12 ref21 ref3">3,11,12,21</xref>
        ]. There have been numerous papers throughout
the last years addressing the social aspects of Twitter and social online networks
in general. Huberman et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] found that the Twitter network basically consists
of two networks: one dense network consisting of all followers and followees and
one sparse network consisting of the actual friends of users. Huberman de nes a
friend of a user as another Twitter user with whom the user exchanged at least
two directed messages. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] contains an analysis of the retweet messages and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is
concerned with how Twitter might be suitable for collaboration by exchanging
direct messages.
      </p>
      <p>
        As for the recommender system facilitated in our approach, many publications
are focused around collaborative ltering. The papers by Resnick [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and
Adomavicius [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provide a very good overview about the eld of collaborative
ltering.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we presented an approach for the recommendation of hashtags
within the Twitter microblogging application. The presented algorithm is based
on the analysis of similar tweets and the hashtags contained in these tweets. Our
evolutions were based on a self-crawled data set consisting of 12 million tweets.
The preliminary evaluations showed promising results as the recall values of
the recommendations are about 45-50%. Future work will include integrating
the social graph of Twitter users for the recommendation. Furthermore, the
ranking of hashtag recommendation candidates is also subject to further research
and improvements. The enhancement of the recommendations of synonymous
hashtags based on a semantic analysis for the exclusion of synonymous hashtags
and their recommendation is also part of future work.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>G.</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Tuzhilin</surname>
          </string-name>
          .
          <article-title>Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions</article-title>
          .
          <source>IEEE transactions on knowledge and data engineering</source>
          ,
          <volume>17</volume>
          (
          <issue>6</issue>
          ):
          <volume>734</volume>
          {
          <fpage>749</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ames</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Naaman</surname>
          </string-name>
          .
          <article-title>Why we tag: motivations for annotation in mobile and online media</article-title>
          .
          <source>In Proceedings of the SIGCHI conference on Human factors in computing systems, CHI '07</source>
          , pages
          <fpage>971</fpage>
          {
          <fpage>980</fpage>
          , New York, NY, USA,
          <year>2007</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Asur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Huberman</surname>
          </string-name>
          , G. Szabo, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Trends in Social Media: Persistence and Decay</article-title>
          .
          <source>Arxiv preprint arXiv:1102.1402</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>D.</given-names>
            <surname>Boyd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Golder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Lotan</surname>
          </string-name>
          .
          <article-title>Tweet, tweet, retweet: Conversational aspects of retweeting on twitter</article-title>
          .
          <source>In hicss, pages</source>
          <volume>1</volume>
          {
          <fpage>10</fpage>
          . IEEE Computer Society,
          <year>1899</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nairn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Chi</surname>
          </string-name>
          .
          <article-title>Short and tweet: experiments on recommending content from information streams</article-title>
          .
          <source>In Proceedings of the 28th international conference on Human factors in computing systems</source>
          , pages
          <volume>1185</volume>
          {
          <fpage>1194</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>N.</given-names>
            <surname>Garg</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.</given-names>
            <surname>Weber</surname>
          </string-name>
          .
          <article-title>Personalized, interactive tag recommendation for ickr</article-title>
          .
          <source>In Proceedings of the 2008 ACM conference on Recommender systems, RecSys '08</source>
          , pages
          <fpage>67</fpage>
          {
          <fpage>74</fpage>
          , New York, NY, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.</given-names>
            <surname>Hannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bennett</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          .
          <article-title>Recommending twitter users to follow using content and collaborative ltering approaches</article-title>
          .
          <source>In RecSys '10: Proceedings of the fourth ACM conference on Recommender systems</source>
          , pages
          <volume>199</volume>
          {
          <fpage>206</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>C.</given-names>
            <surname>Honeycutt</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Herring</surname>
          </string-name>
          .
          <article-title>Beyond microblogging: Conversation and collaboration via twitter</article-title>
          .
          <source>In HICSS</source>
          , pages
          <volume>1</volume>
          {
          <fpage>10</fpage>
          . IEEE Computer Society,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>B.</given-names>
            <surname>Huberman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Romero</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Social networks that matter: Twitter under the microscope</article-title>
          .
          <source>First Monday</source>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ):
          <fpage>8</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R.</given-names>
            <surname>Jaeschke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Marinho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hotho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt-Thieme</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Stumme</surname>
          </string-name>
          .
          <article-title>Tag Recommendations in Folksonomies</article-title>
          . In J. Kok,
          <string-name>
            <given-names>J.</given-names>
            <surname>Koronacki</surname>
          </string-name>
          , R. Lopez de Mantaras,
          <string-name>
            <given-names>S.</given-names>
            <surname>Matwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mladenic</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A</surname>
          </string-name>
          . Skowron, editors,
          <source>Knowledge Discovery in Databases: PKDD</source>
          <year>2007</year>
          , volume
          <volume>4702</volume>
          of Lecture Notes in Computer Science, pages
          <volume>506</volume>
          {
          <fpage>514</fpage>
          . Springer Berlin / Heidelberg,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>A.</given-names>
            <surname>Java</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Finin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Tseng</surname>
          </string-name>
          .
          <article-title>Why we twitter: understanding microblogging usage and communities</article-title>
          .
          <source>In Proceedings of the 9th WebKDD and 1st SNA-KDD 2007 workshop on Web mining and social network analysis</source>
          , pages
          <volume>56</volume>
          {
          <fpage>65</fpage>
          . ACM,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>B.</given-names>
            <surname>Krishnamurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gill</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Arlitt</surname>
          </string-name>
          .
          <article-title>A few chirps about twitter</article-title>
          .
          <source>In Proceedings of the rst workshop on Online social networks</source>
          , pages
          <volume>19</volume>
          {
          <fpage>24</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. H.
          <string-name>
            <surname>Kwak</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Park</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Moon</surname>
          </string-name>
          .
          <article-title>What is Twitter, a social network or a news media</article-title>
          ?
          <source>In Proceedings of the 19th international conference on World wide web</source>
          , pages
          <volume>591</volume>
          {
          <fpage>600</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>M.</given-names>
            <surname>Lipczak</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Milios</surname>
          </string-name>
          .
          <article-title>Learning in e cient tag recommendation</article-title>
          .
          <source>In Proceedings of the fourth ACM conference on Recommender systems, RecSys '10</source>
          , pages
          <fpage>167</fpage>
          {
          <fpage>174</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>C. Marlow</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Davis</surname>
          </string-name>
          . HT06, tagging paper, taxonomy, Flickr, academic article, to read.
          <source>In Proceedings of the seventeenth conference on Hypertext and hypermedia, page 40. ACM</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>O.</given-names>
            <surname>Phelan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McCarthy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          .
          <article-title>Using twitter to recommend real-time topical news</article-title>
          .
          <source>In Proceedings of the third ACM conference on Recommender systems</source>
          , pages
          <volume>385</volume>
          {
          <fpage>388</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>A.</given-names>
            <surname>Rae</surname>
          </string-name>
          , B. Sigurbjornsson, and R. van Zwol.
          <article-title>Improving tag recommendation using social networks</article-title>
          .
          <source>In Adaptivity, Personalization and Fusion of Heterogeneous Information, RIAO '10</source>
          , pages
          <fpage>92</fpage>
          {
          <fpage>99</fpage>
          , Paris, France, France,
          <year>2010</year>
          . Le Centre de Hautes Etudes Internationales d'Informatique Documentaire.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>P.</given-names>
            <surname>Resnick</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Varian</surname>
          </string-name>
          .
          <article-title>Recommender systems</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>40</volume>
          (
          <issue>3</issue>
          ):
          <fpage>58</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>D. M. Romero</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Meeder</surname>
            , and
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Kleinberg</surname>
          </string-name>
          .
          <article-title>Di erences in the mechanics of information di usion across topics: idioms, political hashtags, and complex contagion on twitter</article-title>
          . In S. Srinivasan,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ramamritham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Ravindra</surname>
          </string-name>
          , E. Bertino, and R. Kumar, editors,
          <source>WWW</source>
          , pages
          <volume>695</volume>
          {
          <fpage>704</fpage>
          . ACM,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. B.
          <article-title>Sigurbjornsson and</article-title>
          <string-name>
            <given-names>R. Van</given-names>
            <surname>Zwol</surname>
          </string-name>
          .
          <article-title>Flickr tag recommendation based on collective knowledge</article-title>
          .
          <source>In Proceeding of the 17th international conference on World Wide Web</source>
          , pages
          <volume>327</volume>
          {
          <fpage>336</fpage>
          . ACM,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>S.</given-names>
            <surname>Ye</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Measuring Message Propagation and Social In uence on Twitter. com</article-title>
          . In Social Informatics: Second International Conference,
          <year>Socinfo 2010</year>
          , Laxenburg, Austria,
          <source>October 27-29</source>
          ,
          <year>2010</year>
          , Proceedings, page 216. Springer-Verlag New York Inc,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>