<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>MSM</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>What makes a tweet relevant for a topic?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ke Tao</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabian Abel</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudia Hauff</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Geert-Jan Houben Web Information Systems</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>TU Delft PO Box</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>GA Delft</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>the Netherlands</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <volume>2</volume>
      <fpage>49</fpage>
      <lpage>56</lpage>
      <abstract>
        <p>Users who rely on microblogging search (MS) engines to find relevant microposts for their queries usually follow their interests and rationale when deciding whether a retrieved post is of interest to them or not. While today's MS engines commonly rely on keyword-based retrieval strategies, we investigate if there exist additional micropost characteristics that are more predictive of a post's relevance and interestingness than its keyword-based similarity with the query. In this paper, we experiment with a corpus of Twitter messages and investigate sixteen features along two dimensions: topicdependent and topic-independent features. Our in-depth analysis compares the importance of the different types of features and reveals that semantic features and therefore an understanding of the semantic meaning of the tweets plays a major role in determining the relevance of a tweet with respect to a query. We evaluate our findings in a relevance classification experiment and show that by combining different features, we can achieve a precision and recall of more than 35% and 45% respectively.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Microblogging services such as Twitter1 or Sina Weibo2
have become a valuable source of information particularly
for exploring, monitoring and discussing news-related
information [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Searching for relevant information in such
services is challenging as the number of posts published per
day can exceed several hundred millions3.
      </p>
      <p>
        Moreover, users who search for microposts about a
certain topic typically perform a keyword search. Teevan et
al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] found that keyword queries on Twitter are
significantly shorter than those issued for Web search: on Twitter
people typically use 1.64 words (or 12.0 characters) to search
while on the Web they use, on average, 3.08 words (or 18.8
      </p>
      <sec id="sec-1-1">
        <title>1http://twitter.com/</title>
      </sec>
      <sec id="sec-1-2">
        <title>2http://www.weibo.com/</title>
      </sec>
      <sec id="sec-1-3">
        <title>3http://blog.twitter.com/2011/06/</title>
        <p>200-million-tweets-per-day.html
Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are
nCootpmyardieghotr dcistr2i0b1ut2edhefoldr pbryofit aour tchoomrm(se)r/coiawlnaedrv(asn)t.age and that copies
bPeuarbtlhisishendotiacse apnadr thoeffuthlleci#tatMioSnMon2t0h1e2firsWtpoargkes.hToopcpoproycoetehderiwngisse,, to
raevpauibllaibshle,tonploisnteonasseCrvEerUsRor Vtoorle-d8i3s8tr,ibautt:e htottlips:ts/,/rceequrir-ewssp.roiorrgs/pVeoclific-838
p#erMmSisMsio2n01an2d,/Aorpariflee1.6, 2012, Lyon, France.</p>
        <p>Copyright 200X ACM X-XXXXX-XX-X/XX/XX ...$5.00.
characters). This can be explained by the length of
Twitter messages which is limited to 140 characters so that long
queries easily become too restrictive. Short queries on the
other hand may result in a large (or too large) number of
matching microposts.</p>
        <p>For these reasons, building search algorithms that are
capable of identifying interesting and relevant microposts for
a given topic is a non-trivial and crucial research challenge.
In order to take a first step towards solving this challenge,
in this paper, we present an analysis of the following
question: is a keyword-based retrieval strategy sufficient or can
we identify features that are more predictive of a tweet’s
relevance and interestingness? To investigate this question,
we took advantage of last year’s TREC4 2011 Microblog
Track5, where for the first time an openly accessible search
&amp; retrieval Twitter data set with about 16 million tweets
was published.</p>
        <p>In the context of TREC, the ad-hoc search task on
Twitter is defined as follows: given a topic (identified by a title)
and a point in time pt, retrieve all interesting and relevant
microposts from the corpus that were posted no later than
pt. A subset of the tweets that were retrieved by the research
groups participating in the benchmark were then judged by
human assessors as either relevant to the topic or as
nonrelevant. For example, “Obama birth certificate” is one of
the topics that is part of the TREC corpus. Given the
temporal context, one can infer that this topic title refers the
discussions about Barack Obama’s birth certificate: people
were questioning whether Barack Obama was truly born in
the United States.</p>
        <p>We rely on the judged tweets for our analysis and
investigate topic-dependent as well as topic-independent features.
Examples of topic-dependent features are the retrieval score
derived from retrieval strategies that are based on document
and corpus statistics as well as the semantic overlap score
which determines the extent of overlap between the
semantic meaning of a search topic and a tweet. In addition to
these topic-dependent features, we also studied a number of
topic-independent features: syntactical features (such as the
presence of URLs or hashtags in a tweet), semantic features
(such as the diversity of the semantic concepts mentioned in
a tweet) and social context features (such as the authority
of the user who published the tweet).</p>
        <p>The main contributions of our work can be summarized
as follows:</p>
        <p>• We present a set of strategies for the extraction of
fea4http://trec.nist.gov/
5http://sites.google.com/site/trecmicroblogtrack/
(keyword-based and semantic-based relevance) and
topicinsensitive measures that do not consider the actual topic
but solely exploit syntactical or semantic tweet
characteristics. Finally, we also consider contextual features that, for
example, characterize the creator of a tweet.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Keyword-based Relevance Features</title>
      <p>
        keyword-based relevance score (Indri-based query
relevance): To calculate the retrieval score for pair of (topic,
tweet), we employ the language modeling approach to
information retrieval [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. A language model θt is derived
for each document (tweet). Given a query Q with terms
Q = {q1, ..., qn} the document language models are ranked
with respect to the probability P (θt|Q), which according to
the Bayes theorem can be expressed as:
      </p>
      <p>P (θt|Q) =</p>
      <p>P (Q|θt)P (θt)</p>
      <p>P (Q)
∝ P (θt) Y P (qi|θt).</p>
      <p>
        qi∈Q
This is the standard query likelihood based language
modeling setup which assumes term independence. Usually, the
prior probability of a tweet P (θt) is considered to be
uniform, that is, each tweet in the corpus is equally likely. The
language models are multinomial probability distributions
over the terms occurring in the tweets. Since a maximum
likelihood estimate of P (qi|θt) would result in a zero
probability of any tweet that misses one or more of the query terms
in Q, the estimate is usually smoothed with a background
language model, generated over all tweets in the corpus. We
employed Dirichlet smoothing [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]:
tures from Twitter messages that allow us to predict
the relevance of a post for a given topic.
• Given a set of more than 38,000 tweets that were
manually labeled as relevant or not relevant for a set of
49 topics, we analyze the features and characteristics
of relevant and interesting tweets.
• We evaluate the effectiveness of the different features
for predicting the relevance of tweets for a topic and
investigate the impact of the different features on the
quality of the relevance classification. We also study to
what extent the success of the classification depends on
the type of topics (e.g. topics of short-term vs. topics
of long-term interest) for which relevant tweets should
be identified.
2.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        Since its launch in 2006 Twitter attracted a lot of
attention, both in the general public as well as in the
research community. Researchers started studying
microblogging phenomena to find out what kind of information is
discussed on Twitter [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], how trends evolve on Twitter [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], or
how one detects influential users on Twitter [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
Applications have been researched that utilize microblogging data to
enrich traditional news media with information from
Twitter [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], to detect and manage emergency situations such
as earthquakes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or to enhance search and ranking of
Web sites which possibly have not been indexed yet by Web
search engines.
      </p>
      <p>
        So far, search on Twitter or other microblogging
platforms such as Sina Weibo has not been studied extensively.
Teevan et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] compared the search behavior on Twitter
with traditional Web search behavior. It was found that
keyword queries that people issue to retrieve information from
Twitter are, on average, significantly shorter than queries
submitted to traditional Web search engines (1.64 words vs.
3.08 words). This finding indicates that there is a demand
to investigate new algorithms and strategies for retrieving
relevant information from microblogging streams.
      </p>
      <p>
        Bernstein et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] proposed an interface that allows for
exploring tweets by means of tag clouds. However, their
interface is targeted towards browsing the tweets that have
been published by the people whom a user is following and
not for searching the entire Twitter corpus. Jadhav et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
developed an engine that enriches the semantics of Twitter
messages and allows for issuing SPARQL queries on
Twitter streams. In previous work, we followed such a semantic
enrichment strategy to provide faceted search capabilities
on Twitter [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Duan et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] investigated features such
as Okapi BM25 relevance scores or Twitter specific features
(length of a tweet, presence or absence of a URL or
hashtag, etc.) in combination with RankSVM to learn a ranking
model for tweets (learning to rank). In an empirical study,
they found that the length of a tweet and information about
the presence of a URL in a tweet are important features to
rank relevant tweets. In this paper, we re-visit some of the
features proposed by Duan et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and introduce novel
semantic measures that allow us to estimate whether a
micropost is relevant to a given topic or not.
      </p>
    </sec>
    <sec id="sec-4">
      <title>FEATURES OF MICROPOSTS</title>
      <p>In this section, we provide an overview of the different
features that we analyze to estimate the relevance of a Twitter
message to a given topic. We present topic-sensitive
features that measure the relevance with respect to the topic
#MSM2012
(1)
(2)
(3)
P (qi|θt) =
c(qi, t) + μP (qi|θC ) .</p>
      <p>|t| + μ
Here, μ is the smoothing parameter, c(qi, t) is the count of
term qi in t and |t| is the length of the tweet. The probability
P (qi|θC ) is the maximum likelihood probability of term qi
occurring in the collection language model θC (derived by
concatenating all tweets in the corpus).</p>
      <p>Due to the very small probabilities of P (Q|θt), we utilize
log (P (Q|θt)) as feature scores. Note that this score is always
negative. The greater the score (that is, the less negative),
the more relevant the tweet is to the query.
3.2</p>
      <p>
        Semantic-based Relevance Features
semantic-based relevance score This feature is also a
retrieval score calculated according to Section 3.1 though
with a different set of queries. Since the average length
of search queries submitted to microblog search engines is
lower than in traditional Web search, it is necessary to
understand the information need behind the query. The search
topics provided as part of the TREC data set contain
abbreviations, part of names, and nicknames. One example (cf.
Table 1) is the first name “Jintao” (in the query: “Jintao
visit US”) which refers to the President of the People’s
Republic of China. However, in tweets he is also referred to as
“President Hu”, “Chinese President”, etc. If these semantic
variants of a person’s name and titles would be considered
when deriving an expanded query, a wider variety of
potentially relevant tweets could be found. We utilize the
wellknown Named-Entity-Recognition (NER) service DBPedia
hasHashtag This is a boolean property which indicates
whether a given tweet contains at least one hashtag or not.
Twitter users typically apply hashtags in order to facilitate
the retrieval of the tweet. For example, by using a hashtag
people can join a discussion on a topic that is represented via
that hashtag. Users, who monitor the hashtag, will retrieve
all tweets that contain it. Teevan et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] showed that
such monitoring behavior is a common practice on Twitter
to retrieve relevant Twitter messages. Therefore, we
investigate whether the occurrence of hashtags (possibly without
any obvious relevance to the topic) is an indicator for the
relevance and interestingness of a tweet.
      </p>
      <p>
        Hypothesis H1: tweets that contain hashtags are more likely
to be relevant than tweets that do not contain hashtags.
hasURL Dong et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] showed that people often exchange
URLs via Twitter so that information about trending URLs
can be exploited to improve Web search and particularly the
ranking of recently discussed URLs. Therefore, the presence
of a URL (boolean property) can be an indicator for the
relevance of a tweet.
      </p>
      <p>
        Hypothesis H2: tweets that contain a URL are more likely
to be relevant than tweets that do not contain a URL.
isReply On Twitter, users can reply to the tweets of other
people. This type of communication can, for example, be
used to comment on a certain message, to answer a
question or to chat with other people. Chen et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] studied
the characteristics of reply chains and discovered that one
can distinguish between users who are merely interested in
news-related information and users who are also interested
in social chatter. For deciding whether a tweet is relevant for
a news-related topic, we therefore assume that the boolean
isReply feature, which indicates whether a tweet is a reply
to another tweet, can be a valuable signal.
      </p>
      <p>Hypothesis H3: tweets that are formulated as a reply to
another tweet are less likely to be relevant than other tweets.
length The length of a tweet—measured in the number of
characters—may also be an indicator for the relevance or</p>
      <sec id="sec-4-1">
        <title>6DBpedia Spotlight, http://spotlight.dbpedia.org/</title>
        <p>interestingness. We hypothesize that the length of a Twitter
message correlates with the amount of information that is
conveyed in the message.</p>
        <p>Hypothesis H4: the longer a tweet, the more likely it is to be
relevant and interesting.</p>
        <p>The values of boolean properties are set to 0 (false) and 1
(true) while the length of a Twitter message is measured
by the number of characters divided by 140 which is the
maximum length of a Twitter message.</p>
        <p>There are further syntactical features that can be explored
such as the mentioning of certain character sequences
including emoticons, question marks, exclamation marks, etc. In
line with the isReply feature, one could also utilize
knowledge about the re-tweet history of a tweet, e.g. a boolean
property that indicates whether the tweet is a copy from
another tweet or a numeric property that counts the number
of users who re-tweeted the message. However, in this paper
we are merely interested in original messages that have not
been re-tweeted yet7 and therefore also merely in features
which do not require any knowledge about the history of a
tweet. This allows us to estimate the relevance of a message
as soon as it is published.
3.4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Semantic Features</title>
      <p>In addition to the semantic relevance scores described in
Section 3.2, one can also analyze the semantics of a Twitter
message independently from the topic of interest. We
therefore utilize again the DBpedia entity extraction provided by
DBpedia Spotlight to extract the following features:
#entities The number of DBpedia entities that are
mentioned in a Twitter message may give further evidence about
the potential relevance and interestingness of a tweet. We
assume that the more entities can be extracted from a tweet,
the more information it contains and the more valuable it
is. For example, in the context of the discussion about birth
certificates we find the following two tweets in our dataset:
t1: “Despite what her birth certificate says, my lady is
actually only 27”
t2: “Hawaii (Democratic) lawmakers want release of Obama’s
birth certificate”
When reading the two tweets, without having a particular
topic or information need in mind, it seems that t2 has a
higher likelihood to be relevant for some topic for the
majority of the Twitter users than t1 as it conveys more entities
that are known to the public and available on Wikipedia
and DBpedia respectively. In fact, the entity extractor is
able to detect one entity, db:Birth certificate, for tweet t1
while it detects three additional entities for t2: db:Hawaii,
db:Legislator and db:Barack Obama.</p>
      <p>Hypothesis H5: the more entities a tweet mentions, the more
likely it is to be relevant and interesting.
#entities(type) Similarly to counting the number of
entities that occur in a Twitter message, we also count the
number of entities of specific types. The rationale behind
this feature being that some types of entities might be a
stronger indicator for relevance than others. The
importance of a specific entity type may also depend on the topic.
7This is in line with the relevance judgments provided by
TREC which did not consider re-tweeted messages.
#MSM2012
For example, when searching for Twitter messages that
report about wild fires in a specific area, location-related
entities may be more interesting than product-related entities.
In this paper, we count the number of entity occurrences in a
Twitter message for five different types: locations, persons,
organizations, artifacts and species (plants and animals).
Hypothesis H6: different types of entities are of different
importance for estimating the relevance of a tweet.
diversity The diversity of semantic concepts mentioned in
a Twitter message can also be exploited as an indicator for
the potential relevance and interestingness of a tweet. We
therefore count the number of distinct types of entities that
are mentioned in a Twitter message. For example, for the
two tweets t1 and t2 mentioned earlier, the diversity score
would be 1 and 4 respectively as for t1 only one type of
entity is detected (yago:PersonalDocuments ) while for t2
also instances of db:Person (person), db:Place (location) and
owl:Thing (the role db:Legislator is not further classified) are
detected.</p>
      <p>
        Hypothesis H7: the greater the diversity of concepts
mentioned in a tweet, the more likely it is to be interesting and
relevant.
sentiment Naveed et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] showed that tweets which
contain negative emoticons are more likely to be re-tweeted than
tweets which feature positive emoticons. The sentiment of
a tweet may thus impact the perceived relevance of a tweet.
Therefore, we classify the the semantic polarity of a tweet
into positive, negative or neutral using Twitter Sentiment 8.
Hypothesis H8: the likelihood of a tweet’s relevance is
influenced by its sentiment polarity.
3.5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Contextual Features</title>
      <p>In addition to the aforementioned features, which describe
characteristics of the Twitter messages, we also investigate
features that describe the context in which a tweet was
publish. In our analysis, we investigate the social and temporal
context:
social context The social context describes the creator of
a Twitter message. Different characteristics of the message
creator may increase or decrease the likelihood of her tweets
being relevant and interesting such as the number of
followers or the number of tweets from this user that have been
re-tweeted. In this paper, we apply a light-weight measure
to characterize the creator of a message: we count the
number of tweets which the user has published.</p>
      <p>Hypothesis H9: the higher the number of tweets that have
been published by the creator of a tweet, the more likely it is
that the tweet is relevant.
temporal context The temporal context describes when
a tweet was published. The creation time can be specified
with respect to the time when a user is requesting tweets
about a certain topic (query time) or it can be independent
of the query time. For example, one could specify at which
hour during the day the tweet was published or whether it
was created during the weekend. In our analysis, we utilize
the temporal distance (in seconds) between the query time
and the creation time of the tweet. Hypothesis H10: the
lower the temporal distance between the query time and the
creation time of a tweet, the more likely is the tweet relevant
to the topic.</p>
      <sec id="sec-6-1">
        <title>8http://twittersentiment.appspot.com/</title>
        <p>Contextual features may also refer to characteristics of
Web pages that are linked from a Twitter message. For
example, one could exploit the PageRank scores of the
referenced Web sites to estimate the relevance of a tweet or one
could categorize the linked Web pages to discover the types
of Web sites that usually attract attention on Twitter. We
leave the investigation of such additional contextual features
for future work.
4.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>FEATURE ANALYSIS</title>
      <p>In this section, we describe and characterize the Twitter
corpus with respect to the features that we presented in the
previous section.
4.1</p>
    </sec>
    <sec id="sec-8">
      <title>Dataset Characteristics</title>
      <p>We use the Twitter corpus which was used in the
microblog track of TREC 20119. The original corpus consists
of approximately 16 million tweets, posted over a period
of 2 weeks (January 24 until February 8th, inclusive). We
utilized an existing language detection library10 to identify
English tweets and found that 4,766,901 tweets were
classified as English. Employing NER on the English tweets
resulted in a total over six million named entities among
which we found approximately 0.14 million distinct entities.
Besides the tweets, 49 topics were given as the targets of
retrieval. TREC assessors judged the relevance of 40,855
topic-tweet pairs which we use as ground truth in our
experiments. 2,825 tweets were judged as relevant for a given
topic while the majority of the tweet-topic pairs (37,349)
were marked as non-relevant.
4.2</p>
    </sec>
    <sec id="sec-9">
      <title>Feature Characteristics</title>
      <p>In Table 2 we list the average values and the standard
deviations of the features and the percentages of true instances
for boolean features respectively. It shows that relevant and
non-relevant tweets show, on average, different
characteristics for several features.</p>
      <p>As expected, the average keyword-based relevance score
of tweets, which are judged as relevant for a given topic, is
much higher than the one for non-relevant tweets: -10.709
in comparison to -14.408 (the higher the value the better,
see Section 3.1). Similarly, the semantic-based relevance
score, which exploits the semantic concepts mentioned in
the tweets (see Section 3.2) while calculating the retrieval
rankings, shows the same characteristic. The
isSemanticallyReleated feature, which is a binary measure of the overlap
between the semantic concepts mentioned in the query and
the respective tweets, is also higher for relevant tweets than
for non-relevant tweets. Hence, when we consider the
topicdependent features (keyword-based and semantic-based), we
find first indicators that the hypotheses behind these
features hold.</p>
      <p>For the syntactical features we observe that, regardless of
whether the tweets are relevant to a topic or not, the ratios of
tweets that contain hashtags are almost the same (about 19%).
Hence, it seems that the presence of a hashtag is not
necessarily an indicator for relevance. However, the presence
of a URL is potentially a very good indicator: 81.9% of
the relevant tweets feature a URL whereas only 54.1% of
the non-relevant tweets contain a URL. A possible
explana</p>
      <sec id="sec-9-1">
        <title>9http://trec.nist.gov/data/tweets/</title>
        <p>10Language detection, http://code.google.com/p/
language-detection/
#MSM2012
Category
Relevant
tion for this difference is that the tweets containing URLs
tend to feature also an attractive short title, especially for
breaking news, in order to attract people to follow the link.
Moreover, the actual content of the linked Web site may
also stipulate users when assessing the relevance of a tweet.
In Hypothesis 3 (see Section 3.3), we speculate that
messages which are replies to other tweets are less likely to be
relevant than other tweets. The results listed in Table 2
support this hypothesis: only 3.4% of the relevant tweets
are replies in contrast to 14.2% of the non-relevant tweets.
The length of the tweets that are judged as relevant is, on
average, 90.3 characters, which is slightly longer than for the
non-relevant ones (87.8 characters).</p>
        <p>The comparison of the topic-independent semantic
features also reveals some differences between relevant and
nonrelevant tweets. Overall, relevant tweets contain more
entities (2.4) than non-relevant tweets (1.9). Among the five
most frequently mentioned types of entities, persons,
organizations, and locations occur more often in relevant tweets
than in non-relevant ones. On average, messages are
therefore considered as more likely to be relevant or interesting
for users if they contain information about people, involved
organizations, or places. Artifacts (e.g. tangible things,
software) and species (e.g. plants, animals) are more frequent
in non-relevant tweets. However, counting the number of
entities of type species seems to be a less promising feature
since the fraction of tweets which mention a species is fairly
low.</p>
        <p>
          The diversity of content mentioned in a Twitter message—
i.e. the number of distinct types (only person, organization,
location, artifact, and species are considered)—is potentially
a good feature: the semantic diversity is higher for the
relevant tweets (0.8) than for the non-relevant ones (0.6). In
addition to the entities that are mentioned in the tweets,
we also conducted a sentiment analysis of the tweets (see
Section 3.4). Although most of the tweets are neutral
(sentiment score = 0), the average sentiment score for relevant
tweets is negative (-0.025). This observation is in line with
the finding made by Naveed et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] who found that
negative tweets are more likely to be re-tweeted.
        </p>
        <p>Finally, we also attempted to determine the relationship
between a tweet’s likelihood of relevance and its context.
With respect to the social context, we however do not
observe a significant difference between relevant an non-relevant
tweets: users who publish relevant tweets are, on average,
not more active than publishers of non-relevant tweets (12.3
vs. 12.2). For the temporal context, the average distance
between the time when a user requests tweets about a topic
and the creation time of tweets is 4.85 days for relevant
tweets and 3.98 for non-relevant tweets. However, the
standard deviations of these scores is with 4.53 days (relevant)
and 4.39 days (non-relevant) fairly high. This indicates that
the temporal context is not a reliable feature for our dataset.
Preliminary experiments indeed confirmed the low utility of
the temporal feature. However, this observation seems to be
strongly influenced by the TREC dataset itself which was
collected within a short time period of time (two weeks). In
our evaluations, we therefore do not consider the temporal
context and leave an analysis of the temporal features for
future work.
5.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>EVALUATION OF FEATURES FOR REL</title>
    </sec>
    <sec id="sec-11">
      <title>EVANCE PREDICTION</title>
      <p>Having analyzed the dataset and the proposed features,
we now evaluate the quality of the features for predicting
the relevance of tweets for a given topic. We first outline the
experimental setup before we present our results and analyze
the influence of the different features on the performance for
the different types of topics.
5.1</p>
    </sec>
    <sec id="sec-12">
      <title>Experimental Setup</title>
      <p>We employ logistic regression to classify tweets as
relevant or non-relevant to a given topic. Due to the small size
of the topic set (49 topics), we use 5-fold cross validation
to evaluate the learned classification models. For the final
setup, 16 features were used as predictor variables (all
features listed in Table 2 except for the temporal context). To
conduct our experiments, we rely on the machine learning
toolkit Weka11. As the number of relevant tweets is
considerably smaller than the number of non-relevant tweets, we
employed a cost-sensitive classification setup to prevent the
classifier from following a best match strategy where simply
all tweets are marked as non-relevant. As the estimation for
the negative class achieves a precision and recall both over
90%, we focus on the precision and recall of the relevance
classification (the positive class) in our evaluation as we aim
to investigate the characteristics that make tweets relevant
to a given topic.
11http://www.cs.waikato.ac.nz/ml/weka/
#MSM2012</p>
      <p>Table 3 shows the performances of estimating the
relevance of tweets based on different sets of features. Learning
the classification model solely based on the keyword-based
or semantic-based relevance scoring features leads to an
FMeasure of 0.2981 and 0.2991 respectively. There is thus no
notable difference between the two topic-sensitive features.
However, by combining both features (see topic-sensitive in
Table 3) the F-Measure increases which is caused by a higher
recall, increasing from 0.29 to 0.34. It appears that the
keyword-based and semantic-based relevance scores
complement each other.</p>
      <p>As expected, when solely learning the classification model
based on the topic-independent features—i.e. without
measuring the relevance to the given topic—the quality of the
relevance prediction is poor. The best performance is achieved
when all features are combined. A precision of 36.74% means
that more than a third of all tweets that our approach
classifies as relevant are indeed relevant, while the recall level
(47.36%) implies that our approach discovers nearly half of
all relevant tweets. Since microblog messages are very short,
a significant number of tweets can be read quickly by a user
when presented in response to her search request. In such a
setting, we believe such a classification accuracy to be
sufficient. Overall, the semantic features seem to play an
important role as they lead to a performance improvement with
respect to the F-Measure from 0.3965 to 0.4138. We will
now analyze the impact of the different features in detail.</p>
      <p>One of the advantages of the logistic regression model is,
that it is easy to determine the most important features
of the model by considering the absolute weights assigned
to them. For this reason, we have listed the relevant-tweet
prediction model coefficients for all employed features in
Table 4. The features influencing the model the most are:
• hasURL: Since the feature coefficient is positive, the
presence of a URL in a tweet is more indicative of
relevance than non-relevance. That means, that
hypothesis H2 (Section 3.3) holds.
• isSemanticallyRelated : The overlap between the
identified DBpedia concepts in the topics and the identified
DBpedia concepts in the tweets is the second most
important feature in this model. This is an interesting
observation, especially in comparison to the
keywordbased relevance score, which is only the ninth
important feature among the evaluated ones. It implies that
a standard keyword-based retrieval approach, which
performs well for longer documents, is less suitable for
microposts.
• isReply : This feature, which is true (= 1) if a tweet is
written in reply to a previously published tweet has a
negative coefficient which means that tweets which are
replies are less likely to be in the relevant class than
tweets which are not replies, confirming hypothesis H3
(Section 3.3).
• sentiment : The coefficient of the sentiment feature is
similarly negative, which suggests that a negative
sentiment is more predictive of relevance than a positive
sentiment, in line with our hypothesis H8 (Section 3.4).</p>
      <p>We note that the keyword-based similarity, while being
positively aligned with relevance, does not belong to the
most important features in this model. It is superseded by
syntactic as well as semantic-based features. When we
consider the non-topical features only, we observe that
interestingness (independent of a topic) is related to the
potential amount of additional information (i.e. the presence of
a URL), the clarity of the tweet overall (a tweet in reply
may be only understandable in the context of the
contextual tweets) and the different aspects covered in the tweet (as
evident in the diversity feature). It should also be pointed
out that the negative coefficients assigned to most
topicinsensitive entity count features (#entities(X)) is in line
with the results in Table 2.
5.3</p>
    </sec>
    <sec id="sec-13">
      <title>Influence of Topic Characteristics on Relevance Prediction</title>
      <p>In all reported experiments so far, we have considered the
entire set of topics available to us. In this section, we
investigate to what extent certain topic characteristics play a role
for relevance prediction and to what extent those differences
lead to a change in the logistic regression models.</p>
      <p>Consider the following two topics: Taco Bell filling lawsuit
(MB02012) and Egyptian protesters attack museum (MB010).
While the former has a business theme and is likely to be
mostly of interest to American users, the latter topic belongs
into the politics category and can be considered as being of
global interest, as the entire world was watching the events
in Egypt unfold. Due to these differences we defined a
number of topic splits. A manual annotator then decided for
each split dimension into which category the topic should
fall. We investigated four topic splits, three splits with two
12The identifiers of the topics correspond to the ones used in
the official TREC dataset.
#MSM2012
Feature Category
partitions each and one split with five partitions:
• Popular/unpopular: The topics were split into popular
(interesting to many users) and unpopular (interesting
to few users) topics. An example of a popular topic is
2022 FIFA soccer (MB002) - in total we found 24. In
contrast, topic NIST computer security (MB005) was
classified as unpopular (as one of 25 topics).
• Global/local: In this split, we considered the
interest for the topic across the globe. The already
mentioned topic MB002 is of global interest, since soccer
is a highly popular sport in many countries, whereas
topic Cuomo budget cuts (MB019) is mostly of local
interest to users living or working in New York where
Andrew Cuomo is the current governor. We found 18
topics to be of global and 31 topics to be of local
interest.
• Persistent/occasional: This split is concerned with the
interestingness of the topic over time. Some topics
persist for a long time, such as MB002 (the FIFA world
cup will be played in 2022), whereas other topics are
only of short-term interest, e.g. Keith Olbermann new
job (MB030). We assigned 28 topics to the persistent
and 21 topics to the occasional topic partition.
• Topic themes: The topics were classified as belonging
to one of five themes, either business, entertainment,
sports, politics or technology. While MB002 is a sports
topic, MB019 for instance is considered to be a
political topic.</p>
      <p>Our discussion of the results focuses on two aspects: (i)
the difference between the models derived for each of the
two partitions, and, (ii) the difference between these models
(denoted MsplitName) and the model derived over all topics
(MallT opics) in Table 4. The results for the three binary
topic splits are shown in Table 5.</p>
      <p>Popularity: A comparison of the most important
features of Mpopular and Munpopular shows few differences with
the exception of a single feature: sentiment. While
sentiment, and in particular a negative sentiment, is the third
most important feature in Mpopular, it is ranked eighth in
Munpopular. We hypothesize that unpopular topics are also
partially unpopular because they do not evoke strong
emotions in the users. A similar reasoning can be applied when
considering the amount of relevant tweets discovered for
both topic splits: while on average 67.3 tweets were found to
be relevant for popular topics, only 49.9 tweets were found
to be relevant for unpopular topics (the average number of
relevant tweets across the entire topic set is 58.44).</p>
      <p>Global vs. local: This split did not result in
models that are significantly different from each other or from
MallTopics, indicating that—at least for our currently
investigated features—a distinction between global and local topics
is not useful.</p>
      <p>Temporal persistence: The same conclusion can be
drawn about the temporal persistence topic split; for both
models the same features are of importance which in turn
are similar to MallTopics. However, it is interesting to see
that the performance (regarding all metrics) is clearly higher
for the occasional (short-term) topics in comparison to the
persistent (long-term) topics. For topics that have a short
lifespan recall and precision are notably higher than for the
other types of topics.</p>
      <p>Topic Themes: The results of the topic split
according to the theme of the topic are shown in Table 6. Three
topics did not fit in one of the five categories. Since the
topic set is split into five partitions, the size of some
partitions is extremely small, making it difficult to reach
conclusive results. We can, though, detect trends, such as the
fact that relevant tweets for business topics are less likely to
contain hashtags (negative coefficient), while the opposite
holds for entertainment topics (positive coefficient). The
#MSM2012
Feature Category
entertainment
politics</p>
      <p>technology
semantic similarity has a large impact on all themes but
entertainment. Another interesting observation is that
sentiment, and in particular negative sentiment, is a prominent
feature in Mbusiness and in Mpolitics but less so in the other
models.</p>
      <p>Finally we note that there are also some features which
have no impact at all, independent of the topic split
employed: the length of the tweet and the social context of
the user posting the message. The observation that certain
topic splits lead to models that emphasize certain features
also offers a natural way forward: if we are able to determine
for each topic in advance to which theme or topic
characteristic it belongs to, we can select the model that fits the
topic best.</p>
    </sec>
    <sec id="sec-14">
      <title>6. CONCLUSIONS</title>
      <p>In this paper, we have analyzed features that can be used
as indicators of a tweet’s relevance and interestingness to
a given topic. To achieve this, we investigated features
along two dimensions: topic-dependent features and
topicindependent features. We evaluated the utility of these
features with a machine learning approach that allowed us to
gain insights into the importance of the different features for
the relevance classification.</p>
      <p>Our main discoveries about the factors that lead to
relevant tweets are the following: (i) The learned models which
take advantage of semantics and topic-sensitive features
outperform those which do not take the semantics and
topicsensitive features into account. (ii) The length of tweets and
the social context of the user posting the message have little
impact on the prediction. (iii) The importance of a feature
differs depending on the characteristics of the topics. For
example, the sentiment-based feature is more important for
popular than for unpopular topics and the semantic
similarity does not have a significant impact on entertaining topics.</p>
      <p>The work presented here is beneficial for search &amp; retrieval
of microblogging data and contributes to the foundations of
engineering search engines for microposts. In the future, we
plan to investigate the social and the contextual features in
depth. Moreover, we would like to investigate to what
extent personal interests of the users (possibly aggregated from
different Social Web platforms) can be utilized as features
for personalized retrieval of microposts.</p>
    </sec>
    <sec id="sec-15">
      <title>7. REFERENCES</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Abel</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Celik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Siehndel</surname>
          </string-name>
          .
          <article-title>Leveraging the Semantics of Tweets for Adaptive Faceted Search on Twitter</article-title>
          .
          <source>In ISWC '11</source>
          , Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Suh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kairam</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          .
          <article-title>Eddi: interactive topic-based browsing of social status streams</article-title>
          .
          <source>In UIST '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nairn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          .
          <article-title>Speak Little and Well: Recommending Conversations in Online Social Streams</article-title>
          .
          <source>In CHI '11</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , P. Kolari,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Zha</surname>
          </string-name>
          .
          <article-title>Time is of the essence: improving recency ranking using twitter data</article-title>
          .
          <source>In WWW '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.-Y.</given-names>
            <surname>Shum</surname>
          </string-name>
          .
          <article-title>An empirical study on learning to rank of tweets</article-title>
          .
          <source>In COLING '10</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jadhav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Purohit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kapanipathi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ananthram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ranabahu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cooney</surname>
          </string-name>
          , ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Sheth</surname>
          </string-name>
          .
          <article-title>Twitris 2.0 : Semantically Empowered System for Understanding Perceptions From Social Data</article-title>
          .
          <source>In Semantic Web Challenge</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kwak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Moon</surname>
          </string-name>
          .
          <article-title>What is twitter, a social network or a news media?</article-title>
          <source>In WWW '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mathioudakis</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Koudas</surname>
          </string-name>
          .
          <article-title>Twittermonitor: trend detection over the twitter stream</article-title>
          .
          <source>In SIGMOD '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Naveed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gottron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kunegis</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Alhadi</surname>
          </string-name>
          .
          <article-title>Bad news travel fast: A content-based analysis of interestingness on twitter</article-title>
          .
          <source>In WebSci '11</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sakaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Okazaki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsuo</surname>
          </string-name>
          .
          <article-title>Earthquake shakes Twitter users: real-time event detection by social sensors</article-title>
          .
          <source>In WWW '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramage</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Morris</surname>
          </string-name>
          . #
          <article-title>TwitterSearch: a comparison of microblog search and web search</article-title>
          .
          <source>In WSDM '11</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.-P.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jiang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>He</surname>
          </string-name>
          .
          <article-title>Twitterrank: finding topic-sensitive influential twitterers</article-title>
          .
          <source>In WSDM '10</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Lafferty</surname>
          </string-name>
          .
          <article-title>A study of smoothing methods for language models applied to ad hoc information retrieval</article-title>
          .
          <source>In SIGIR '01</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          ,
          <year>2001</year>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>