<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of RepLab 2013: Evaluating Online Reputation Monitoring Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrique Amigo</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Carrillo de Albornoz</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irina Chugur</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adolfo Corujo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tamara Mart n</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edgar Meij</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maarten de Rijke</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Damiano Spina</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ISLA, University of Amsterdam Science Park 904</institution>
          ,
          <addr-line>1098 XH Amsterdam</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Llorente &amp; Cuenca Lagasca</institution>
          ,
          <addr-line>88. 28001 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>UNED NLP &amp; IR Group Juan del Rosal</institution>
          ,
          <addr-line>16. 28040 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Yahoo! Research Diagonal 177</institution>
          ,
          <addr-line>08018 Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We summarize the goals, organization, and results of the second RepLab competitive evaluation campaign for Online Reputation Management systems (RepLab 2013). RepLab 2013 focuses on the process of monitoring the reputation of companies and individuals, and asks participating systems to annotate di erent types of information on tweets containing the names of several companies. First, tweets have to be classi ed as related or unrelated to the entity; relevant tweets have to be classi ed according to their polarity for reputation (Does the content of the tweet have positive or negative implications for the reputation of the entity?), clustered in coherent topics, and clusters have to be ranked according to their priority (potential reputation problems had to come rst). The gold standard consists of more than 140,000 tweets annotated by a group of trained annotators supervised and monitored by reputation experts.</p>
      </abstract>
      <kwd-group>
        <kwd>RepLab</kwd>
        <kwd>Reputation Management</kwd>
        <kwd>Evaluation Methodologies and Metrics</kwd>
        <kwd>Test Collections</kwd>
        <kwd>Text Clustering</kwd>
        <kwd>Sentiment Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In a world of online networked information, where its control has moved to users
and consumers, every move of a company and every act of a public gure are
subject, at all times, to the scrutiny of a powerful global audience. While
traditional reputation analysis is mostly manual, online media one allow to process,
understand and aggregate large streams of facts and opinions about a company
or individual. In this context, natural language processing plays a key, enabling
role and we are witnessing an unprecedented demand for text mining software for
ORM. Although opinion mining has made signi cant advances in recent years,
most of the work has focused on products. However, mining and interpreting
opinions about companies and individuals is, in general, a much harder and less
understood problem, since unlike products or services, opinions about people and
organizations cannot be structured around any xed set of features or aspects,
requiring a more complex modeling of these entities.</p>
      <p>
        RepLab is an initiative promoted by the EU project LiMoSINe5 which aims
at enabling research on reputation management as a \living lab": a series of
evaluation campaigns in which task design and evaluation are jointly carried
out by researchers and the target user communities (reputation management
experts). Like its rst edition in 2012 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], RepLab 2013 has been organized as a
CLEF lab, and the results of the exercise are discussed at CLEF 2013 in Valencia,
Spain, on 23{26th September.
      </p>
      <p>RepLab 2013 has been focused on the task of monitoring the reputation of
entities (companies, organizations, celebrities, etc.) on Twitter. The monitoring
task for analysts consists of searching the stream of tweets for potential mentions
to the entity, ltering those that do refer to the entity, detecting topics (i.e.,
clustering tweets by subject) and ranking them based on the degree to which
they are potential reputation alerts (i.e., issues that may have a substantial
impact on the reputation of the entity, and must be handled by reputation
management experts).
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Tasks</title>
      <sec id="sec-2-1">
        <title>Task De nition</title>
        <p>Following the outline given above, the RepLab 2013 task is de ned as
(multilingual) topic detection combined with priority ranking of the topics, as input
for reputation monitoring experts. The detection of polarity for reputation (does
the tweet have negative/positive implications for the reputation of the entity?)
is an essential step to assign priority, and is evaluated as a standalone subtask.</p>
        <p>Participants were welcome to present systems that attempt the full
monitoring task ( ltering + topic detection + topic ranking) or modules that contribute
only partially to solve the problem. Subtasks that are explicitly considered in
RepLab 2013 are:
{ Filtering. Systems are asked to determine which tweets are related to the
entity and which are not. For instance, distinguishing between tweets that
contain the word \Stanford" referring to the University of Stanford and
ltering out tweets about Stanford as a place. Manual annotations are provided
with two possible values: related/unrelated.
{ Polarity for reputation classi cation. The goal is to decide if the tweet
content has positive or negative implications for the company's reputation.
Manual annotations are: positive/negative/neutral.
5 http://www.limosine-project.eu
{ Topic detection. Systems are asked to cluster related tweets about the entity
by topic with the objective of grouping together tweets referring to the same
subject/event/conversation.
{ Priority assignment. The full task involves detecting the relative priority of
topics. So as to be able to evaluate priority independently from the clustering
task, we will evaluate the subtask of predicting the priority of the cluster a
tweet belongs to.</p>
        <p>A substantial di erence between RepLab 2013 and its rst edition in 2012 is
that, in 2013, the training and test entities are the same, and therefore
conventional machine learning techniques are readily applicable. RepLab 2013 models
a scenario where reputation experts are constantly tracking and annotating
information about a client (entity), and therefore it is likely to have manual
annotations for data related to the entity of interest. RepLab 2012, on the other
hand, modeled the scenario of a web application that can be used by anyone, at
any time, using any entity name as keyword. In that case, training material was
referred to entities other than those in the training set.</p>
        <p>In RepLab 2013 it was possible to present systems that address only
ltering, only polarity identi cation, only topic detection or only priority assignment.
Another di erence with 2012 is that in its second edition, the RepLab
organization provided baseline components for all of the four subtasks. This way any
participant was able to participate in the full task regardless of his particular
contribution or expertise.</p>
        <p>Some relevant details on the polarity for reputation and topic detection tasks
follow. Polarity for reputation is substantially di erent from standard sentiment
analysis. First, when analyzing polarity for reputation, both facts and opinions
have to be considered. For instance, \Barclays plans additional job cuts in the
next two years" is a fact with negative implications for reputation. Therefore,
systems will not be explicitly asked to classify tweets as factual vs.
opinionated: the goal is to nd polarity for reputation, that is, what implications a
piece of information might have on the reputation of a given entity, regardless of
whether the content is opinionated or not. Second, negative sentiments do not
always imply negative polarity for reputation and vice versa. For instance, \R.I.P.
Michael Jackson. We'll miss you" has a negative associated sentiment (sadness,
deep sorrow), but a positive implication for the reputation of Michael
Jackson. And the other way around, a tweet such as \I LIKE IT..... NEXT...MITT
ROMNEY...Man sentenced for hiding millions in Swiss bank account," has a
positive sentiment (joy about a sentence) but has a negative implication for the
reputation of Mitt Romney.</p>
        <p>As for the topic detection + topic ranking process, a three-valued classi
cation was applied to assess the priority of each entity-related topic: alert (the
topic deserves immediate attention of reputation managers), mildly relevant (the
topic contributes to the reputation of the entity but does not require
immediate attention) and unimportant (the topic can be neglected from a reputation
management perspective). Some of the factors that play a role in the priority
assessments are:
{ Polarity. Topics with polarity (and, in particular, with negative polarity,
where action is needed) usually have more priority.
{ Centrality. A high priority topic is very likely to have the company as the
main focus of the content.
{ User's authority. A topic promoted by an in uential (for example, in terms of
the number of followers or the expertise) user has better chances of receiving
high priority.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Baselines</title>
        <p>The baseline approach consists of tagging tweets (in the test set) with the same
tags of the closear tweet in the (entity) training set according to the Jaccard
word distance. The baseline is, therefore, a simple version of memory-based
learning. We have selected this approach for several reasons: (i) it is easy to
understand; (ii) it can be applied to every subtask in RepLab 2013; (iii) it keeps
the coherence between tasks: if a tweet is annotated as non-related, it will not
receive any priority or topic tag; (iv) it exploits the training data set per entity.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Evaluation Measures</title>
        <p>All subtasks consist of tagging single tweets according to their relatedness,
priority, polarity or topic. However, each one corresponds to a particular arti cial
intelligence problem: binary classi cation (relatedness), three-level classi cation
(polarity and priority), clustering (topic detection), and their concatenation (full
task). A common feature for all tasks is that the classes, levels or clusters can be
unbalanced. This entails challenges for the de nition of our evaluation
methodology. First, in classi cation tasks, a non informative system (i.e., all tweets to
the same class) can achieve high scores without providing useful information.
Second, in three-level classi cation tasks, a system could sort tweets correctly
without a perfect correspondence between predicted and true tags. Third, an
unbalanced cluster distribution across entities produces an important trade-o
between precision/recall oriented evaluation metrics (precision or cluster entropy
versus recall or class entropy) and that makes the measure combination function
crucial for system ranking.</p>
        <p>
          In evaluation, there is a hidden trade-o between interpretability and
strictness. For instance, Accuracy is easy to interpret: it simply reports how frequently
the system makes the correct decision. However, it is also easy to be cheated
under unbalanced test sets. For instance, returning all tweets in the same class,
cluster or level, may have high accuracy if the set is unbalanced. Other measures
based on information theory are more strict when penalizing non informative
outputs, but at the cost of interpretability. In this evaluation campaign we employ
Accuracy as a high interpretable measure, and the combination of Reliability
and Sensitivity (R&amp;S) as a strict and theory grounded measure [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Basically, R&amp;S assumes that any organization task consists of a bag of
relationships between documents. In our tasks, two documents are related if they
have di erent priority, polarity or relatedness level, or when they appear in the
same cluster. In brief, R&amp;S computes the precision and recall of relationships
produced by the systems with respect to the goldstandard. In order to avoid the
quadratic efect of document pairwise, R&amp;S is computed for each document
relationships and averaged in a second step. Reliability and Sensitivity are computed
as, being I the set of tweets considered in the evaluation:</p>
        <p>R(system) = Avgi2I R(i) S(system) = Avgi2I S(i)</p>
        <p>R(i) = Pj2I (relgold(i; j) = relsys(i; j)jrelsys(i; j))</p>
        <p>S(i) = Pj2I (relgold(i; j) = relsys(i; j)jrelgold(i; j));
where relgold(i; j) represents that i has a higher or lower polarity, priority or
relatedness than i, or that i and j belong to the same cluster. Relsys(i; j) is
analogous but applied to the system output.</p>
        <p>
          R&amp;S has three main strengths. First, it can be applied to ranking, ltering,
organization by levels and grouping tasks. This matches all the RepLab 2013
tasks. In addition, it gives the possibility to evaluate the full task as a whole.
Second, it covers simultaneously the desirable formal properties satis ed by other
measures in each particular task [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Third, according to experimental results
that we corroborate with RepLab 2013 data, R&amp;S is strict with respect to other
measures: a high score according to R&amp;S ensures a high score according to
any traditional measure. In other words, a low score according to one particular
traditional measure produces a low R&amp;S score, even when the system is rewarded
by other measures.
        </p>
        <p>
          R and S are combined with the F measure, i.e., a weighted harmonic mean
of R and S. This combining function is grounded in measure theory and satis es
a set of desirable constraints. One of the most useful is that a low score
according to one of the two measures strongly penalizes the combined score. However,
specially in clustering tasks, the F measure is seriously a ected by the relative
weight of partial measures (the parameter). In order to solve this, we
complement the evaluation results with the Unanimous Improvement Ratio, which has
been proved to be the only weighting independent combining criterion [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. UIR
is computed over the test cases (entities in RepLab) in which all measures
corroborates a di erence between runs. Let S1 and S2 be two runs and N&gt;8(S1; S2)
the amount of test cases for which S1 improves S2 for all measures, then:
U IR(S1; S2) =
        </p>
        <p>N&gt;8(S1; S2) N&gt;8(S2; S1)</p>
        <p>Amount of cases
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Dataset</title>
      <p>RepLab 2013 uses Twitter data in English and Spanish. The balance between
both languages depends on the availability of data for each of the entities
included in the dataset. The collection comprises tweets about 61 entities from
four domains: automotive, banking, universities and music. The domain
selection was done to o er a variety of scenarios for reputation studies. To this aim
we included entities whose reputation largely relies on their products
(automotive), entities for which transparency and ethical side of their activity are the
most decisive reputation factors (banking), entities for which the reputation of
which depends on a very broad and intangible set of products (universities) and,
nally, entities where the reputation is based almost equally on their products
and personal qualities (music bands and artists). Table 1 summarizes the
description of the corpus, as well as the number of tweets for both training and
test sets, and the distribution by language.</p>
      <p>Crawling was performed from 1 June, 2012 until 31 Dec, 2012 using each
entity's canonical name as query. For each entity, at least 2,200 tweets were
collected: the rst 700 were reserved for the training set and the last 1,500 for
the test collection. This distribution was set in this way to obtain a temporal
separation (ideally of several months) between the training and test data. The
corpus also comprises additional background tweets for each entity (up to 50,000,
with a large variability across entities). These are the remaining tweets situated
between the training (earlier tweets) and test material (the latest tweets) in the
timeline.
These data sets were manually labelled by thirteen annotators who were trained,
guided and constantly monitored by experts in ORM. Each tweet is annotated
as follows:
{ RELATED/UNRELATED: the tweet is/is not about the entity.
{ POSITIVE/NEUTRAL/NEGATIVE: the information contained in the tweet
has positive, neutral or negative implications for the entity's reputation.
{ Identi er of the topic cluster the tweet has been assigned to.
{ ALERT/MILDLY IMPORTANT/UNIMPORTANT: the priority of the topic
cluster the tweet belongs to.</p>
      <p>Table 2 shows statistics about the ltering subtask. The collection contains
110,352 tweets related with the entities, out of which 34,882 are in the training
set and 75,470 are in the test set. The 32,175 unrelated tweets of the dataset are
distributed as follows: 10,797 tweets in the training set and 21,378 in the test
set. The table also shows the distributions by domain.</p>
      <p>Table 3 shows the distribution of polarity classes in the RepLab 2013 dataset.
The RepLab 2013 dataset contains 63,442 tweets classi ed as positive by the
annotator, 30,493 classi ed as neutral and 16,415 classi ed as negative. The
distribution in the training set is 19,718 tweets classi ed as positive, 9,753 as
neutral and 5,409 as negative, while the test set contains 63,442 positive tweets,
30,493 neutral tweets and 16,415 negatives.</p>
      <p>Table 4 displays the number of topics per set as well as the average number
of tweets per topic, which is 17.77 for the whole collection but goes from 14.40 in
the training set to 21.14 in the test set. The training set contains 3,813 di erent
topics, the test set 5,757 di erent topics, for a total of 9,570 di erent topics in
the RepLab 2013 dataset.</p>
      <p>Finally, Table 5 summarizes the distributions of tweets in priority classes.
The less representative class is alert, with 4,780 tweets classi ed as a possible
reputation alert in the whole corpus. Mildly Important has 56,578 tweets and
Unimportant receives 48,992 tweets.</p>
      <p>In order to determine inter-annotator agreement we perform two di erent
experiments. First, 14 entities (4 automotive, 3 banking, 3 universities, 4
music) have been labeled by two annotators. This subset contains 31,381 tweets
that represent 22% of the RepLab 2013 dataset covering all domains. Second,
three annotators labeled 3 entities of the automotive domain. Table 6 shows
the results of the rst experiment of agreement using percentage of agreement
and Kappa metrics (both Cohen and Fleiss) for ltering, polarity and priority
detection tasks, and F measure of Reliability and Sensitivity for topic detection
task. As can be observed, the percentage of agreement for the ltering subtask
is near 100%, while taking into account the class distribution with the kappa
metrics the inter agreement between annotator decreases. The values obtained
for reputation polarity in terms of percentage of agreement are quite similar to
other studies over sentiment analysis task. As in the ltering subtask, the value
obtained with kappa in the reputation polarity subtask decrease with respect of
percentage of agreement. For the topic detection subtask, we do not compute
inter agreement between annotators for the whole RepLab 2013 dataset. This
is due to the organization of the labeling process. The annotators consider the
training and test set as two di erent sets, so cannot group tweets of both sets.
The agreement for the topic detection task is higher than expected, taking into
account the complexity of the subtask.</p>
      <p>As expected, the results obtained in the experiment with three annotators are
lower. As can be seen in Table 7, the inter agreement for the ltering task is quite
similar to that obtained in the experiment with two annotators, while the results
for the reputational polarity decrease considerably in all metrics. Concerning
the topic detection subtask, the table shows the average of F measure over
all combinations between annotators. Notably, this task is the one with a lower
decrease with respect to the experiment with two annotators, even if this subtask
depends on the organization behavior of the annotators. Similarly to the previous
experiments of two annotators, as the training and test are considered as two
sets by the annotator, the topic detection inter agreement for the whole RepLab
2013 dataset is not computed. Finally, the values obtained for the priority task
for three annotators decrease more than for topic detection comparing with the
previous experiment, but are still similar.
4</p>
      <p>Participation
44 groups signed up for RepLab 2013, although only 15 of them submitted runs
to the o cial evaluation.6 This year the task was de ned in such a way that using
the baselines provided by the organizers, every group, besides participating in
a concrete subtask, could submit its system to the full task. Nevertheless, only
4 systems explicitly used this possibility.7 Overall, 5 groups participated in the
topic detection subtask, 11 in the reputation polarity classi cation subtask, 14
in the ltering subtask and 4 in the priority assignment subtask. Below we list
the participants and brie y describe the approaches used by each group. Table
8 shows the acronyms and a liations of the research groups that took part in
RepLab 2013.</p>
      <p>CIRGDISCO participated in the ltering subtask. They exploited \context
phrases" found in tweets and Wikipedia disambiguated articles for a particular
entity in an SVM classi er that utilizes features extracted from the Wikipedia
graph structure, i.e. incoming and outcoming links from and to Wikipedia
articles. They used, in addition, features derived from term-speci city and
termcollocation features derived from the Wikipedia article of the analysed entity.
6 One additional group sent their results two days after the deadline, and their runs
are reported here as \uno cial." An asterisk in tables indicates an uno cial result.
7 Daedalus, GAVKTH, SZTE NLP, and UNED ORM.
Daedalus submitted speci c runs for the ltering and polarity subtasks, apart
from the full task. Their approach to the ltering subtask is based on the
use of linguistic processing modules to detect and disambiguate named
entities at several levels. The 4 submitted runs are de ned by a combination of
morphosyntactic-based vs. semantic disambiguation and a case
sensitive/insensitive processing of the tweets. On the other hand, the polarity classi cation uses
a lexicon-based approach to sentiment analysis, improved with a full syntactic
analysis and detection of negation and polarity modi ers, which also provides
the polarity at entity level.</p>
      <p>DIUE applied a supervised Machine Learning (ML) approach for the polarity
classi cation subtask. The Python NLTK has been used for preprocessing,
including le parsing, text analysis and feature extraction. The best run combines
bag-of-words with a set of 18 features related to presence of the polarized term,
negation before the polarized expression, as well as entity reference based on
sentiment lexicons and shallow text analysis.</p>
      <p>GAVKTH used its commercially available system for the ltering and reputation
polarity subtasks. The system, designed for large scale analysis of streaming
text and measuring the public attitude towards targets of interest, has been
used with no adjustment for the speci c subtasks. The basic approach relies on
distributional semantics represented in a semantic space by means of a patented
implementation of the Random Indexing processing framework.
LIA applied a large variety of ML methods mainly based on exploiting tweet
contents to ltering, polarity classi cation, topic detection, and priority
assignment. In several experiments some metadata were added and a fewer number
of runs incorporated external information by using provided links to Wikipedia
and entities' o cial web sites.
NLP&amp;IR GROUP UNED focused on addressing ltering and reputation
polarity classi cation using an IR method. Viewing these two subtasks as the same
problem, i.e. nding the most relevant class to annotate a given tweet, a classical
IR approach was applied, using the tweet content as query against an index with
the models of the classes used to annotate tweets. The classes were modelled by
means of the Kullback Leibler Divergence (KLD), in order to extract their most
representative terminology. For topic detection, instead of a clustering based
technique, this group resorted to Formal Concept Analysis (FCA) to represent
the contents in a lattice structure. Topics were extracted from the lattice using
a FCA concept, stability.
popstar participated in the ltering and reputation polarity classi cation
subtasks. For ltering, these researchers explored di erent learning algorithms
considering a variety of features describing the relationship between an entity and
a tweet, such as text, keyword similarity scores between entities metadata and
tweets, the Freebase entity graph and Wikipedia.</p>
      <p>REINA used classical systems for the similarity matrix and community detection
techniques for topic detection. No distinction was made between languages of
the tweets, doing a uniform lexical analysis of all tweets, applying a simple
sstemmer and removing the words with less than 4 characters. Additionally, the
discarded emoticons were considered as well as hashtags and some entities terms.
The urls shared by two tweets were deemed as another important feature of the
tweets, assuming this is indicative of topic similarity.</p>
      <p>SZTE NLP presented a system to tackle the ltering and reputation polarity
classi cation subtasks using supervised ML techniques. Several Twitter speci c
text preprocessing and features engineering methods were applied. Besides
supervised methods, they also experimented with incorporating clustering
information.</p>
      <p>UAMCLYR adopted Distributional Term Representations (DTR) to tackle the
ltering and reputation polarity classi cation subtasks. Terms were represented
by means of contextual information given by the term co-occurrence statistics.
For topic detection and priority assignment, these researchers explored clustering
and classi cation methods as well as term selection techniques working with two
settings: single tweets and tweets extended with derived posts.</p>
      <p>UNED ORM submitted runs to the full task and all the subtasks testing several
approaches. First, Instance-based learning using Heterogeneity Based Ranking
to combine seven di erent similarity measures was applied to all the subtasks.
The ltering subtask was also tackled by automatically discovering positive and
negative lter keywords, i.e. terms present in a tweet that reliably predict the
relatedness or non-relatedness of the message to the analysed entity. The topic
detection subtask was attempted with three approaches: agglomerative
clustering over Wiki ed tweets, co-occurrence term clustering and an LDA-based model
that uses temporal information. Finally, the polarity subtask was tackled by
generating domain speci c semantic graphs in order to automatically expand the
general purpose lexicon SentiSense.</p>
      <p>UNED-READERS* applied an unsupervised knowledge-based approach to lter
relevant tweets for a given entity. The method exploits a new way of
contextualizing entity names from relatively large collections of texts using probabilistic
signature models, i.e., discrete probability distributions of words lexically related
to the knowledge or topic underlying the set of entities in background text
collections. The contextualization is intended to recover relevant information about
the entity, particularly, lexically related words, from background knowledge.
UNEDTECNALIA submitted a ltering algorithm that takes advantage of the
Web of Data in order to create a context for every entity. The semantic context
of the analysed entities is generated by querying di erent data sources (modelled
by a set of ontologies) provided by the Linked Open Data Cloud. The extracted
context is then compared to the terms contained in the tweet.</p>
      <p>UVA UNED, a collaborative participation of UvA and UNED, focused on
applying an active learning approach to the ltering subtask. It consisted of exploiting
features based on the detected semantics in the tweet (using Entity Linking with
Wikipedia), as well as tweet-inherent features such as hashtags and usernames.
The tweets manually inspected during the active learning process were at most
1% of the test data.
volvam participated in polarity classi cation and applied one supervised and two
unsupervised approaches, combining ML and lexicon-based techniques with an
emotional concept model. These methods had been properly adapted to English
and Spanish depending on the resources available for each language. The rst,
unsupervised, approach made use of fuzzy lexicons in order to catch informal
variants that are common in Twitter texts. The supervised method extended
the rst approach with ML techniques and an emotion concept model, while
the last one also employed ML but incorporating the bag-of-concepts approach
using SenticNet common-sense a ective knowledge.
5
5.1</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Results</title>
      <sec id="sec-4-1">
        <title>Polarity</title>
        <p>Polarity has been evaluated according to Accuracy and R&amp;S. Only entity-related
tweets in the test set have been assessed. In order to keep evaluation independent
from the ltering task, we do not penalize polarity annotations made on
nonrelated tweets. That is, only related tweets are considered in the Accuracy and
R&amp;S computation. The related tweets without system response are penalized.
The system results, sorted by accuracy are shown in Table 9. The table includes
only the best system, according to R&amp;S or Accuracy, for each team. The second
column contains the ratio of tweets for which the output gives results.</p>
        <p>The majority class is the dataset is \POSITIVE". The baseline approach
appears in the middle of the ranking. SZTE and POPSTAR teams improve, in
general, most systems according to both accuracy and R&amp;S. Note that some
systems achieve a low accuracy (below the baseline) but with competitive R&amp;S.
As R&amp;S only looks at the relative ordering between tweets (rather than the</p>
        <p>Fig. 2. Polarity: Accuracy EN vs ES.
actual tags), a possible reason is that, while many tags are not correct, the
ordinal polarity relationship between them is correct. Figure 1 illustrates the
correspondence between Accuracy and R&amp;S. Note that a high R&amp;S tends to be
associated with a high accuracy.</p>
        <p>Another important aspect of polarity detection for ORM is the ability to
predict the average polarity of an entity with respect to other entities. To evaluate
this ability, we have computed the Pearson correlation between the average
estimated and real polarity levels across entities.8 An interesting result is that some
approaches are able to estimate the average polarity reputation for an entity
with a 0.9 correlation with the ground truth.</p>
        <p>Finally, Figure 2 shows the correlation between Accuracy scores over English
versus Spanish tweets. In most cases there is a high correspondence. The accuracy
for Spanish seems to be upper bounded by the accuracy over English tweets.
In this task, tweets must be classi ed as related or unrelated to the entity of
interest. R&amp;S in ltering tasks (two levels) correspond with the products of
8 For the correlation computation, we assign 0, 1 and 2 for each class respectively.
precision in both classes and the product or recall scores respectively. Table
10 shows the Accuracy and R&amp;S results for the ltering task. Again, we have
included only the best run according to Accuracy or R&amp;S for each team. Most
tweets are related (77%). As in the polarity tasks, the baseline approach appears
in the middle of the ranking for both R&amp;S and Accuracy. Figure 3 shows the
correspondence between Accuracy and R&amp;S. As in the polarity task, a high
R&amp;S score ensures a high Accuracy score. As in the polarity task, there are no
important di erences in system scores when considering the Spanish vs. English
tweets. There is a 0.94 Pearson Correlation) between scores over both kind of
tweets. In general, the top scores are much higher than in RepLab 2012; this is
explained by the fact that in this new dataset the training and test entities are
the same.
The Priority task consists of classifying tweets into three levels. Reliability
represents the ratio of correct priority relationships per tweet, while Sensitivity
represents the ratio of captured relationships per tweet. In this case, as well as
in polarity, only the related tweets (according to assessors) are considered in the
evaluation process. Table 11 shows the results. Only the best Accuracy and R&amp;S
score per team is included. Not all systems have annotated all tweets (see the</p>
        <p>
          S
last column). The best run achieves a high score for both R&amp;S and Accuracy
measures. The baseline approach is improved substantially for both measures.
Topic detection is a clustering task which has been evaluated according to R&amp;S,
which correspond with the popular measures Bcubed precision and Recall [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Table 12 displays the results. Only the best F measure is considered for each
team. Figure 4 shows that there is an important trade-o between R and S in
this task. In these circumstances, the F measure weighted with = 0:5 rewards
the runs located in the diagonal axis. But this choice of is, to some extent,
arbitrary. For this reason, we check the evaluation results according to UIR (see
previous section). UIR is a complementary measure that indicates to what extent
run improvements are sensitive to variations in the measure weighting scheme
(i.e. in ). Table 13 shows for all runs, the other runs which are improved by
the rst with UIR 0; 2. This implies that there is a di erence higher than 0.2
between the cases in which the rst run improves the other for R and S and
vice versa. Interestingly, although UAMCLYR 7 is not the best system in the
F =0:5 ranking, it improves robustly a great amount of runs. Some team runs like
LIA are not comparable to each other. Probably, they have di erent grouping
thresholds.
The full task joins ltering, priority and topic detection tasks. The use of R&amp;S
allows us to apply the same evaluation criterion to all subtasks and therefore,
to combine all of them. It is possible to apply R&amp;S directly over the set of
relationships (priority, ltering and clustering) but then the most frequent binary
relationships dominate the evaluation results (in our case, priority relationships
would dominate). We decided to use a weighted harmonic mean (F measure) of
the six Reliability and Sensitivity measures corresponding to the three subtasks
embedded in the full task. In cases of empty partial outputs, we have completed
runs with the baseline approach as speci ed in the guidelines.
        </p>
        <p>Table 14 shows the team ranking in terms of F. However, this evaluation is
highly sensitive to the relative importance of measures in the combining function.
For this reason, we have also computed UIR between each pair of runs. Here we
consider as an unanimous improvement of system A over system B to those
test cases (entities) for which all the six measures are better for A than for B.
Results of the UIR analysis are shown in Table 15. The third and fourth columns
represent how many entities one run improves or is improved by the other. It
only includes those run pairs for which UIR is bigger than 0.2. As the table
shows, actually, runs from di erent teams are not comparable to each other:
improvements in F are dependent on the relative weighting scheme. However,
there are a number of signi cant improvements (in terms of UIR) between runs
from the same teams.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>Perhaps the main outcome of RepLab 2013 is its dataset, which comprises more
than 142,000 tweets in two languages with four types of high-quality manual
annotations, covering all essential aspects of the reputation monitoring process.
We expect this dataset to become a useful resource for researchers not only in the
eld of reputation management, but also for researchers in Information Retrieval
and Natural Language Processing in general. Just to give an example, the topics
(tweet clusters) together with their relative ranking can be directly mapped into
a test collection to evaluate search with diversity algorithms over Twitter.</p>
      <p>Compared to RepLab 2012, availability of training data for the entities in the
test set naturally improves system results and also allows for a more
straightforward application of machine learning techniques. But the tasks themselves
are still far from solved; even with plenty of entity-speci c training material the
RepLab tasks|polarity, topic detection, and ranking|have proved challenging
for state-of-the-art systems.</p>
      <sec id="sec-5-1">
        <title>Acknowledgements</title>
        <p>This research was partially supported by the European Community's FP7
Programme under grant agreement nrs 258191 (PROMISE Network of Excellence)
and 288024 (LiMoSINe), the ESF Research Network Program ELIAS, the
Spanish Ministry of Education (FPU grant AP2009-0507 and FPI grant
BES-2011044328), the Spanish Ministry of Science and Innovation (Holopedia Project,
TIN2010-21128-C02), and the Regional Government of Madrid under
MA2VICMR (S2009/TIC-1542), the Netherlands Organisation for Scienti c Research
(NWO) under project nrs 640.004.802, 727.011.005, 612.001.116, HOR-11-10, the
Center for Creation, Content and Technology (CCCT), the QuaMerdes project
funded by the CLARIN-nl program, the TROVe project funded by the
CLARIAH program, the Dutch national program COMMIT, the Elite Network Shifts
project funded by the Royal Dutch Academy of Sciences (KNAW), the
Netherlands eScience Center under project number 027.012.105 and the Yahoo! Faculty
Research and Engagement Program.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artiles</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A comparison of extrinsic clustering evaluation metrics based on formal constraints</article-title>
          .
          <source>Information Retrieval</source>
          <volume>12</volume>
          (
          <issue>4</issue>
          ),
          <volume>461</volume>
          {
          <fpage>486</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corujo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meij</surname>
          </string-name>
          , E., de Rijke, M.: Overview of RepLab 2012:
          <article-title>Evaluating Online Reputation Management Systems</article-title>
          .
          <source>In: CLEF 2012 Labs and Workshop Notebook</source>
          Papers (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Artiles</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Combining evaluation metrics via the unanimous improvement ratio and its application to clustering tasks</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          <volume>42</volume>
          (
          <issue>1</issue>
          ),
          <volume>689</volume>
          {
          <fpage>718</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verdejo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A General Evaluation Measure for Document Organization Tasks</article-title>
          .
          <source>In: Proceedings of SIGIR</source>
          <year>2013</year>
          (jul
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>