<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of RepLab 2014: Author Pro ling and Reputation Dimensions for Online Reputation Management</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Enrique Amigo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Carrillo-de-Albornoz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Irina Chugur</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adolfo Corujo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julio Gonzalo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edgar Meij</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maarten de Rijke</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Damiano Spina</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Llorente &amp; Cuenca Lagasca</institution>
          ,
          <addr-line>88. 28001 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>UNED NLP &amp; IR Group Juan del Rosal</institution>
          ,
          <addr-line>16. 28040 Madrid, Spain, nlp.uned.es</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Amsterdam Science Park 904</institution>
          ,
          <addr-line>1098 XH Amsterdam</addr-line>
          ,
          <country>The</country>
          <addr-line>Netherlands, ilps.science.uva.nl</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Yahoo Labs Avinguda Diagonal 177</institution>
          ,
          <addr-line>08018 Barcelona, Spain, labs.yahoo.com</addr-line>
        </aff>
      </contrib-group>
      <fpage>1438</fpage>
      <lpage>1457</lpage>
      <abstract>
        <p>This paper describes the organisation and results of RepLab 2014, the third competitive evaluation campaign for Online Reputation Management systems. This year the focus lied on two new tasks: reputation dimensions classi cation and author pro ling, which complement the aspects of reputation analysis studied in the previous campaigns. The participants were asked (1) to classify tweets applying a standard typology of reputation dimensions and (2) categorise Twitter pro les by type of author as well as rank them according to their in uence. New data collections were provided for the development and evaluation of systems that participated in this benchmarking activity.</p>
      </abstract>
      <kwd-group>
        <kwd>RepLab</kwd>
        <kwd>Reputation Management</kwd>
        <kwd>Evaluation Methodologies and Metrics</kwd>
        <kwd>Test Collections</kwd>
        <kwd>Reputation Dimensions</kwd>
        <kwd>Author Pro ling</kwd>
        <kwd>Twitter</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>RepLab is a competitive evaluation exercise supported by the EU project
LiMoSINe.5 It aims at encouraging research on Online Reputation Management and
providing a framework for collaboration between academia and practitioners in
the form of a \living lab": a series of evaluation campaigns in which task design
and evaluation are jointly carried out by researchers and the target user
community (in our case, reputation management experts). Similar to the previous
campaigns [1,2], RepLab 2014 was organized as a CLEF lab.6</p>
      <p>Previous RepLab editions focused on problems such as entity resolution
(resolving name ambiguity), topic detection (what are the issues discussed about
the entity?), polarity for reputation (which statements and opinions have
negative/positive implications for the reputation of the entity?) and alert detection
(which are the issues that might harm the reputation of the entity?). Although
online monitoring pervades all online media (news, social media, blogosphere,
etc.), RepLab has always been focused on Twitter content, as it is the key media
for early detection of potential reputational issues.</p>
      <p>In 2014, RepLab focused on two additional aspects of reputation analysis
{ reputation dimensions classi cation and author pro ling { that complement
the tasks tackled in the previous campaigns. As we will see below, reputation
dimensions contribute to a better understanding of the topic of a tweet or group
of tweets, whilst author pro ling provides important information for priority
ranking of tweets, as certain characteristics of the author can make a tweet (or
a group of tweets) an alert, requiring special attention of reputation experts.
Section 2 explains the tasks in more detail. A description of the data collections
created for RepLab 2014 and chosen evaluation methodology can be found in
Sections 3 and 4, respectively. In Section 5, we brie y review the list of
participants and employed approaches. Section 6 is dedicated to the display and
analysis of the results, based on which we, nally, draw conclusions in Section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Tasks De nition</title>
      <p>
        In 2014, RepLab o ered its participants the following tasks: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) classi cation
of Twitter posts by reputation dimension and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) classi cation and ranking of
Twitter pro les.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Reputation Dimensions Classi cation</title>
        <p>The aim of this task is to assign tweets to one of the seven standard reputation
dimensions of the RepTrak Framework7 developed by the Reputation Institute.
These dimensions re ect the a ective and cognitive perceptions of a company by
di erent stakeholder groups. The task can be viewed as a complement to topic
detection, as it provides a broad classi cation of the aspects of the company
under public scrutiny. Table 1 shows the de nition of each reputation dimension,
supported by an example of a labelled tweet:
7 http://www.reputationinstitute.com/about-reputation-institute/
the-reptrak-framework
The company's acknowledgement of the social and
environmental responsibility, including ethical aspects of business: integrity,
transparency and accountability.</p>
        <p>Find out more about Santander Universities scholarships, grants,
awards and SME Internship Programme bit.ly/1mMl2OX
Related to the relationship between the company and the public
authorities.</p>
        <p>Judge orders Barclays to reveal names of 208 staff linked to Libor
probe via @Telegraph soc.li/mJVPh1R
Related to the working environment and the company's ability to
attract, form and keep talented and highly quali ed people.</p>
        <p>Goldman Sachs exec quits via open letter in The New York Times, brands
bank working environment ``toxic and destructive'' ow.ly/9EaLc
The innovativeness shown by the company, nurturing novel ideas
and incorporating them into products.</p>
        <p>Eddy Merckx Cycles announced a partnership with Lexus to develop their</p>
        <p>ETT Hme trial bike. More info at...http://fb.me/1VAeS3zJP
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Author Pro ling</title>
        <p>This task is composed of two subtasks that were evaluated separately.
Author Categorisation. The task was to classify Twitter pro les by type of
author: Company (i.e., corporate accounts of the company itself ), Professional
(in the economic domain of the company), Celebrity, Employee, Stockholder,
Investor, Journalist, Sportsman, Public Institution, and Non-Governmental
Organisation (NGO). The systems' output was expected to be a list of pro le
identi ers with the assigned categories, one per pro le.</p>
        <p>Author Ranking. Using as input the same set of Twitter pro les as in the task
above, systems had to nd out which authors had more reputational in uence
(who the in uencers or opinion makers are) and which pro les are less in uential
or have no in uence at all. For a given domain (e.g., automotive or banking),
the systems' output was a ranking of pro les according to their probability of
being an opinion maker with respect to the concrete domain, optionally including
the corresponding weights. Note that, because the number of opinion makers is
expected to be low, we modelled the task as a search problem (hence the system
output is a ranked list) rather than as a classi cation problem.</p>
        <p>Some aspects that determine the in uence of an author in Twitter (from
a reputation analysis perspective) can be the number of followers, number of
comments on a domain or type of author. As an example, below is the pro le
description of an in uential nancial journalist:</p>
        <p>Description: New York Times Columnist &amp; CNBC Squawk Box
(@SquawkCNBC) Co-Anchor. Author, Too Big To Fail. Founder,
@DealBook. Proud father. RTs endorsements
Location: New York, New York nytimes.com/dealbook
Tweets: 1,423</p>
      </sec>
      <sec id="sec-2-3">
        <title>Tweet examples:</title>
        <p>\Whitney Tilson: Evaluating the Dearth of Female Hedge Fund Managers
http://nyti.ms/1gpClRq @dealbook"
\Dina Powell, Goldman's Charitable Foundation Chief to Lead the Firm's
Urban Investment Group http://nyti.ms/1fpdTxn @dealbook"
Shared PAN-RepLab Author Pro ling: Participants were also o ered the
opportunity to attempt the shared author pro ling task RepLab@PAN.8 In
order to do so, systems had to classify Twitter pro les by gender and age. Two
categories, female and male, were used for gender. Regarding age, the following
classes were considered: 18-24, 25-34, 35-49, 50-64, and 65+.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data Sets</title>
      <p>This section brie y describes the data collections used in each task. Note that
the current amount of available tweets may be lower, as some posts may have
been deleted or made private by the authors: in order to respect the Twitter's
terms of service (TOS), we did not provide the contents of the tweets, but only
tweet ids and screen names. Tweet texts can be downloaded using any of the
following tools:
1. TREC Microblog Track9
2. SemEval-2013 Task 2 Download script10
3. A Java tool provided by the RepLab organisers11
8 http://pan.webis.de/
9 https://github.com/lintool/twitter-tools
10 http://www.cs.york.ac.uk/semeval-2013/task2/index.php?id=data
11 http://nlp.uned.es/replab2013/replab2013_twitter_texts_downloader_
latest.tar.gz
3.1</p>
      <sec id="sec-3-1">
        <title>Reputation Dimensions Classi cation Data Set</title>
        <p>This data collection is based on the RepLab 2013 corpus12 and contains over
48,000 manually labelled English and Spanish tweets related to 31 entities from
the automotive and banking domains. The tweets were crawled from the 1st
June 2012 to the 31st Dec 2012 using the entity's canonical name as query. The
balance between languages depends on the availability of data for each entity.
The distribution between the training and test sets was established as follows.
The training set was composed of 15,562 Twitter posts and 32,446 tweets were
reserved for the test set. Both data sets were manually labelled by annotators
trained and supervised by experts in Online Reputation Management from the
online division of a leading Public Relations consultancy Llorente &amp; Cuenca.13</p>
        <p>The tweets were classi ed according to the RepTrak dimensions14 listed in
Section 2. In case a tweet could not be categorised into any of these dimensions,
it was labelled as \Unde ned".</p>
        <p>The reputation dimensions corpus also comprises additional background tweets
for each entity (up to 50,000, with a large variability across entities). These are
the remaining tweets temporally situated between the training (earlier tweets)
and test material (the latest tweets) in the timeline.</p>
        <p>Figure 1 shows the distribution of the reputation dimensions in the training
and test sets, and in the whole collection. As can be seen, the Products &amp; Services
dimension is the majority class in both data sets, followed by the Citizenship
and Governance. The large number of tweets associated with the Unde ned
dimension in both sets is noteworthy, which suggests the complexity of the task,
as even human annotators could not specify the category of 6,577 tweets.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Author Pro ling Data Set</title>
        <p>This data collection contains over 7,000 Twitter pro les (all with at least 1,000
followers) that represent the automotive, banking and miscellaneous domains.
The latter includes pro les from di erent domains. The idea of this extra set is
to evaluate if approaches designed for a speci c domain are suitable for a broader
multi-domain scenario. Each pro le contains (i) screen name; (ii) pro le URL,
and (iii) the last 600 tweets published by the author at crawling time.</p>
        <p>The collection was split into training and test sets: 2,500 pro les in the
training set and 4,991 pro les in the test set. Reputation experts performed manual
annotations for two subtasks: Author Categorisation and Author Ranking. First,
they categorised pro les as company (i.e., corporate accounts of companies),
professional, celebrity, employee, stockholder, journalist, investor, sportsman, public
institution, and non-governmental organisation (NGO). In addition, reputation
experts manually identi ed the opinion makers (i.e., authors with reputational
in uence) and annotated them as \In uencer". The pro les that were not
considered opinion makers were assigned the \Non-In uencer" label. Those pro les
that could not be classi ed into one of these categories, were labelled as
\Undecidable".</p>
        <p>The distribution by classes in the Author Categorisation data collection is
shown in Figure 2. As can be seen, Professional and Journalist are the
majority classes in both training and test sets, followed by the Sportsman, Celebrity,
Company and NGO. Surprisingly, the number of authors in the categories
Stockholder, Investor and Employee is considerably low. One possible explanation is
that such authors are not very active on Twitter, and more specialized forums
need to be considered in order to monitor these types of users.</p>
        <p>Regarding the distribution of classes in the Author Ranking dataset, Table
?? shows the number of authors labelled as In uencer and Non-In uencer in
the training and test sets. The proportion of in uencers is much higher than
we expected, and calls for a revision of our decision to cast the problem as
search ( nd the in uentials) rather than classi cation (classify as in uential or
non-in uential).
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Shared PAN-RepLab Author Pro ling Data Set</title>
        <p>For the shared PAN-RepLab author pro ling task, 159 Twitter pro les from
several domains were annotated with gender (female and male) and age
(1824, 25-34, 35-49, 50-64, and 65+). The pro les were selected from the RepLab
2013 test collection and from a list of in uential authors provided by Llorente
&amp; Cuenca.</p>
        <p>
          131 pro les were included into the miscellaneous data set of the RepLab
author pro ling data collection accompanied by the last 600 tweets published by
the authors at crawling time. 28 users had to be discarded because more than
50% of their tweets were written in languages other than English or Spanish.
The selected 131 pro les, in addition to age and gender, were manually tagged
by reputation experts as explained in Section 3.2: with (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) type of author and
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) opinion-maker labels.
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Methodology</title>
      <sec id="sec-4-1">
        <title>Baselines</title>
        <p>For both classi cation tasks | Reputation Dimensions and Author
Categorisation | a simple Bag-of-Words (BoW) classi er was proposed as o cial baseline.
We used Support Vector Machines,15 with a linear kernel. The penalty
parameter C was automatically adjusted by weights inversely proportional to class
frequencies in the training data. We used the default values for the rest of
parameters.</p>
        <p>For the Reputation Dimensions task, a di erent multi-class tweet classi er
was built for each entity. Tweets were represented as BoW with binary
occurrence (1 if the word is present in the tweet, 0 if not). The BoW representation was
generated by removing punctuation, lowercasing, tokenizing by white spaces,
reducing multiple repetitions of characters (from n to 2) and removing stopwords.</p>
        <p>For the Author Categorisation task, a di erent classi er was built for each
domain in the training set (i.e., banking and automotive). Here, each Twitter
pro le was represented by the latest 600 tweets provided with the collection.
Then, the built pseudo-documents were preprocessed as described before.</p>
        <p>Finally, the number of followers of each Twitter pro le has been used as
baseline for the Author Ranking task.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation Measures</title>
        <p>Reputation Dimensions Categorisation This task is a multi-class
classication problem and its evaluation is an open issue. The traditional Accuracy
measure presents drawbacks for unbalanced data. On the other hand, the
commonly used F-measure over each of the classes does not allow to produce a global
system ranking. In this evaluation campaign we chose Accuracy as the o cial
measure for the sake of interpretability. It is worth mentioning that, in the
Reputation Dimensions task, systems outperformed a most-frequent baseline which
always selects the majority class labels (see Section 6.1).</p>
        <p>Author Categorisation Similar to the Reputation Dimensions, the rst
subtask of Author Pro ling is a categorization task. We also used Accuracy as the
o cial evaluation measure. However, the obtained empirical results suggest that
Accuracy is not able to discriminate system outputs from the majority class
baseline. For this reason, the results were complemented with Macro Average
Accuracy (MAAC ), which penalizes non-informative runs.</p>
        <p>Author Ranking The second subtask of Author Pro ling is a ranking problem.
In uential authors must be located at the top of the system output ranking.
This is actually a traditional information retrieval problem, where relevant and
irrelevant classes are not balanced. Studies on information retrieval measures
can be applied in this context, although author pro ling di ers from information
retrieval in a number of aspects. The main di erence (which is a post-annotation
nding) is that the ratio of relevant authors is much higher than the typical ratio
of relevant documents in a traditional information retrieval scenario.
15 http://scikit-learn.org/stable/modules/svm.html</p>
        <p>Another di erentiating characteristic is that the set of potentially in
uential authors is rather small, while information retrieval test sets usually consist
of millions of documents. This has an important implication for the evaluation
methodology. All information retrieval measures state a weighting scheme which
re ects the probability of users to explore a deepness level in the system's output
ranking. In the Online Reputation Management scenario, this deepness level is
still not known. We decided to use MAP (Mean Average Precision) for two
reasons. First, because it is a well-known measure in information retrieval. Second,
because it is recall-oriented and also considers the relevance of authors at lower
ranks.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Participation</title>
      <p>49 groups signed in for RepLab 2014, although only 11 of them (from 9 di erent
countries) nally submitted results in time for the o cial evaluation. Overall, 8
groups participated in the Reputation Dimensions task, and 5 groups submitted
their results to Author Pro ling (all of them submitted to the author ranking
subtask, and all but one to the author categorization subtask).</p>
      <p>Table 3 shows the acronyms and a liations of the research groups that
participated in RepLab 2014. In what follows, we list the participants and brie y
describe the approaches they used.</p>
      <p>Country
Ireland
Spain
France
Spain
Sweden
Mexico
Iran
The Netherlands
Switzerland
CIRGIRDISCO participated in the Reputation Dimensions task. They used
dominant Wikipedia categories related to a reputation dimension in a Random
Forest classi er. Additionally, they also applied tweet-speci c, language-speci c
and similarity-based features. The best run signi cantly improved over the
baseline accuracy.</p>
      <p>DAE attempted the Reputation Dimensions Classi cation. Their initial idea
was to evaluate the best combination strategy of a machine learning classi er
with a rule-based algorithm that uses logical expressions of terms. However, the
baseline experiment employing just Naive Bayes Multinomial with a term vector
model representation of the tweet text was ranked second among runs from all
participants in terms of Accuracy.</p>
      <p>LIA carried out a considerable number of experiments for each task. The
proposed approaches rely on a large variety of machine learning methods. The main
accent was put on exploiting tweet contents. Several methods also included
selected metadata. Marginally, external information was considered by using
provided background messages.</p>
      <p>LyS attempted all the tasks. For Dimensions Classi cation and Author
Categorisation a supervised classi er was employed with di erent models for each
task and each language. A NLP perspective was adopted, including
preprocessing, PoS tagging and dependency parsing, relying on them to extract features
for the classi er. For author ranking, their best performance was obtained by
training a bag-of-words classi er fed with features based on the Twitter pro le
description of the users.</p>
      <p>ORM UNED proposed a learning system based on voting model for the
Author Pro ling task. They used a small set of features based on the information
that can be found in the text of tweets: POS tags, number of hashtags or number
of links.</p>
      <p>SIBtex integrated several tools into a complete system for tweet monitoring
and categorisation which uses instance-based learning (K-Nearest Neighbours).
Dealing with the domain (automotive or banking) and the language (English
or Spanish), their experiments showed that even with all data merged into one
single Knowledge Base (KB), the observed performances were close to those with
dedicated KBs. Furthermore, English training data in addition to the sparse
Spanish data were useful for Spanish categorisation.</p>
      <p>STAVICTA devised an approach based on the textual content of tweets
without considering metadata and the content of URLs for the reputation
dimensions classi cation. They experimented with di erent feature sets including bag
of n-grams, distributional semantics features, and deep neural network
representations. The best results were obtained with bag of bi-gram features with
minimum frequency thresholding. Their experiments also show that semi-supervised
recursive auto-encoders outperform other feature sets used in the experiments.
UAMCLYR participated in the Author Pro ling task. For Author
Categorisation they used a supervised approach based on the information found in Twitter
users' pro les. Employing attribute selection techniques, the most representative
attributes from each user's activity domain were extracted. For Author Ranking
they developed a two-step chained method based on stylistics attributes (e.g.,
lexical richness, language complexity) and behavioural attributes (e.g., posts'
frequency, directed tweets) obtained from the users' pro les and posts. These
attributes were used in conjunction with a Markov Random Fields to improve
an initial ranking given by the con dence of Support Vector Machine classi
cation algorithm.
uogTr investigated two approaches to the Reputation Dimensions classi cation.
Firstly, they used a term's Gini-index score to quantify the term's
representativeness of a speci c class and constructed class pro les for tweet classi cation.
Secondly, they performed tweet enrichment using a web scale corpus to derive
terms representative of a tweet's class, before training a classi er with the
enriched tweets. The tweet enrichment approach proved to be e ective for this
classi cation task.</p>
      <p>UTDBRG participated in the Author Ranking subtask. The presented system
utilizes a Time-sensitive Voting algorithm. The underlying hypothesis is that
in uential authors tweet actively about hot topics. A set of topics was extracted
for each domain of tweets and a time-sensitive voting algorithm was used to rank
authors in each domain based on the topics.</p>
      <p>UvA addressed the Reputation Dimensions task by using corpus-based methods
to extract textual features from the labelled training data to train two
classiers in a supervised way. Three sampling strategies were explored for selecting
training examples. All submitted runs outperformed the baseline, proving that
elaborate feature selection methods combined with balanced datasets help
improve classi cation performance.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Evaluation Results</title>
      <p>This section reports and analyses the results of the RepLab 2014 tasks, except
for the shared PAN-RepLab author pro ling, for which no submissions were
received.
6.1</p>
      <sec id="sec-6-1">
        <title>Reputation Dimensions Classi cation</title>
        <p>Eight groups participated in the Reputation Dimensions task. 31 runs were
submitted. Most approaches employed di erent machine learning algorithms such
as Support Vector Machine (UvA, uogTr), Random Forest (CIRGIRDISCO,
uogTr), Naive Bayes (DAE, UvA, STAVICTA), distance to class vectors (uogTr),
LibLinear (LyS). SIBtex focussed on instance based learning techniques.</p>
        <p>Regarding the employed features, some approaches considered information
beyond the tweet textual content. For instance, uogTr expanded tweets with
pseudo-relevant document sets and Wikipedia entries, CIRGIRDISCO employed
Wikipedia categories, LyS considered psychometric dimensions and linguistic
information such as dependency trees and part of speech. STAVICTA applied
Distributional Semantic Models to expand tweets.</p>
        <p>Table 4 shows the nal ranking for the Reputation Dimensions task in terms
of Accuracy. The last column represents the ratio of classi ed tweets from the
set of tweets that were available at the time of evaluation. Note that tweets
manually tagged as \Unde ned" were excluded from the evaluation and tweets
tagged by systems as \Unde ned" were considered as non-processed.</p>
        <p>Accuracy Ratio of processed tweets
Run
0.99
0.96
0.91
0.95
0.95
0.95
0.89
0.98
0.92
0.88
0.94
0.99
0.89
0.95
0.86
0.91
0.96
0.91
0.95
0.94
0.86</p>
        <p>1
0.96</p>
        <p>1
0.98
0.94
0.98
0.82
0.82
0.91</p>
        <p>1
0.99
1</p>
        <p>Besides participant systems, we included a baseline that employs Machine
Learning (SVM) using words as features. Note that classifying every tweet as the
most frequent class (majority class baseline) would get an accuracy of 56%. Most
runs are above this threshold and provide, therefore, some useful information
beyond a non-informative run.</p>
        <p>There is no clear correspondence between performance and algorithms. The
top systems used a variety of methods such as a basic Naive Bayes approach
(DAE RD 1), enrichment with pseudo-relevant documents (uogTR RD 4) or
multiple features including dependency relationships, POS tags, and psycometric
dimensions (Lys RD 1).</p>
        <p>Given that tweets labelled as \Unde ned" in the gold standard were not
considered for evaluation purposes, systems that tagged tweets as \Unde ned"
had a negative impact on their performance. In order to check to what extent
this a ects the evaluation results, we computed Accuracy without considering
this label. The leftmost graph in Figure 3 shows that there is a high
correlation between both evaluation results across single runs. Moreover, replacing the
\Undecidable" labels by \Product and Services" (majority class) also produces
similar results (see rightmost graph in Figure 3).</p>
        <p>Figure 4 illustrates the distribution of classes across the systems annotations
and the goldstandard. As the gure shows, most of the systems tend to assign the
majority class \Products and services" to a greater extent than the goldstandard.
6.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Author Categorisation</title>
        <p>Four groups participated in this task providing 10 o cial runs. Most of the
runs are based on some kind of Machine Learning method over Twitter pro les.
For instance, LIA employed Hiden Markov Models, Cosine distances with
TFIDF and Gini purity criteria, as well as Poisson modelling. UAMCLYR and LyS
applied Support Vector Machine, and LyS used a combination of four algorithms:
ZeroR, Random Tree, Random Forest and Naive Bayes.</p>
        <p>As for features, the proposal of LyS includes term expansion with
WordNet. ORM UNED considered di erent metadata (e.g., pro le domain, number
of mentions, hashtags), and LyS included psychometric properties related to
psychological dimensions (e.g., anger, happiness) and to topics such as money,
sports, or religion.</p>
        <p>Automotive Banking Miscellaneous Average (Aut.&amp;Bank.)</p>
        <p>Table 5 shows the ranking for the Author Categorisation task. Two uno cial
runs (submitted shortly after the deadline) are marked with an asterisk (*). The
Run
LIA AC 1
Baseline-SVM
Most frequent
UAMCLYR AC 2
UAMCLYR AC 1
ORM UNED AC 1
UAMCLYR AC 3*
ORM UNED AC 3
UAMCLYR AC 4*
LIA AC 2
ORM UNED AC 2
LIA AC 3
LyS AC 1
LyS AC 2
0.46</p>
        <p>0.51
0.39
0.42
0.39
0.22
0.18
0.19
0.38
0.3
0.37
0.25
0.22
0.47
0.46
0.44
0.41
0.4
0.39
0.39
0.39
0.39
0.38
0.37
0.3
0.15
0.13</p>
        <p>Fig. 5: Accuracy and MAAC for the Author Categorisation task.
Accuracy values were computed separately for each domain (automotive, banking
and miscellaneous). We included two baselines: Machine Learning (SVM) using
words as features, and a baseline that assigns the most frequent class (in the
training set) to all authors. Average Accuracy of the banking and automotive
domains was used to rank systems.</p>
        <p>Interestingly, there is a high correlation between system scores in the
automotive vs. banking domains (0.97 Pearson coe cient ). The low Accuracy values
in the case of LyS are due to the fact that more than half of the authors were
not included in the output le.</p>
        <p>The most relevant aspect of these results is that, in terms of Accuracy,
assigning the majority class outperforms most runs, although, of course, this output
is not informative. The question, then, is how much information the systems are
able to produce. In order to answer this question we have computed the Macro
Average Accuracy (MAAC ), which has the characteristic of assigning the same
low score to any non informative classi er (e.g., random classi cation or one
label for all instances). Figure 5 shows that most systems are able to improve
the majority class baseline according to MAAC. This means that systems are
able to produce information about classes, although they reduce the number of
accurate decisions with respect to the majority class baseline.</p>
        <p>From the grouping point of view, the majority class baseline relates all author
to each other in the same class. On the other hand, systems try to identify more
classes, increasing the correctness of grouping relationships at the cost of losing
relationships. In order to analyse this aspect, we also calculated Reliability and
Sensitivity of author relationships (Bcubed Precision and Recall ), as if it was a
clustering problem. Reliaiblity re ects the correctness of grouping relationships
between authors. Sensitivity shows how many of these relationships are captured
by the system.</p>
        <p>The graphs in Figure 6 show the relationship between grouping precision and
recall (R and S ). The majority class baseline achieves the maximum Sensitivity,
given that all authors are assigned to one class and, therefore, all relationships
are captured, but including noisy relationships. As the graphs show, in general,
systems are able to increase slightly the correctness of the produced relationships
(Reliability increase), but at the cost of losing relationships (lower Sensitivity ).
Five groups participated in this task, for a total of 14 runs. The author in uence
estimation is grounded on di erent hypotheses. The approach proposed by LIA
assumes that in uencers tend to produce more opinionated terms in tweets.
UTDBRG assumed that in uential authors tweet more about hot topics. This
requires a topic retrieval step and a time sensitive voting algorithm to rank
authors. Some participants trained their systems over the biography text (LyS,
UAMCLYR), binary pro le metadata such as the presence of URLs, veri ed
account, user image (LyS), quantitative pro le metadata such as the number
of followers (LyS, UAMCLYR ), style-behaviour features such as the number of
URLs, hashtags, favourites, retweets etc. (UAMCLYR).</p>
        <p>Table 6 shows the results for the Author Ranking task produced with the
TREC EVAL tool. In the table, systems are ordered according to the average
MAP between the automotive and banking domains. Unfortunately, some
participants returned their results in the gold standard format (binary classi cation
as in uencers or non in uencers) instead of using the prescribed ranking format.
We did not discard those submissions and turned their results into the o cial
format by locating pro les marked as in uencers at the top, otherwise respecting
the original list order.</p>
        <p>The followers baseline simply ranks the authors by descending number of
followers. It is clearly outperformed by most runs, indicating that additional
signals provide useful information. The exception is the miscellaneous domain,
where probably additional requirements over the number of followers, such as
expertise in a given area, do not clearly apply.</p>
        <p>On the other hand, runs from three participants exceeded 0.5 MAP, using
very di erent approaches. Therefore, current results do not clearly point to one
particular technique.</p>
        <p>Figure 7 shows the correlation between the MAP values achieved by the
systems in the automotive vs. banking domains. There seems to be little
correspondence between results in both domains, suggesting that the performance of
systems is highly biased by the domain. For future work, it is probably necessary
to consider multiple domains to extract robust conclusions.</p>
        <p>Figures 8, 9 and 10 illustrate the precision recall curves in the automotive,
banking and miscellaneous data sets respectively. We tried to group systems in
three levels (black, grey and discontinuous lines) according to their performance.
The baseline approach based on followers is represented by the thick dashed line.</p>
        <p>Systems improve the followers based baseline in both the automotive and
banking domains in all recall levels. This suggests that the number of followers
is not the most determinant feature even for the most followed authors. However,</p>
        <p>Fig. 7: Correlation of MAP values: Automotive vs. Banking.
this is not the case of the miscellaneous data set, in which the author compilation
were biased to high popular writers.
After two evaluation campaigns on core Online Reputation Management tasks
(name ambiguity resolution, reputation polarity, topic and alert detection),
Rep</p>
        <p>
          Fig. 9: Precision/Recall curves for Author Ranking in the banking domain.
Lab 2014 developed an evaluation methodology and test collections for two
different reputation management problems: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) classi cation of tweets according to
the reputation dimensions, and (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) identi cation and categorisation of opinion
makers. Once more, the manual annotations were provided by reputation experts
from Llorente &amp; Cuenca (48,000 tweets and 7,000 author pro les annotated).
        </p>
        <p>Being the rst shared evaluation on these tasks, participants explored a wide
range of approaches in each of them. The classi cation of tweets according to
their reputation dimensions seems to be feasible, although it is not yet clear
which are the best signals and techniques to optimally solve it. Author
categorisation, on the other hand, proved to be challenging in this initial approximation.</p>
        <p>Current results represent simply a rst attempt to understand and solve the
tasks. Nevertheless, we expect that the data set we are releasing will allow for
further experimentation and for a substantial improvement of the state of the
art in the near future, as has been the case with the RepLab 2012 and RepLab
2013 data sets.</p>
        <p>Acknowledgements. This research was partially supported by the European
Community's Seventh Framework Programme (FP7/2007-2013) under grant
agreements nr 288024 (LiMoSINe) and nr 312827 (VOX-Pol), ESF grant ELIAS,
the Spanish Ministry of Education (FPU grant AP2009-0507), the Spanish
Ministry of Science and Innovation (Holopedia Project, TIN2010-21128-C02), the
Regional Government of Madrid under MA2VICMR (S2009/TIC-1542), Google
Award (Axiometrics), the Netherlands Organisation for Scienti c Research
(NWO) under project nrs 727.011.005, 612.001.116, HOR-11-10, 640.006.013,
the Center for Creation, Content and Technology (CCCT), the QuaMerdes
project funded by the CLARIN-nl program, the TROVe project funded by the
CLARIAH program, the Dutch national program COMMIT, the ESF Research
Network Program ELIAS, the Elite Network Shifts project funded by the Royal
Dutch Academy of Sciences (KNAW), the Netherlands eScience Center under
project number 027.012.105, the Yahoo! Faculty Research and Engagement
Program, the Microsoft Research PhD program, and the HPC Fund.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrillo-de-Albornoz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chugur</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corujo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mart n</surname>
          </string-name>
          , T.,
          <string-name>
            <surname>Meij</surname>
            , E., de Rijke,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spina</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Overview of RepLab 2013:
          <article-title>Evaluating Online Reputation Management Systems</article-title>
          .
          <source>In: CLEF 2013 Working Notes (Sep</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Amigo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corujo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meij</surname>
          </string-name>
          , E., de Rijke, M.: Overview of RepLab 2012:
          <article-title>Evaluating Online Reputation Management Systems</article-title>
          .
          <source>In: CLEF 2012 Labs and Workshop Notebook</source>
          Papers (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>