<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Feature Selection and Data Sampling Methods for Learning Reputation Dimensions</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Cristina Gaˆrbacea</institution>
          ,
          <addr-line>Manos Tsagkias, and Maarten de Rijke</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>G.C.Garbacea</institution>
          ,
          <addr-line>E.Tsagkias, deRijke</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Amsterdam</institution>
          ,
          <addr-line>Amsterdam</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>1479</fpage>
      <lpage>1490</lpage>
      <abstract>
        <p>We report on our participation in the reputation dimension task of the CLEF RepLab 2014 evaluation initiative, i.e., to classify social media updates into eight predefined categories. We address the task by using corpus-based methods to extract textual features from the labeled training data to train two classifiers in a supervised way. We explore three sampling strategies for selecting training examples, and probe their effect on classification performance. We find that all our submitted runs outperform the baseline, and that elaborate feature selection methods coupled with balanced datasets help improve classification accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Products &amp; Services Related to the products and services offered by the company and reflecting
customers’ satisfaction.</p>
      <p>Innovation The innovativeness displayed by the company, nurturing novel ideas and
incorporating these ideas into products.</p>
      <p>Workplace Related to the employees’ satisfaction and the company’s ability to attract,
form and keep talented and highly qualified people.</p>
      <p>Governance Capturing the relationship between the company and the public
authorities.</p>
      <p>Citizenship The company’s acknowledgement of community and environmental
responsibility, including ethic aspects of the business: integrity, transparency
and accountability.</p>
      <p>Leadership Related to the leading position of the company.</p>
      <p>Performance Focusing on long term business success and financial soundness.
Undefined In case a tweet cannot be classified into none of the above dimensions, it
is labelled as “Undefined.”
as she has been successful in moving forward into a. . . http://fb.me/18FKDLQIr”
belongs to the “Workplace” reputation dimension, while “HSBC to upgrade 10,000 POS
terminals for contactless payments. . . http://bit.ly/K9h6 QW” is related to “Innovation.”</p>
      <p>The author profiling task aims at profiling Twitter users with respect to their
domain of expertise and influence for identifying the most influential opinion makers in a
particular domain of expertise. The task is further divided into two subtasks: (i) author
categorization, and (ii) author ranking. The first subtask aims at the classification of
Twitter profiles according to the type of author, i.e., journalist, professional, authority,
activist, investor, company or celebrity. The second subtask aims at identifying user
profiles with the biggest influence on a company’s reputation.</p>
      <p>We focus on the reputation dimensions task. Our main research question is how
we can use machine learning to extract and select discriminative features that can help
us learn to classify the reputation dimension of a tweet. In our approach we exploit
corpus-based methods to extract textual features that we use for training a Support
Vector Machine (SVM) and a Naive Bayes (NB) classifier in a supervised way. For
training the classifiers we use the provided annotated tweets in the training set and
explore three strategies for sampling training examples: (i) we use all training examples
for all classes, (ii) we downsample classes to match the size of the smallest class, (iii) we
oversample classes to match the size of the largest class. Our results show that our runs
consistently outperform the baseline, and demonstrate that elaborate feature extraction
and oversampling the training data peak classification accuracy at 0.6704.</p>
      <p>The rest of paper is organized as follows. In Section 2 we present related work, in
Section 3 we introduce our feature extraction approach, in Section 5 we describe our
experimental setup, in Section 6 we report on our results. We follow up with an error
analysis and reflections in Section 7 and conclude in Section 8.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The field of Online Reputation Management (ORM) is concerned with the development
of automatic ways for tracking online content that can impact the reputation of a
company. This involves non-trivial aspects from natural language processing, opinion
mining, sentiment analysis and topic detection. Generally, opinions expressed about
individuals or organizations cannot be structured around a predefined set of features/aspects.
Entities require complex modeling, which is a less understood process, and this turns
ORM into a challenging field of research and study.</p>
      <p>
        The RepLab campaigns address the task of detecting the reputation of entities on
social media (Twitter). Each year there are new tasks defined by the organizers.
Replab 2012 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] focused on profiling, that is filtering the stream of tweets for detecting
those microblog posts which are related to a company and their implications on the
brand’s image, and monitoring, i.e., topical clustering of tweets for identifying topics
that harm a company’s reputation and therefore, require the immediate attention of
reputation management experts. Replab 2013 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] built upon the previously defined tasks
and proposed a full reputation monitoring system consisting of four individual tasks.
First, the filtering task asked systems to detect which tweets are related to an
organization by taking entity name disambiguation into account. Second, the polarity detection
for reputation classification task, required systems to decide on whether the content of
a social media update has positive, neutral or negative implications for the company’s
reputation. Third, the topic detection task aimed at grouping together tweets that are
about the same topic. Four, the priority assignment task aimed at ranking the previous
topics based on their potential for triggering a reputation alert.
      </p>
      <p>
        Replab proposes an evaluation test bed made up of multilingual tweets in English
and Spanish with human annotated data for a significant number of entities. The best
systems from previous years addressed the majority of the above presented tasks as
classification tasks by the use of conventional machine learning techniques, and focused on
the extraction of features that encapsulate the main characteristics of a specific
reputation related class. For the filtering and polarity detection tasks, Hangya and Farkas [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
reduce the size of the vocabulary following an elaborate sequence of data
preprocessing steps and create an n-gram based supervised model, which was found previously
successful on short messages like tweets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Graph-based semantic approaches for
assembling domain specific affective lexicon seem not to yield very accurate results given
the inherent short and noisy content of social media updates [
        <xref ref-type="bibr" rid="ref19">20</xref>
        ]. The utility of topic
modeling algorithms and unsupervised techniques based on clustering are explored in
[
        <xref ref-type="bibr" rid="ref5 ref7">7,5</xref>
        ], both addressing the topic detection task. Peetz et al. [
        <xref ref-type="bibr" rid="ref14">15</xref>
        ] show how active learning
can maximize performance for entity name disambiguation by systematically
interacting with the user and updating the classification model.
      </p>
      <p>The reputation dimensions task stems from the hypothesis that customer satisfaction
is easier to measure and manage when we understand the key drivers of reputation that
actively influence a company’s success. The Reptrak system was designed to identify
these drivers by evaluating how corporate reputation emerges from the emotional
connection that an organization develops with its stakeholders. In this scenario, reputation
is measured on a scale from 0–100 and considers the degree of admiration, trust, good
feeling and overall esteem investors display about the organization. Reptrak defines
seven key aspects that define reputation and the reputation dimensions task uses them
to define the reputation dimensions that we listed in Table 1 except the “Undefined”
category, which is an extra class local to the reputation dimensions task.</p>
      <p>
        In our study we follow [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and in particular Gaˆrbacea et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], who presented a
highly accurate system on the task of predicting reputation polarity. We build on their
approaches but we focus on the task of reputation dimensions and we also explore the
effect of balanced and unbalanced training sets on classification accuracy.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Feature Engineering</title>
      <p>
        Classifying tweets by machine learning techniques imposes the need to represent each
document as a set of features based on the presence, absence or frequency of terms
occurring inside the text. Frequency distribution, tf.idf or 2 calculations are common
approaches in this respect. In addition, identifying the semantic relations between
features can capture the linguistic differences across corpora [
        <xref ref-type="bibr" rid="ref20">21</xref>
        ].
      </p>
      <p>
        In our approach we consider textual features that we extract using corpus-based
methods for frequency profiling. We build on the assumption that more elaborate
feature extraction methods can help us identify discriminative features relevant for
characterizing a reputation dimension class. We hypothesize that frequency profiling using the
log-likelihood ratio method (LLR) [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ], which is readily used to identify discriminative
features between corpora, can also yield discriminative features specific to each of our
reputation dimension classes. We extract unigrams and bigrams (the latter because they
can better capture the context of a term) from our training data after having it split into
eight annotated reputation dimensions, each corresponding to one of the given labels.
Our procedure for extracting textual features is described in what follows.
      </p>
      <p>Given two corpora we want to compare, a word frequency list is first produced
for each corpus. Although here a comparison at word level is intended, part of speech
(POS) or semantic tag frequency lists are also common. The log-likelihood statistic is
performed by constructing a contingency table that captures the frequency of a term
as compared to the frequency of other terms inside two distinct corpora. We build our
first corpus out of all the annotated tweets for our target class and our second corpus
out of all the tweets found in the rest of the reputation dimension classes. For
example, for finding discriminative terms for the class “Products &amp; Services,” we compare
pairs of corpora of the form: “Products &amp; Services” vs. “Innovation” and “Workplace”
and “Governance” and “Citizenship” and “Leadership” and “Performance” and
“Undefined.” We repeat this process for each of the eight classes and rank terms by their LLR
score in descending order. We only keep terms that have higher frequency in the target
class than for all the rest of the classes. This results in using as features only terms
expected to be highly discriminative for our target class.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Strategies for Sampling Training Data</title>
      <p>
        Machine learning methods are sensitive to the class distribution in the training set; this
is a well described issue [
        <xref ref-type="bibr" rid="ref21">22</xref>
        ]. Some of RepLab’s datasets, such as the one used for
detecting reputation polarity, have different distributions among classes and between
training and test set. These differences can potentially impact the classification
effectiveness of a system. To this extent, we are interested in finding out what the effect
is of balancing the training set of a classifier on its classification accuracy. Below, we
describe three strategies for sampling training data.
      </p>
      <p>Unbalanced strategy. This strategy uses the original class distribution in the training
data, and it uses all of the training data.</p>
      <p>Downsampling. This strategy downsamples the training examples of each class to
match the size of the smallest class. Training examples are removed at random. We
evaluate the system using ten fold cross validation on the training data, and we repeat
the process ten times. We select the model with the highest accuracy.</p>
      <p>Oversampling. This strategy oversamples the training examples of each class to match
the size of the largest class. For each class, training examples are selected at random
and are duplicated. Similarly as before, we evaluate the system using ten fold cross
validation on the training data, we repeat the process ten times, and we select the model
with the highest accuracy.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experimental Setup</title>
      <p>We conduct classification experiments to assess the discriminative power of our
features for detecting the reputation dimension of a tweet. We are particularly interested
in knowing the effectiveness of our extracted textual LLR features for each of the eight
reputation dimension classes, and the effect of the three sampling strategies for selecting
training data.</p>
      <p>We submitted a total of 5 supervised systems where we probe the usefulness of
machine learning algorithms for the current task. We list our runs in Table 2. We train our
classifiers regardless of any association with a given entity, since there are cases in the
training data when not all classes are present for an entity (see Table 3). In UvA RD 1
we choose to train an SVM classifier using all the available training tweets for each
reputation dimension class, which implies that our classes are very unbalanced at this
stage. In our next runs, UvA RD 2 and UvA RD 3, we randomly sample 214 tweets
from each reputation dimension class and train a NB classifier, and respectively an
SVM classifier. We also consider that using more training data can help our classifiers
become more robust and better learn the distinguishing features of our classes. We
explore a bootstrapping approach in runs UvA RD 4 and UvA RD 5, which train an NB
classifier and an SVM classifier, respectively, using 7,738 tweets for each reputation
dimension class. For under-represented classes we randomly sample from the labeled
training data until we reach the defined threshold.</p>
      <p>Dataset The Replab 2014 dataset is based on the Replab 2013 dataset and consists of
automotive and banking related Twitter messages in English and Spanish, targeting a
total number of 31 entities. Crawling the messages was performed in the period June
December 2012 using the entity’s canonical name as query. For each entity, there are
around 2,200 tweets collected: around 700 tweets at the beginning of the timeline used
as training set, and approximately 1,500 tweets collected at a later stage reserved as
test set. The corpus also comprises additional unlabeled background tweets for each
entity (up to 50,000, with a large variability across entities). We make use of labeled
tweets only and do not process messages for which the text content is not available
or users profiles went private. The training set consists of 15,562 tweets. Out of these,
we can access 15,294 (11,657 English, 3,637 Spanish) tweets. The test set consists of
32,446 tweets, out of we which we make predictions for 32,019 (24,254 English, 7,765
Spanish) non-empty tweets. Table 3 summarizes our training and test datasets.
Preprocessing Normalization techniques help to reduce the large vocabulary size of the
standard unigram model. Social media posts are known for the lack of language
regularity, typically containing words in multiple forms, in upper and lower case, with
character repetitions and misspellings. The presence of blogging annotations, abundance of
hashtags, emoticons, URLs, and heavy punctuation can be interpreted as possible
indicators of the rich meaning conveyed. We apply uniform lexical analysis to English</p>
      <p>LLR Unigrams
LLR Bigrams
and Spannish tweets. Our preprocessing steps are basic and aim to normalize text
content: we lowercase the tweets, remove language specific stopwords and replace Twitter
specific mentions @user, URLs and numbers with the [USER], [URL] and [NUMBER]
placeholder tags. We consider hashtags of interest since users generally supply them to
categorize and increase the visibility of a tweet. For this reason we delete hashmarks and
preserve the remaining token, i.e., #BMW is converted to BMW, so that Twitter specific
words cannot be distinguished from other words. We reduce character repetition inside
words to at most 3 characters to differentiate between the regular and the emphasized
usage of a word. All unnecessary characters ["#$%&amp;()?!*+,./:;&lt;=&gt;\ˆ{}˜] are
discarded. We apply Porter stemming algorithm to reduce inflectional forms of related
words to a basic common form.</p>
      <p>Feature selection We select our textual features by applying the LLR approach; see
Table 4 for the distribution of features over reputation dimensions in the training set.
We represent each feature as a boolean value based on whether or not it occurs inside
the tweet content. There is a bias towards extracting more features from the “Products
&amp; Services” reputation dimension class, since the majority of tweets in the training
data have this label. At the opposite end, the “Innovation” and “Leadership” classes are
among the least represented in the training set, which explains their reduced presence
inside our list of LLR extracted features.</p>
      <p>
        Training We use supervised methods for text classification and choose to train an
entity independent classifier. For our classifiers we consider Naive Bayes (NB) and
a Support Vector Machines (SVM). We motivate our choice of classifiers based on their
performance on text classification tasks that involve many word features [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref18">10,11,12,19</xref>
        ].
We train them using the different scenarios described in Section 4: making use of all
training data, balancing classes to account for the least represented class (“Innovation”,
214 tweets) and bootstrapping to consider for the most represented class (“Products &amp;
Services”, 7,738 tweets). We conduct our experiments using the natural language toolkit
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and the scikit-learn framework [
        <xref ref-type="bibr" rid="ref13">14</xref>
        ]. We use NB with default nltk.classify settings;
for SVM we choose a linear kernel.
      </p>
      <p>Evaluation Replab 2014 allowed participants to send up to 5 runs per task. For the
Reputation dimension task systems were asked to classify tweets into 7 reputation
dimension classes (see Table 1); samples tagged as “Undefined” according to human assessors
are not considered in the evaluation. Performance is measured in terms of accuracy (%
of correctly annotated cases), and precision, recall and F-measure over each class are
reported for comparison purposes. In the evaluation of our system we also take into
account the predictions made for the “Unknown” class. The predictions for this class
were ignored in the official evaluation, and therefore the absolute numbers between the
two evaluations do not match.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <p>We list the performance of our official runs in Tables 5 and 6. All our runs perform
better than the baseline (0.6221). We highlight the fact that we make predictions for
all 8 classes, including the “Undefined” category which is not considered in the official
evaluation. We also decide to ignore empty tweets, even though these are taken into
consideration by the official evaluation script!</p>
      <p>
        Our most effective run is UvA RD 4, where we train a NB classifier using a
bootstrapping approach to balance our classes. It is followed closely by UvA RD 5, which
suggests that oversampling to balance classes towards the most representative class
is a more sensible decision than using all training data or downsampling classes
towards the least represented one. When we use all training data (UvA RD 1) we provide
the SVM classifier with more informative features than when dropping tweets (in runs
UvA RD 2, UvA RD 3), which confirms the usefulness of our LLR extracted features.
Balancing classes with only 214 tweets per class can still yield competitive results,
which are rather close to our previous approaches, and a lot more accurate in the case
of the SVM classifier. We notice that NB constantly outperforms SVM. NB’s better
accuracy might be due to independence assumptions it makes among features, which is in
line with other research carried on text classification tasks where NB classifiers output
other methods with very competitive accuracy scores [
        <xref ref-type="bibr" rid="ref16 ref17 ref8">17,18,8</xref>
        ].
      </p>
      <p>Looking at the performance of our system per class, we find the following. The
“Citizenship” and “Leadership” reputation dimension classes present high precision,
followed by “Governance” and “Products &amp; Services.” Recall is very high for the latter
class, which comes as no surprise given the large number of features we extract with this
label that tend to bias the predictions of our classifier towards “Products &amp; Services.”
The F1-measure is remarkably low for the “Innovation” class, since there are only few
“Innovation” annotated tweets in the training set.</p>
      <p>Detailed statistics of the number of tweets classified per reputation dimension class
by our best system are presented in Table 7.
In our analysis section, we perform a further experiment to assess how much including
empty tweets in the evaluation and making predictions for the “Unknown” class
influences results in terms of accuracy. We regenerated our top two best runs excluding the
“Undefined” features and removing the 427 empty annotated tweets from the gold
standard test file. We report on an almost 3% improvement in accuracy for run UvA RD 4
(from 0.6704 to 0.6897) and a 2% increase in accuracy for run UvA RD 5 (from 0.6604
to 0.6739).</p>
      <p>On the one hand, we believe it is difficult to assess the performance of
submitted systems and compare methods for the task of detecting Reputation Dimensions on
Twitter data among RepLab participants since making predictions for only 7 reputation
dimension classes outperforms systems that consider the “Undefined” category. We are
not convinced that including empty tweets in the evaluation is a good idea and we were
expecting the test corpus to be re-crawled beforehand so as to ignore non-relevant
entries from the gold standard file.</p>
      <p>Finally, our suggestion is that results could be more reliable and useful if the ratio
of classified tweets would actually be considered when establishing a hierarchy of
submitted runs. We were surprised to see systems with high accuracy scores ranking high
up in the charts despite classifying fewer tweets than other runs with lower accuracy
scores and more test set samples considered. It is well-known that accuracy is highly
dependent upon the percentage of instances classified.
8</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>We have presented a corpus-based approach for inferring textual features from labeled
training data in addressing the task of detecting reputation dimensions in tweets at
CLEF RepLab 2014. Our results show that machine learning techniques can perform
reasonably accurate on text classification if the text is well modeled using appropriate
feature selection methods. Our unigram and bigram LLR features combined with an
NB classifier trained on balanced data confirm steady increases in performance when
the classification model is inferred from more example documents with known class
labels. In future work we plan to use Wikipedia pages and incorporate entity linking
methods for to improve the detection of concepts underlying the reputation dimensions
inside tweets. We would also like to probe the utility of some other classifiers, like
Random Forests, at an entity level, and consider tweets separately by language.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgements</title>
      <p>This research was partially supported by the European Community’s Seventh
Framework Programme (FP7/2007-2013) under grant agreements nr 288024 (LiMoSINe) and
nr 312827 (VOX-Pol), the Netherlands Organisation for Scientific Research (NWO)
under project nrs 727.011.005, 612.001.116, HOR-11-10, 640.006.013, the Center for
Creation, Content and Technology (CCCT), the QuaMerdes project funded by the
CLARIN-nl program, the TROVe project funded by the CLARIAH program, the Dutch
national program COMMIT, the ESF Research Network Program ELIAS, the Elite
Network Shifts project funded by the Royal Dutch Academy of Sciences (KNAW), the
Netherlands eScience Center under project number 027.012.105, the Yahoo! Faculty
Research and Engagement Program, the Microsoft Research PhD program, and the
HPC Fund.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>A.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Vovsha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Rambow</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Passonneau</surname>
          </string-name>
          .
          <article-title>Sentiment analysis of Twitter data</article-title>
          .
          <source>In Proceedings of the Workshop on Language in Social Media</source>
          ,
          <year>2011</year>
          ,
          <fpage>30</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. E. Amigo´,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Chugur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Corujo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Meij</surname>
          </string-name>
          , M. de Rijke and
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          .
          <source>Overview of Replab</source>
          <year>2013</year>
          :
          <article-title>Evaluating online reputation monitoring systems</article-title>
          .
          <source>In Information Access Evaluation. Multilinguality, Multimodality and Visualization</source>
          , volume
          <volume>8138</volume>
          <source>of LNCS</source>
          , pages
          <fpage>333</fpage>
          -
          <lpage>352</lpage>
          , Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. E. Amigo´,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrillo-de-Albornoz</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Chugur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Corujo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , E. Meij, M. de Rijke and
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          .
          <source>Overview of RepLab</source>
          <year>2014</year>
          :
          <article-title>author profiling and reputation dimensions for Online Reputation Management</article-title>
          .
          <source>In Proceedings of the Fifth International Conference of the CLEF Initiative, Sept</source>
          .
          <year>2014</year>
          , Sheffield, UK.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. E. Amigo´,
          <string-name>
            <given-names>A.</given-names>
            <surname>Corujo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Meij and M. de Rijke</surname>
          </string-name>
          .
          <source>Overview of Replab</source>
          <year>2012</year>
          :
          <article-title>Evaluating online reputation management systems</article-title>
          .
          <source>In CLEF (Online Working Notes/Labs/Workshop)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J. L. A.</given-names>
            <surname>Berrocal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. G.</given-names>
            <surname>Figuerola</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. Z.</given-names>
            <surname>Rodriguez</surname>
          </string-name>
          . Reina at
          <article-title>RepLab 2013 topic detection task: Community detection</article-title>
          .
          <source>In CLEF 2013 Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Klein</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Loper. Natural Language Processing with Python. O'Reilly Media Inc</surname>
          </string-name>
          ., Sebastopol, California,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>A.</given-names>
            <surname>Castellanos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cigarran</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Garcia-Serrano</surname>
          </string-name>
          .
          <article-title>Modelling techniques for Twitter contents: A step beyond classification based approaches</article-title>
          .
          <source>In CLEF 2013 Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>C. Gaˆrbacea</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Tsagkias and M. de Rijke</surname>
          </string-name>
          .
          <article-title>Detecting the reputation polarity of microblog posts</article-title>
          .
          <source>In ECAI 2014: 21st European Conference on Artificial Intelligence</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>V.</given-names>
            <surname>Hangya</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Farkas</surname>
          </string-name>
          .
          <article-title>Filtering and polarity detection for reputation management on tweets</article-title>
          .
          <source>In CLEF 2013 Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <article-title>Text categorization with Support Vector Machines: Learning with many relevant features</article-title>
          .
          <source>In 10th European Conference on Machine Learning</source>
          , pages
          <fpage>137</fpage>
          -
          <lpage>142</lpage>
          , Springer Verlag,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>T.</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <article-title>A statistical learning model of text classification for Support Vector Machines</article-title>
          .
          <source>In 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , pages
          <fpage>128</fpage>
          -
          <lpage>136</lpage>
          , ACM,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>A. McCallum</surname>
            and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Nigam</surname>
          </string-name>
          .
          <article-title>A comparison of event models for Naive Bayes text classification</article-title>
          .
          <source>In AAAI-98 Workshop on learning for text categorization</source>
          , pages
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          , AAAI Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          14.
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          et al..
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          ,
          <year>2011</year>
          ,
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          15.
          <string-name>
            <surname>M. H. Peetz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo and M. de Rijke</surname>
          </string-name>
          .
          <article-title>Towards and active learning system for company name disambiguation in microblog streams</article-title>
          ,
          <source>In CLEF 2013 Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          16.
          <string-name>
            <given-names>P.</given-names>
            <surname>Rayson</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Garside</surname>
          </string-name>
          .
          <article-title>Comparing corpora using frequency profiling</article-title>
          .
          <source>In Workshop on comparing corpora</source>
          , volume
          <volume>9</volume>
          , pages
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <fpage>2000</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          17.
          <string-name>
            <surname>I. Rish.</surname>
          </string-name>
          <article-title>An empirical study of the Naive Bayes classifier</article-title>
          .
          <source>International Joint Conference in Artificial Intelligence. Workshop on Empirical Methods in Artificial Intelligence</source>
          , Vol.
          <volume>3</volume>
          , No.
          <volume>22</volume>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          18. I. Rish,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Thathachar</surname>
          </string-name>
          .
          <article-title>An analysis of data characteristics that affect Naive Bayes performance</article-title>
          .
          <source>International Conference on Machine Learning</source>
          .
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          19.
          <string-name>
            <surname>K. M. Schneider</surname>
          </string-name>
          .
          <article-title>Techniques for improving the performance of Naive Bayes for text classification</article-title>
          .
          <source>In CICLing</source>
          , pages
          <fpage>682</fpage>
          -
          <lpage>693</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          20.
          <string-name>
            <given-names>D.</given-names>
            <surname>Spina</surname>
          </string-name>
          , J. C. de Albornoz, T. Martin,
          <string-name>
            <given-names>E.</given-names>
            <surname>Amigo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Giner</surname>
          </string-name>
          .
          <article-title>UNED online reputation monitoring team at RepLab 2013</article-title>
          .
          <source>In CLEF 2013 Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          21.
          <string-name>
            <given-names>C.</given-names>
            <surname>Whitelaw</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Patrick</surname>
          </string-name>
          .
          <article-title>Selecting systemic features for text classification</article-title>
          .
          <source>In Australasian Language Technology Workshop</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          22.
          <string-name>
            <given-names>I.H.</given-names>
            <surname>Witten</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Frank</surname>
          </string-name>
          .
          <source>Data Mining: Practical Machine Learning Tools and Techniques</source>
          . Morgan Kaufmann Publishers Inc.,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>