<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Domain-Based Lexicon Enhancement for Sentiment Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aminu Muhammad</string-name>
          <email>a.b.muhammad1@rgu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nirmalie Wiratunga</string-name>
          <email>n.wiratunga@rgu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Lothian</string-name>
          <email>r.m.lothian@rgu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Richard Glassey</string-name>
          <email>r.j.glassey@rgu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IDEAS Research Institute, Robert Gordon University</institution>
          ,
          <addr-line>Aberdeen</addr-line>
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>7</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>General knowledge sentiment lexicons have the advantage of wider term coverage. However, such lexicons typically have inferior performance for sentiment classification compared to using domain focused lexicons or machine learning classifiers. Such poor performance can be attributed to the fact that some domain-specific sentiment-bearing terms may not be available from a general knowledge lexicon. Similarly, there is difference in usage of the same term between domain and general knowledge lexicons in some cases. In this paper, we propose a technique that uses distant-supervision to learn a domain focused sentiment lexicon. The technique further combines general knowledge lexicon with the domain focused lexicon for sentiment analysis. Implementation and evaluation of the technique on Twitter text show that sentiment analysis benefits from the combination of the two knowledge sources. The technique also performs better than state-of-the-art machine learning classifiers trained with distantsupervision dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Introduction
Sentiment analysis concerns the study of opinions expressed in text. Typically, an
opinion comprises of its polarity (positive or negative), the target (and aspects) to which the
opinion was expressed and the time at which the opinion was expressed [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Sentiment
analysis has a wide range of applications for businesses, organisations, governments
and individuals. For instance, a business would want to know customer’s opinion about
its products/services and that of its competitors. Likewise, governments would want to
know how their policies and decisions are received by the people. Similarly, individuals
would want make use of other people’s opinion (reviews or comments) to make
decisions [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Also, applications of sentiment analysis have been established in the areas
of politics [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], stock markets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], economic systems [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and security concerns [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
among others.
      </p>
      <p>
        Typically, sentiment analysis is performed using machine learning or lexicon-based
methods; or a combination of the two (hybrid). With machine learning, an algorithm is
trained with sentiment labelled data and the learnt model is used to classify new
documents. This method requires labelled data typically generated through labour-intensive
human annotation. An alternative approach to generating labelled data called
distantsupervision has been proposed [
        <xref ref-type="bibr" rid="ref23 ref9">9, 23</xref>
        ]. This approach relies on the appearance of
certain emoticons that are deemed to signify positive (or negative) sentiment to tentatively
labelled documents as positive (or negative). Although, training data generated through
distant-supervision have been shown to do well in sentiment classification [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], it is hard
to integrate into a machine learning algorithm, knowledge which is not available from
its training data. Similarly, it is hard to explain the actual evidence on which a machine
learning algorithm based its decision.
      </p>
      <p>
        The lexicon-based, on the other hand, involves the extraction and aggregation of
terms’ sentiment scores offered by a lexicon (i.e prior polarities) to make sentiment
prediction. Sentiment lexicons are language resources that associate terms with
sentiment polarity (positive, negative or neutral) usually by means of numerical score that
indicate sentiment dimension and strength. Although sentiment lexicon is necessary for
lexicon-based sentiment analysis, it is far from enough to achieve good results [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
This is because the polarity with which a sentiment-bearing term appears in text (i.e.
contextual polarity) could be different from its prior polarity. For example in the text
“the movie sucks”, although the term ’sucks’ seems highly sentiment-bearing, this may
not be reflected by a sentiment lexicon. Another problem with sentiment lexicons is that
they do not contain domain-specific, sentiment-bearing terms. This is especially more
common when a lexicon generated from standard formal text is applied in sentiment
analysis of informal text.
      </p>
      <p>In this paper, we introduce lexicon enhancement technique (LET) to address the
the afore-mentioned problems of lexicon-based sentiment analysis. LET leverages the
success of distant-supervision to mine sentiment knowledge from a target domain and
further combines such knowledge with the one obtained from a generic lexicon.
Evaluation of the technique on sentiment classification of Twitter text shows performance
gain over using either of the knowledge sources in isolation. Similarly, the techniques
performs better than three standard machine learning algorithms namely Support
Vector Machine, Naive Bayes and Logistic Regression. The main contribution of this paper
is two-fold. First, we introduce a new fully automated approach of generating social
media focused sentiment lexicon. Second, we propose a strategy to effectively combine
the developed lexicon with a general knowledge lexicon for sentiment classification.</p>
      <p>The remainder of this paper is organised as follows. Section 2 describes related
work. The proposed technique is presented in Section 3. Evaluation and discussions
appear in Section 4, followed by conclusions and future work in Section 5.
2</p>
      <p>
        Related Work
Typically, three methods have been employed for sentiment analysis namely machine
learning, lexicon based and hybrid. For machine learning, supervised classifiers are
trained with sentiment labelled data commonly generated through labour-intensive
human annotation. The trained classifiers are then used to classify new documents for
sentiment. Prior work using machine learning include the work of Pang et al [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], where
three classifiers namely, Na¨ıve Bayes (NB), Maximum Entropy (ME) and Support
Vector Machines (SVMs) were used for the task. Their results show that, like topic-based
text classification, SVMs perform better than NB and ME. However, performance of
all the three classifiers in sentiment classification is lower than in topic-based text
classification. Document representation for machine learning is an unordered list of terms
that appear in the documents (i.e. bag-of-words). A binary representation based on term
presence or absence attained up to 87.2% accuracy on a movie review dataset [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
The addition of phrases that are used to express sentiment (i.e. appraisal groups) as
additional features in the binary representation resulted in further improvement of 90.6%
[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] while best result of 96.9% was achieved using
term-frequency/inverse-documentfrequency (tf/idf) weighting [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Further sentiment analysis research using machine
learning attempt to improve classification accuracy with feature selection mechanisms.
An approach for selecting bi-gram features was introduced in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Similarly, feature
space reduction based on subsumption hierarchy was introduced in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The
aforementioned works concentrate on sentiment analysis of reviews, therefore, they used
star-rating supplied with reviews to label training and test data instead of hand-labelling.
This is typical with reviews, however, with other forms of social media (e.g. discussion
forums, blogs, tweets e.t.c.), star-rating is typically unavailable. Distant-supervision has
been employed to generate training data for sentiment classification of tweets [
        <xref ref-type="bibr" rid="ref23 ref9">9, 23</xref>
        ].
Here, emoticons supplied by authors of the tweets were used as noisy sentiment
labels. Evaluation results on NB, ME and SVMs trained with distant-supervision data
but tested on hand-labelled data show the approach to be effective with ME attaining
the highest accuracy of 83.0% on a combination of unigram and bigram features. The
limitation of machine learning for sentiment analysis is that it is difficult to integrate
into a classifier, general knowledge which may not be acquired from training data.
Furthermore, learnt models often have poor adaptability between domains or different text
genres because they often rely on domain specific features from their training data.
Also, with the dynamic nature of social media, language evolves rapidly which may
render a previous learning less useful.
      </p>
      <p>
        The lexicon based method excludes the need for labelled training data but requires
sentiment lexicon which several are readily available. Sentiment lexicons are
dictionaries that associate terms with sentiment values. Such lexicons are either manually
generated or semi-automatically generated from generic knowledge sources. With manually
generated lexicons such as General Inquirer [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and Opinion Lexicon [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], sentiment
polarity values are assigned purely by humans and typically have limited coverage. As
for the semi-automatically generated lexicons, two methods are common, corpus-based
and dictionary-based. Both methods begin with a small set of seed terms. For example,
a positive seed set such as ‘good’, ‘nice’ and ‘excellent’ and a negative seed set could
contain terms such as ‘bad’, ‘awful’ and ‘horrible’. The methods leverage on language
resources and exploit relationships between terms to expand the sets. The two methods
differ in that corpus-based uses collection of documents while the dictionary-based uses
machine-readable dictionaries as the lexical resource. Corpus-based was used to
generate sentiment lexicon [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Here, 657 and 679 adjectives were manually annotated
as positive and negative seed sets respectively. Thereafter, the sets were expanded to
conjoining adjectives in a document collection based on the connectives ‘and’ and ‘but’
where ‘and’ indicates similar and ‘but’ indicates contrasting polarities between the
conjoining adjectives. Similarly, a sentiment lexicon for phrases generated using the web as
a corpus was introduced in [
        <xref ref-type="bibr" rid="ref29 ref30">29, 30</xref>
        ]. Dictionary-based was used to generate sentiment
lexicon in [
        <xref ref-type="bibr" rid="ref2 ref31">2, 31</xref>
        ]. Here, relationships between terms in WordNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] were explored
to expand positive and negative seed sets. Both corpus-based and dictionary-based
lexicons seem to rely on standard spelling and/or grammar which are often not preserved
in social media [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        Lexicon-based sentiment analysis begins with the creation of a sentiment lexicon or
the adoption of an existing one, from which sentiment scores of terms are extracted and
aggregated to predict sentiment of a given piece of text. Term-counting approach has
been employed for the aggregation. Here, terms contained in the text to be classified
are categorised as positive or negative and the text is classified as the class with highest
number of terms [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. This approach does not account for varying sentiment intensities
between terms. An alternative approach is the aggregate-and-average strategy [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
This classifies a piece of text as the class with highest average sentiment of terms. As
lexicon-based sentiment analysis often rely on generic knowledge sources, it tends to
perform poorly compared to machine learning.
      </p>
      <p>
        Hybrid method, in which some elements from machine learning and lexicon based
are combined, has been used in sentiment analysis. For instance, sentiment polarities
of terms obtained from lexicon were used as additional features to train machine
learning classifiers [
        <xref ref-type="bibr" rid="ref17 ref5">5, 17</xref>
        ]. Similarly, improvement was observed when multiple classifiers
formed from different methods are used to classify a document [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Also, machine
lsecaorrneinfogr wtearsm esܿ݋ݎݐ݃ܵ݊ܽ݁݅ݒ,mapslsoigynededtomሺ aoݐnpሻutൌiamlߙ൬liyz݂ݐ earሺseݐא e݌݋ݏ݅ݐݒ݁i׫n݂ݐ݁݊݃ܽݐ݅ݒ݀݋ܿݑ݉ݏnt ciሺmrݐאe݊݁݃ܽݐ݅ݒ݀݋ܿݑ݉ݏeansetdscoorrdesecirneaaseldexbiacሻ soend[o2ሻ൰ n8൅ͳሺെ ]ߙݎܵݒݏܲ݊݋ܿ݅ݔ݈݁ሻݐሺo.bHseerrvee,dinciltaisa-l
sification accuracies. ʹ
3
      </p>
      <p>Lexicon Enhancement Technique
Lexicon enhancement technique (LET) addresses the semantic gap between generic
and domain knowledge sources. As illustrated in Fig. 1, the technique involves
obtaining scores from a generic lexicon, automated domain data labelling using
distantsupervision, domain lexicon generation and aggregation strategy for classification.
Details of these components is presented in the following sub sections.</p>
    </sec>
    <sec id="sec-2">
      <title>Unlabelled data</title>
    </sec>
    <sec id="sec-3">
      <title>Domain lexicon generation</title>
    </sec>
    <sec id="sec-4">
      <title>Aggregation for sentiment classification</title>
      <p>
        3.1
We use SentiWordNet [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as the source of generic sentiment scores for terms.
SentiWordNet is a general knowledge lexicon generated from WordNet [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Each synset (i.e.
a group of synonymous terms based on meaning) in WordNet is associated with three
numerical scores indicating the degree of association of the synset with positive,
negative and objective text. In generating the lexicon, seed (positive and negative) synsets
were expanded by exploiting synonymy and antonymy relations in WordNet, whereby
synonymy preserves while antonymy reverses the polarity with a given synset. As there
is no direct synonym relation between synsets in WordNet, the relations: see also,
similar to, pertains to, derived from and attribute were used to represent synonymy relation
while direct antonym relation was used for the antonymy. Glosses (i.e. textual
definitions) of the expanded sets of synsets along with that of another set assumed to be
composed of objective synsets were used to train eight ternary classifiers. The
classifiers are used to classify every synset and the proportion of classification for each
class (positive, negative and objective) were deemed as initial scores for the synsets.
The scores were optimised by a random walk using the PageRank [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] approach. This
starts with manually selected synsets and then propagates sentiment polarity (positive
or negative) to a target synset by assessing the synsets that connect to the target synset
through the appearance of their terms in the gloss of the target synset. SentiWordNet
can be seen to have a tree structure as shown in Fig. 2. The root node of the tree is a
term whose child nodes are the four basic PoS tags in WordNet (i.e. noun, verb,
adjective and adverb). Each PoS can have multiple word senses as child nodes. Sentiment
scores illustrated by a point within the triangular space in the diagram are attached to
word-senses. Subjectivity increases (while objectivity decreases) from lower to upper,
and positivity increases (while negativity decreases) from right to the left part of the
triangle.
      </p>
      <p>
        We extract scores from SentiWordNet as follows. First, input text is broken into unit
tokens (tokenization) and each token is assigned a lemma (i.e. corresponding dictionary
entry) and PoS using Stanford CoreNLP library1. Although in SentiWordNet scores are
associated with word-senses, disambiguation is usually not performed as it does not
seem to yield better results than using either the average score across all senses of a
term-PoS or the score attached to the most frequent sense of the term (e.g. in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]). In this work, we use average positive (or negative) score at PoS level as the positive
(or negative) for terms as shown in Equation 1.
      </p>
      <p>gs(t)dim =
|senses(t,PoS)|
∑
i=1</p>
      <p>ScoreSensei(t, PoS)dim
|senses(t, PoS)|
(1)</p>
      <p>Where gs(t)dim is the score of term t (given its part-of-speech, PoS) in the sentiment
dimension of dim (dim is either positive or negative). ScoreSensei(t, PoS)dim is the
sentiment score of the term t for the part-of-speech (PoS) at sense i. Finally, |senses(t, PoS)|
is number of word senses for the part-of-speech (PoS) of term t.
Noun</p>
      <p>Verb</p>
      <p>Adjective</p>
      <p>
        Adverb
s1
sn1
s1
sn2
s1
sn3
s1
sn4
Distant-supervision offers an automated approach to assigning sentiment class labels to
documents. It uses emoticons as noisy labels for documents. It is imperative to have as
many data as possible at this stage as this affects the reliability of scores to be
generated at the subsequent stage. Considering that our domain of focus is social media, we
assume there will be many documents containing such emoticons and, therefore, large
dataset can be formed using the approach. Specifically, in this work we use Twitter as
a case study. We use a publicly available distant-supervision dataset for this stage [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]2.
This dataset contains 1,600,000 tweets balanced for positive and negative sentiment
classes. We selected first 10,000 tweets from each class for this work. This is because
the full dataset is too big to conveniently work with. For instance, building a single
machine learning model on the full dataset took several days on a machine with 8GB
RAM, 3.2GHZ Processor and 64bit Operating System. However, we aim to employ ”big
data” handling techniques to experiment with larger datasets in the future. The dataset
is preprocessed to reduce feature space using the approach introduced in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. That is,
all user names (i.e. words that starts with the @ symbol) are replaced with the token
‘USERNAME’. Similarly all URLs (e.g. “http://tinyurl.com/cvvg9a”) are replaced with
the token ‘URL’. Finally, words consisting of sequence of three or more repeated
character (e.g. ”haaaaapy”) are normalised to contain only two of such repeated character
in sequence.
      </p>
      <p>2The dataset available from Sentiment140.com
3.3</p>
      <p>Domain Lexicon Generation
Domain sentiment lexicon is generated at this stage. Each term from the
distantsupervision dataset is associated with positive and negative scores. Positive (or
negative) score for a term is determined as the proportion of the term’s appearance in positive
(or negative) documents given by equation 2. Separate scores for positive and negative
classes are maintain in order to suit integration with the scores obtained from the generic
lexicon (SentiWordNet). Table 1 shows example terms extracted from the dataset and
their associated positive and negative scores.</p>
      <p>Where ds(t)dim is the sentiment score of term t for the polarity dimension dim
(positive or negative) and tf(t) is document term frequency of t.
At this stage, scores from generic and domain lexicons for each term t are combined
for sentiment prediction. The scores are combined so as to complement each other
according to the following strategy.</p>
      <p> 0,
Score(t)dim =  gs(t)dim,</p>
      <p>ds(t)dim,
 α × gs(t)dim + (1 − α) × ds(t)dim,
if gs(t)dim = 0 and ds(t)dim = 0
if ds(t)dim = 0 and gs(t) &gt; 0
if gs(t)dim = 0 and ds(t) &gt; 0
if gs(t)dim &gt; 0 and ds(t)dim &gt; 0</p>
      <p>The parameter, α, controls a weighted average of generic and domain scores for t
when both scores are non-zero. In this work we set α to 0.5 thereby giving equal weights
to both scores. However, we aim to investigate an optimal setting for the parameter in
the future. Finally, sentiment class for a document is determined using
aggregate-andaverage method as outlined in Algorithm 1.
(2)
Algorithm 1 Sentiment Classification
1: INPUT: Document
2: OUTPUT: class
3: Initialise: posScore, negScore
4: for all t ∈ Document do
5: if Score(t)pos &gt; 0 then
6: posScore ← posScore + Score(t)pos
7: nPos ← nPos + 1
8: end if
9: if Score(t)neg &gt; 0 then
10: negScore ← negScore + Score(t)neg
11: nNeg ← nNeg + 1
12: end if
13: end for
14: if posScore/nPos &gt; negScore/nNeg then return positive
15: else return negative
16: end if
⊲ document sentiment class
⊲ increment number of positive terms
⊲ increment number of negative terms
4</p>
      <p>
        Evaluation
We conduct a comparative study to evaluate the proposed technique (LET). The aim of
the study is three fold, first, to investigate whether or not combining the two knowledge
sources (i.e. LET) is better than using each source alone. Second, to investigate
performance of LET compared to that of machine learning algorithms trained with
distantsupervision data since that is the state-of-the-art use of distant-supervision for sentiment
analysis. Lastly, to study the behaviour of LET on varying dataset sizes. We use
handlabelled Twitter dataset, introduced in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]3 for the evaluation. The dataset consists of
182 positive and 177 negative tweets.
4.1
      </p>
      <p>LET Against Individual Knowledge Sources</p>
      <sec id="sec-4-1">
        <title>Here, the following settings are compared:</title>
      </sec>
      <sec id="sec-4-2">
        <title>1. LET: The proposed technique (see Algorithm 1)</title>
        <p>2. Generic: A setting that only utilises scores obtained from the generic lexicon
(SentiWorNet). In Algorithm 1, Score(t)pos (line 5) and Score(t)neg (line 9) are replaced
with gs(t)pos and gs(t)neg respectively.
3. Domain: A setting that only utilises scores obtained from the domain lexicon. In
Algorithm 1, Score(t)pos (line 5) and Score(t)neg (line 9) are replaced with ds(t)neg
and ds(t)neg respectively.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3The dataset is available from Sentiment140.com</title>
        <p>
          been omitted by Generic. Also the result shows that the generated domain lexicon
(Domain) is more effective than the general knowledge lexicon (Generic) for sentiment
analysis.
Three machine learning classifiers namely Na¨ıve Bayes (NB), Support Vector Machine
(SVM) and Logistic Regression (LR) are trained with the distant-supervision dataset
and then evaluated with the human-labelled test dataset. These classifiers are selected
because they are the most commonly used for sentiment classification and typically
perform better than other classifiers. We use presence and absence (i.e. binary) feature
representation for documents and Weka [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] implementation for the classifiers.
Furthermore, we use subsets of the distant-supervision dataset (16000, 12000, 8000 and
4000; also balanced for positive and negative classes) in order to test the effect of
varying distant-supervision dataset sizes for LET (in domain lexicon generation, see Section
3.3) and the machine learning classifiers.
SVM
NB
LR
LET
4000
8000
        </p>
        <p>12000
Data size
16000
20000
In this paper, we presented a novel technique for enhancing generic sentiment
lexicon with domain knowledge for sentiment classification. The major contributions of
the paper are that we introduced a new approach of generating domain-focused lexicon
which is devoid of human involvement. Also, we introduced a novel strategy to
combine generic and domain lexicons for sentiment classification. Experimental evaluation
shows that the technique is effective and better than state-of-the-art machine learning
sentiment classification trained the same dataset from which our technique extracts
domain knowledge (i.e. distant-supervision data).</p>
        <p>As part of future work, we plan to conduct an extensive evaluation of the technique
on other social media platforms (e.g. discussion forums) and also, to extend the
technique for subjective/objective classification. Similarly, we intend perform experiment
to find an optimal setting for α and improve the aggregation strategy presented.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Arnold</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vrugt</surname>
          </string-name>
          , E.:
          <article-title>Fundamental uncertainty and stock market volatility</article-title>
          .
          <source>Applied Financial Economics</source>
          <volume>18</volume>
          (
          <issue>17</issue>
          ),
          <fpage>1425</fpage>
          -
          <lpage>1440</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baccianella</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esuli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sebastiani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining</article-title>
          .
          <source>In: Proceedings of the Annual Conference on Language Resouces and Evaluation</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Baron</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Competing for the public through the news media</article-title>
          .
          <source>Journal of Economics and Management Strategy</source>
          <volume>14</volume>
          (
          <issue>2</issue>
          ),
          <fpage>339</fpage>
          -
          <lpage>376</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The anatomy of a large-scale hypertextual web search engine</article-title>
          .
          <source>In: Seventh International World-Wide Web Conference (WWW</source>
          <year>1998</year>
          )
          <article-title>(</article-title>
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Dang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , H.:
          <article-title>A lexicon-enhanced method for sentiment classification: An experiment on online product reviews</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>25</volume>
          ,
          <fpage>46</fpage>
          -
          <lpage>53</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Denecke</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Using sentiwordnet for multilingual sentiment analysis</article-title>
          .
          <source>In: ICDE Workshop</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Esuli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baccianella</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sebastiani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Sentiwordnet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining</article-title>
          .
          <source>In: Proceedings of the Seventh conference on International Language Resources and Evaluation (LREC10)</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet: An Electronic Lexical Database</article-title>
          . MIT Press (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Go</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhayani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Twitter sentiment classification using distant supervision</article-title>
          . Processing pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>The weka data mining software: an update</article-title>
          .
          <source>SIGKDD Explor. Newsl</source>
          .
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          (
          <year>Nov 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hatzivassiloglou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McKeown</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          :
          <article-title>Predicting the semantic orientation of adjectives</article-title>
          .
          <source>In: Proceedings of the 35th Annual Meeting of the ACL and the 8th Conference of the European Chapter of the ACL</source>
          . pp.
          <fpage>174</fpage>
          -
          <lpage>181</lpage>
          . New Brunswick, NJ (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Mining and summarizing customer reviews</article-title>
          .
          <source>In: Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <fpage>168</fpage>
          -
          <lpage>177</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Karlgren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sahlgren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olsson</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Espinoza</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamfors</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Usefulness of sentiment analysis</article-title>
          .
          <source>In: 34th European Conference on Information Retrieval</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Sentiment Analysis and Subjectivity, chap</article-title>
          .
          <source>Handbook of Natural Language Processing</source>
          , pp.
          <fpage>627</fpage>
          -
          <lpage>666</lpage>
          . Chapman and Francis, second edn. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Ludvigson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Consumer confidence and consumer spending</article-title>
          .
          <source>The Journal of Economic Perspectives</source>
          <volume>18</volume>
          (
          <issue>2</issue>
          ),
          <fpage>29</fpage>
          -
          <lpage>50</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Mukras</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiratunga</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lothian</surname>
          </string-name>
          , R.:
          <article-title>Selecting bi-tags for sentiment analysis of text</article-title>
          .
          <source>In: Proceedings of the Twenty-seventh SGAI International Conference on Innovative Techniques and Applications of Artificial Intelligence</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Ohana</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tierney</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Sentiment classification of reviews using sentiwordnet</article-title>
          .
          <source>In: 9th IT&amp;T Conference</source>
          , Dublin, Ireland (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <source>Polarity dataset v2.0</source>
          ,
          <year>2004</year>
          . online (
          <year>2004</year>
          ), http://www.cs.cornell.edu/People/pabo/movie-review
          <article-title>-data/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Opinion mining and sentiment analysis</article-title>
          .
          <source>Foundations and Trends in Information Retrieval</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaithyanathan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Thumbs up? sentiment classification using machine learning techniques</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods on Natural Language Processing</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Pera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qumsiyeh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>Y.K.</given-names>
          </string-name>
          :
          <article-title>An unsupervised sentiment classifier on summarized or full reviews</article-title>
          .
          <source>In: Proceedings of the 11th International Conference on Web Information Systems Engineering</source>
          . pp.
          <fpage>142</fpage>
          -
          <lpage>156</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Prabowo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>: sentiment analysis: A combined approach</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <fpage>143</fpage>
          -
          <lpage>157</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Read</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Using emoticons to reduce dependency in machine learning techniques for sentiment classification</article-title>
          .
          <source>In: Proceedings of the ACL Student Research Workshop</source>
          . pp.
          <fpage>43</fpage>
          -
          <lpage>48</lpage>
          . ACLstudent '05,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Riloff</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patwardhan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiebe</surname>
          </string-name>
          , J.:
          <article-title>Feature subsumption for opinion analysis</article-title>
          .
          <source>In: Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing (EMNLP-06)</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Stone</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dexter</surname>
            ,
            <given-names>D.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daniel</surname>
            ,
            <given-names>O.M.</given-names>
          </string-name>
          :
          <article-title>The General Inquirer: A Computer Approach to Content Analysis</article-title>
          . MIT Press, Cambridge, MA (
          <year>1966</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brooke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tofiloski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voll</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stede</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Lexicon-based methods for sentiment analysis</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>37</volume>
          ,
          <fpage>267</fpage>
          -
          <lpage>307</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paltoglou</surname>
          </string-name>
          , G.:
          <article-title>Sentiment strength detection for the social web</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>63</volume>
          (
          <issue>1</issue>
          ),
          <fpage>163</fpage>
          -
          <lpage>173</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kappas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Sentiment strength detection in short informal text</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>61</volume>
          (
          <issue>12</issue>
          ),
          <fpage>2444</fpage>
          -
          <lpage>2558</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , et al.:
          <article-title>Mining the web for synonyms: Pmi-ir versus lsa on toefl</article-title>
          .
          <source>In: Proceedings of the twelfth european conference on machine learning (ecml-2001)</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          :
          <article-title>Thumbs up or thumbs down? semantic orientation applied to unsupervised classification of reviews</article-title>
          .
          <source>In: Proceedings of the Annual Meeting of the Association for Computational Linguistics</source>
          . pp.
          <fpage>417</fpage>
          -
          <lpage>424</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Valitutti</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Wordnet-affect: an affective extension of wordnet</article-title>
          .
          <source>In: In Proceedings of the 4th International Conference on Language Resources and Evaluation</source>
          . pp.
          <fpage>1083</fpage>
          -
          <lpage>1086</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Whitelaw</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Argamon</surname>
            .,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Using appraisal groups for sentiment analysis</article-title>
          .
          <source>In: 14th ACM International Conference on Information and Knowledge Management (CIKM</source>
          <year>2005</year>
          ). pp.
          <fpage>625</fpage>
          -
          <lpage>631</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>