<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Refinement of Russian Sentiment Lexicons Using RuThes Thesaurus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>© N. V. Loukachevitch © I. I. Chetviorkin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Proceedings of the 16th All-Russian Conference “Digital Libraries: Advanced Methods and Technologies</institution>
          ,
          <addr-line>Digital Collections” ― RCDL-2014, Dubna</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>61</fpage>
      <lpage>65</lpage>
      <abstract>
        <p>The paper describes a combined approach to extraction of a domain-specific sentiment lexicon. At first, an initial version of a domainspecific lexicon is obtained by application of a supervised model. At the second stage, the ordered list of sentiment words is refined using the thesaurus informioant. This combined model is applied to several domains and at last the domain-specific sentiment lexicons are united to create an improved version of the Russian sentiment lexicon in the generalized domain of products.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1 Introduction At the second stage, the ordered list of sentiment
words is refined using the thesaurus information, in our</p>
      <p>
        Automatic sentiment analysis of texts is a fascta-se, newly published thesaurus of Russian language
developing technology in natural language processing. RuThes1. We trained a supervised model and tuned a
The task of automatic sentiment lexicon construction combined model in the movie domain. Then
and improvement is a basic tsak for sentiment analysis augmented model was utilized in four other domains.
of texts. There are fnreoely available sentiment Finally, extracted sentiment lexicons from five domains
lexicons for many languages or the quality of sucahre united to generate a high quality lexicon in the
lexicons is desired to be better. For example, in Russian general product domain for
only one automatically extracetd sentiment lexicon has (ProductSentiRus+).
been published [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>The reminder of this article is organized as follows.</title>
      <p>Besides, sentiment analysis of domain-specific texts In Section 2 we review methods for generating
requires adaptation of machine-learning models oserntiment lexicons. Section 3 briefly presents the
sentiment lexicons to the target domain [6]. So, some structure of RuThes thesaurus, the Russian newly
sentiment words can loss their polarity in specipfiucblished thesaurus intended for natural language
domains. For example, such word asevil in the movie processing. Section 4 presents an approach for
domain usually refers to the movie plot, but not a user extracting sentiment words in various domains. Section
opinion. 5 describes the refinement of the lexicon in the general</p>
      <p>Other words can obtain the sentiment polarity in a product domain. To evaluate the quality of the obtained
specific domain. For exa,mpwleord киношный general resource extrinsically, we conduct
(adjective to Russian wordкино (movie)) can have the experiments on the tweet subjectivity classification task.
negative polarity with the meaning"far from the real
life". Another example - word атмосферный (adjective 2 Related Work
to word атмосфера (atmosphere)) has the positive
polarity in art-related domains denoting "creation of a There are two main approaches to sentiment lexicon
special mood or feeling" (asatmospheric in English) – extraction: corpus-based and dictionary-based methods.
this is a relatively newsense of this word for Russian, of
not described in Russian dictionaries.</p>
      <p>this</p>
    </sec>
    <sec id="sec-3">
      <title>Corpus-based methods utilize co-occurrence</title>
      <p>
        words with each other [5, 9, 10], or appearance them in
specific collocations or lexico-syntactic patterns
Contemporary corpus-based approaches exploit a large
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>the</p>
      <sec id="sec-3-1">
        <title>1 labinform.ru/ruthes/index.htm</title>
        <p>R
reviews as in [6].</p>
        <p>relations, which includes terms
or hundreds of thousands of useerconomic, military, sports and other fields [7].
from
political,</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Ambiguous words in RuThes are described similar</title>
      <p>Dictionary-based methods utilize avatiolableWordNet-style resources through attachment
electronic dictionaries and thesauri and usually begin several concepts. For example, in the current version of
their work from a set of seed Inwor[d3s]. RuThes word пресный is attached to three concepts:
SentiWordNet resource is described. It is the result of  ПРЕСНАЯ ВОДА (fresh water);
the automatic annotation of all the synsets of WordNet
where each synset is associated to three numerical  ПРЕСНЫЙ, БЕЗВКУСНЫЙ (tasteless, bland in
scores that indicate how positive, negative, and neutral taste);
the terms contained in the synset are. Different senses of  ПРЕСНЫЙ (НЕИНТЕРЕСНЫЙ) (uninteresting).
the same word may thus have different opinion-related
properties.</p>
    </sec>
    <sec id="sec-5">
      <title>In this section an algorithm for extraction</title>
      <p>sentiment words in a specific domain is described. The
results of this algorithm are refined using the iterative
procedure on the basis of RuThes thesaurus to obtain a
high quality domain-specific sentiment lexicon.</p>
      <p>In many studies domain-specific sentiment lexicons
are created with corpus-based approaches using various
types of propagation from a seed set of words, usually a Such a method is applied to four other domains
general sentiment lexicon [6]. An important problem of without additional manual labeling and the results are
such approaches is to determine an appropriate seedcombined in a sentiment lexicon in a generalized
lexicon, which can depend on the domain. product domain ProductSentiRus+.
to</p>
      <p>In our study we create adomain-specific sentiment Table 1. Domain-specific collection statistics
lexicon from medium-size datasets using multiple
features of words and several collections without any Domain Reviews Descriptions
co-occurrences between worsd. Then we improve an Movies 28, 773 17, 680
initial sentiment lexicon using sentiment labeling of the Books 23, 883 22, 321
thesaurus concepts in a specific domain practically Games 7, 928 1, 853
without pre-determined seed words. We use only two Digital Cameras 10, 208 920
fixed seed opinionated wordbsad, ( good), other Mobile Phones 30, 620 890
potential sentiment words are obtained automatically
from a ranked list of a sentiment lexicon (words ordered 4.1 Extraction of domain-specific sentiment lexicon
by the probability of their sentiment orientatbioans)ed on multiple features
extracted from domain-specific collections.</p>
    </sec>
    <sec id="sec-6">
      <title>At the first stage sentiment words are extracted with</title>
      <p>3 RuThes Linguistic Ontology a corpus-based method utilizing a trained
machinelearning model applied to several domain-specific text</p>
      <p>In our study we use RuThes Thesaurus of Russian collections.
language. RuThes is a linguistic ontology for natural The first domain-speccifi collection (with high
language processing, i.e. an ontology, where cothnecentration of sentiment words) is a collection of user
majority of concepts are introduced on the basis orefviews in the domain (review collection)with numeric
actual language expressions. For a long time RuThes scores specified by their author.sIn these experiments
has been manually developed within various NLP and collections were gathered from the neonliservices
information-retrieval projects, and now it is available imhonet.ru and market.yandex.ru in five domains:
for public use. The publicly available version of RuThes movies, books, computer games, mobile phones and
contains around 100 thousand Russian words adnidgital cameras. The second domain-specific collection
expressions [7]. (with low concentration of sentiment words) is a text</p>
      <p>
        If compared to WordNet-style resources RuThes is collection of object descriptions (e.g. plots for movies).
organized as a united semantic net where different parts The overall collection statistics can be found in Table 1.
of speech (nouns, verbs, adjectives) can be text entries Another contrast corpus was a collection of two million
of the same concepts. Each concept has a uninqeuwe s documents. Such a collection is useful for correct
unambiguous name. Concepts can be connected with classification of general neutral words frequent in news.
several types of conceptual relations. In addition,
RuThes includes a lot of multiword expressions useful Using such collections the feature representation is
for applications and terms of so-called Sociopolitical calculated for each word. The set of features includes
domain – a broad domain of contemporary socthiael following feature types [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]:
of
 Score-based: deviation from the average score,
word score variance, sentiment category likelihood for
each (word, category) pair;
 Linguistic: Four binary features indicating the
word part of speech, two binary features reflecting POS
ambiguity, predefined list of prefixes of a word.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Movie</title>
      <p>Books
Games
worTdos trwaiinthsuptheervisferdeqmueancchyineglreeaartneirng tahlagnoritthhmrese, aliln thCeDaimgietraals 85 92 65.8 66.3
amsosevsiesorrse.viewIf coltlheecrteion wwaesre laabeleddisamgareneumalelnyt by atbwoout tPMhheoobnieles 85 97 73.2 78.6
sentiment of a specific word, the collective judgment General
after discussion was used as the final ground truth. As a Product 100 100 90.5 95.2
result of the assessment procedure the list of 4079Domain
sentiment words was obtained. The best quality of
classification using labeled data was shown by the  
ensemble of three classifiers: Logistic Regression, Words from the ranked sentiment list are quite
LogitBoost and Random Forest from WdifEfeKreAnt relative to RuThes descriptions. Some words
programming package. are not described in RuTh,ese.g. three of the most
probable sentiment wordsin the movie domain are</p>
      <p>The result of this corpus-based method is a ranked absent in RuThes, others are mentioned in
list of domain-specific words ordered by the probability collections exactly in the same senses as described in
of their sentiment orientation – fusrtehnetriment RuThes, the thirds (e.g.atmospheric) are described in
weights. The algorithm boosts sentiment words to have RuThes but have an additional (or the other) sentiment
high weights (to be closer to the beginning of the list) polarity. So we should try to correct the word order in
and neutral words to have low weights. the sentiment list carefully applying</p>
      <p>So in the movie domain in the list of more than 18 descriptions.
thousand words the following words are located in the
first positions:</p>
      <p> Frequency-based: collection frequencyw,ith the model described in the previous subsection;
document frequency, frequenyc of capitalized words, however, a similar input can be also generated with
frequency of co-occurrence with polarity shifters (no, other methods.
not), TFIDF;</p>
    </sec>
    <sec id="sec-8">
      <title>RuThes as text The</title>
      <sec id="sec-8-1">
        <title>Word атмосферный (atmospheric) takes</title>
        <p>high-opinionated position in the list.</p>
        <p>The main idea of the lexicon refinement is to label
conceptual subgraphs of the thesaurus network
sentiment or neutral and use this labeling to reorder the
initial sentiment list. This process in contrast to such a
method as Label Propagation [8, 11s]hould be also
regulated with previously obtained sentiment weights of
830th, words.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Let us denote a domain-specific lexicon wiWthD</title>
      <p>Evident sentiment adjectives of the movie domain where all words are orderedby their sentiment weights
пресный and безвкусный (both are translated into (sw). Initially the algorithm forms two sets of thesaurus
English as tasteles)s take even higher opinionatedconcepts using words from the both sides of the list WD:
positions: 139th and 193th . Bu ttheir noun derivations Ls – concepts supposed to be opinionatedL,n – neutral
пресность, безвкусие, безвкусность, and  безвкусица concepts. With this aim the initial average sentiment
are less successfПulр.есность, безвкусие, weights csw for all concepts containing words fromWD
безвкусность, are absent from the list because of low are calculated. Then the algorithm adds toLs concepts
frequency; безвкусица takes 1515th place in the list. So with the high average weight (csws &gt; 0.85) and also two
thesaurus-based improvements may be quite possible. pre-defined concepts, corresponding to senses of words</p>
      <p>The obtained model was applied to four othbeard and good.
domains (books, games, digital cameras, mobileConcepts with the low average weight (cswn &lt; 0.05)
phones) without any additional manual efforts. Thaere added to the set of neutral concLepn,ts which
quality of extracted sentiment lexicons was measured formed without any pre-defined concepts.
using precision measures and presented in the Baseline thresholds for csws and cswn are obtained from
columns of Table 2. experiments.
4.2 Refinement of domain-sentiment sentiment
lexicons using RuThes thesaurus</p>
    </sec>
    <sec id="sec-10">
      <title>To increase the quality of extracted lexicons we refine them with general thesaurus Russian language RuThes [7]. The input refinement algorithm is a ranked sentiment list obtained</title>
      <p>Further, every setLs (and Ln) is iteratively
augmented with concepts using two conditions: the
average sentiment weight threshold and the number of
sentimdeinrtect thesaurus relations to the existing sets. Formally,
foLrs and Ln are calculated as shown in Algorithm 1
of listtihneg. The algorithm uses also the following additional
notation:
 Adj (L) is a set of direct-link neighbor concepts to
set of concepts L;</p>
      <p> nlink (C, L) is a function returning the number of
direct thesaurus relations between concept C and set L.</p>
    </sec>
    <sec id="sec-11">
      <title>In the last stsewp weights of all</title>
      <p>corresponding to Ls concepts are modified
multiplying them by factork1 (k1 &gt; 1) and all words
corresponding to Ln are multiplied by fackt2or
(0 &lt; k2 &lt; 1). The resulting list is reordered by weight.</p>
    </sec>
    <sec id="sec-12">
      <title>In that paper the lexicons of five domains were</title>
      <p>Low-frequent words (with the frequency less than 3) summed up using a formula intended to boost words
of the source domain collection are absent in the initial that occur in many different domains and have high
ranked sentiment list and therefore do not have anyweights in each of them.
sentiment weights. The initial sentiment weights of such Thus, for combining multiple weighted word lists
words are calculated as the average sentiment weights the following formula was used:
fcorofenqcuecepontntsc,seypntiosnnytmthuesryno,r frreaolrmaeteadvcearlatcogu.elawteTedhigehftsrowomefingehoitgsthhebroo,rf mthoersee R(w)  max( pdroDbd (w)) dD D1 1 posdd (w)  ,
concepts in the labeling process.</p>
      <p>5 Improvement of General Sentiment
Lexicon Using RuThes Thesaurus</p>
      <p>
        Integrating sentiment lexicons from various
productworodrsiented domains it is possible to create a general
sentiment lexicon in the broad domain of products and
by
services. Such a lexicon for Russian was described in
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], it was called ProductSentiRus2.
      </p>
      <sec id="sec-12-1">
        <title>2 http://www.cir.ru/SentiLexicon/ProductSentiRus.txt</title>
        <sec id="sec-12-1-1">
          <title>Algorithm 1. Weights+Relations</title>
          <p>Input: concept list with sentiment
weights csw
Output: Ls, Ln</p>
          <p>Ls = {Cbad, Cgood}  Chigh, Chigh={Ci:
csw(Ci)&gt;0.85},</p>
          <p>Ln = Ln  Clow, Clow = {Ci: csw(Ci)&lt;0.05}
θ = 0.1, Nlink = 3, Ls_iter = Ls, Ln_iter
= Ln
while θ&lt;0.6
for C  Adj(Ls)</p>
          <p>if nlink(C, Ls) &gt; Nlink &amp;&amp;
csw(C)&gt;0.7-θ</p>
          <p>then Include(C,Ls);
end
for C  Adj(Ln)</p>
          <p>if nlink(C, Ln) &gt; Nlink &amp;&amp; csw(C)&lt;
θ
end</p>
          <p>then Include(C,Ln);
if Ln == Ln_iter &amp;&amp; Ls == Ls_iter</p>
          <p>then Nlink = Nlink-1;
if Nlink == 0</p>
          <p>then θ = θ + 0.05, Nlink = 3;</p>
          <p>Ln = Ln_iter, Ls = Ls_iter
end
where D – is the domain set with five domains,d is the
sentiment word list for a particular domain and d is the
total number of words in this list. Functionpsrobd (w)
and posd (w) are the sentiment probability and position
of the word   thine lisdt . Precision@1000 of
ProductSentiRus was eporrted as 90.5%. Similar
combination of improved sentiment lexicons in the new
resource (ProductSentiRus+) yields 95.2% in terms of
Precision@1000 (Table 2).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>We took 5000 of the most probable sentiment words of ProductSentiRus+ lexicon for further work (the same amount as in a previous version) and evaluated it in the tweet subjectivity classification task.</title>
    </sec>
    <sec id="sec-14">
      <title>The evaluation is based on TEST data set described</title>
      <p>in [10], which include two thousand tweets in Russian.
We assumed that ProductSentiRus+ comprises
sentiment units of Internet language. A tweet was
classified as subjective if it contained at least one word
from the lexicon. Table 3 demonstrates that such a
generalized lexicon can be useful also in tweet
subjectivity analysis.</p>
      <p>After application of this algorithm in the
domain our example wordпsресный, пресность,
безвкусный, пресность, безвкусица have the following
places in the generated sentiment lisпt:ресный – 81,
пресность – 86, безвкусный – 115, безвкусность –
172, безвкусие – 173, безвкусица – 943.
movie</p>
    </sec>
    <sec id="sec-15">
      <title>In this paper we described a combined approach to</title>
      <p>extraction of domain-specific sentiment lexicons. At
first, an initial version of a domain-specific lexicon is</p>
      <p>The words related to the neutral sense of woorbdtained by application of a supervised model. At the
пресный – ПРЕСНАЯ ВОДА (fresh water) preserved second stage, the ordered list of sentiment words is
their very low positions in the sentimleisntt: вода refined using information described in RuThes
(water) – 23059, айсберг (iceberg) – 26124.</p>
      <p>P
–
–</p>
      <p>R
–
–
thesaurus of Russian language, which was late[l5y] Vasileios Hatzivassiloglou and Kathleen R
published. McKeown. Predicting the semantic orientation of</p>
      <p>This combined model is applied to several domains adjectives. In Proceedings of the 35th Annual
and at last domain-specific sentiment lists are united to Meeting of the Association for Computational
create a sentiment word listin the generalized domain Linguistics, 1997. P. 174–181.
of products – ProductSentiRus+, which is an improved [6] Raymond Lau, Chun-Lam Lai, Peter Bruza,
Kamversion of the only published Russian sentiment lexicon Fai Wong. Leveraging web 2.0 data for scalable
and will be also publicly available. The proposed semi-supervised learning of domain-specific
approach can be applied to other languages and can sentiment lexicons. Proceedings of the 20th ACM
utilize other thesauri. international conference on Information and
knowledge management. ACM. 2011.</p>
      <p>Acknowledgments [7] Natalia Loukachevitch and Boris Dobrov. RuThes
07-0T0h6is82w. ork is partially supported by RFBR grant 14- PLrinocgeueisdtiincgOs notfoGlolgoybavlsW.RoursdsniaentCWoonrfderneentsc.eI.n2013.</p>
      <p>[8] Delip Rao and Deepak Ravichandran.
SemiReferences supervised polarity lexicon induction. In
Proceedings of the 12th Conference of the European
Chapter of the ACL, EACL-2009, 2009. P. 675–
682.
[9] Leonid Velikovich, Sasha Blair-Goldensohn,</p>
      <p>Kerry Hannan, and Ryan McDonald. The viability
of web-derived polarity lexicons. In Human
Language Technologies: The 2010 Annual Conference
of the North American Chapter of the Association
for Computational Linguistics, 2010. P. 777–785.
[10] Svitlana Volkova, Theresa Wilson, and David</p>
      <p>Yarowsky. Exploring sentiment in social media:
Bootstrapping subjectivity clues from multilingual
twitter streams. In Proceedings of the 51st Annual
Meeting of the Association of Computational</p>
      <p>Linguistics (ACL13), 2013. P. 505–510.
[11] Xiaojin Zhu and Zoubin Ghahramani. Learning
from labeled and unlabeled data with label
propagation. Technical Report
CMU-CALD-02107, Carnegie Mellon University. 2002.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ilia</given-names>
            <surname>Chetviorkin</surname>
          </string-name>
          and Natalia V Loukachevitch.
          <article-title>Extraction of Russian sentiment lexicon for product meta-domain</article-title>
          .
          <source>In COLING</source>
          ,
          <year>2012</year>
          . P.
          <volume>593</volume>
          -
          <fpage>610</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Claire</given-names>
            <surname>Cardie</surname>
          </string-name>
          .
          <article-title>Adapting a polarity lexicon using integer linear programming for domain-specific sentiment classification</article-title>
          .
          <source>In Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing:</source>
          Vol.
          <volume>2</volume>
          ,
          <year>2009</year>
          . P.
          <volume>590</volume>
          -
          <fpage>598</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Esuli</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fabrizio</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          .
          <article-title>Sentiwordnet: A publicly available lexical resource for opinion mining</article-title>
          .
          <source>In Proceedings of LREC</source>
          , vol.
          <volume>6</volume>
          ,
          <year>2006</year>
          . P.
          <volume>417</volume>
          -
          <fpage>422</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Song</given-names>
            <surname>Feng</surname>
          </string-name>
          , Jun Seok Kang, Polina Kuznetsova,
          <string-name>
            <given-names>Yejin</given-names>
            <surname>Choi</surname>
          </string-name>
          .
          <article-title>Connotation Lexicon: A Dash of Sentiment Beneath the Surface Meaning</article-title>
          .
          <source>In Proceedings of the 51th Annual Meeting of the Association for Computational Linguistics ACL2013</source>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>