<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Lexicon-based methods and BERT model for sentiment analysis of Russian text corpora*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anastasiya Kotelnikova</string-name>
          <email>kotelnikova.av@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danil Paschenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Razova</string-name>
          <email>razova.ev@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vyatka State University</institution>
          ,
          <addr-line>36, Moskovskaya st., Kirov, 610000, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article discusses two approaches to solving the problem of sentiment analysis: lexicon-based approach and deep learning. Within the first approach two well-known lexicon-based methods has been adapted for the Russian language - SO-CAL and SentiStrength. For these methods a unified sentiment lexicon has been prepared using a voting procedure based on the existing lexicons. The second approach has used the RuBERT deep learning model. The SentiRuEval-2015 corpora, which provides reviews and tweets, has been used as training and test data. Analysis of the results showed that deep learning model demonstrates higher accuracy compared to lexicon-based methods.</p>
      </abstract>
      <kwd-group>
        <kwd>Sentiment analysis</kwd>
        <kwd>deep learning</kwd>
        <kwd>BERT</kwd>
        <kwd>RuBERT</kwd>
        <kwd>sentiment lexicons</kwd>
        <kwd>SO-CAL</kwd>
        <kwd>SentiStrength</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Sentiment analysis is a field of computational linguistics aimed at automated research
of people’s opinions and assessments in relation to various objects mentioned in the
text, for example, products, services, organizations, persons, events [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Sentiment is
represented as a value on a certain scale: binary (positive / negative attitude), ternary
(adding neutral or contradictory), n-ary or real (for example, [
        <xref ref-type="bibr" rid="ref5">–5, 5</xref>
        ]).
      </p>
      <p>
        There are many studies in the field of sentiment analysis, mainly for the English
language [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2–5</xref>
        ], but in the last decade works for the Russian language have also
appeared [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        There are three main approaches to the sentiment analysis in texts – lexicon-based,
machine learning, and hybrid, in which the two indicated approaches are combined
[
        <xref ref-type="bibr" rid="ref2 ref3">2–3</xref>
        ].
      </p>
      <p>
        Examples of existing lexicon-based sentiment analysis methods are SO-CAL [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
and SentiStrength [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Both methods use sentiment lexicons and assess the strength of
positive and negative sentiments in texts. The methods take a text as input and
produce a numerical value that averages the sentiment of the words in text found from
the sentiment lexicon. They were originally designed for only English texts. In
general, a lexicon-based approach requires high-quality sentiment lexicon, the analysis
process is quite fast, it doesn’t need training data, but the accuracy is often not high
enough.
      </p>
      <p>
        The second approach to sentiment analysis is machine learning, within which there
are two areas: traditional machine learning (for example, such methods as SVM,
Naïve Bayes, and Gradient Boosting) and deep learning (for example, models based on
the Transformer architecture such as BERT) [
        <xref ref-type="bibr" rid="ref8 ref9">8–9</xref>
        ]. The best results have recently
been obtained based on deep learning models [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. However, such models, having
high accuracy, are poorly interpretable compared to lexicon-based methods [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], their
application requires high-quality labeled training data, while a significant amount of
time is spent on the training procedure. In addition, deep models do not take into
account the knowledge about sentiment words contained in the corresponding lexicons.
      </p>
      <p>The purpose of this work is to compare the performance of the lexicon-based
methods SO-CAL and SentiStrength adapted for the Russian language and the BERT
deep learning model for sentiment analysis, as well as to explore the possibility of
adding information from sentiment lexicons to deep learning models.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Materials and methods</title>
      <sec id="sec-2-1">
        <title>Sentiment analysis methods</title>
        <p>
          SO-CAL† (Semantic Orientation CALculator) is a method developed by Maite
Taboada that determines the sentiment of texts [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. SO-CAL works for English and
Spanish. We have adapted it to the Russian language.
        </p>
        <p>The first change affected preprocessing – for Russian texts it is required to
determine not only the part of speech of tokens, as in the original version, but also their
initial form. The rnnmorph‡ module was used for this.</p>
        <p>Secondly, we used a lexicon created by combining existing sentiment lexicons
(described below). With the help of rnnmorph the combined lexicon was split into four
separate lexicons for different parts of speech: nouns, adjectives, verbs, adverbs. If an
element was attributed to several of the specified parts of speech, it was used in
several lexicons.</p>
        <p>Thirdly, Russian-language lists of special words and expressions, used in the
algorithm, were formed, in particular, a lexicon of modifiers (words that affect the
sentiment of the words to which they belong, for example, less, absolutely) and lists of
negations (for example, nothing, without).</p>
        <p>As in the original version of the method, the Russian version ignores the sentiment
words in the sentence if there is a condition word from the special list in the sentence.
Such list includes conditional markers (for example, if), some verbs (like expect,
doubt), questions and words enclosed in quotation marks (which may be factual, but
† https://github.com/sfu-discourse-lab/SO-CAL.
‡ https://github.com/IlyaGusev/rnnmorph.
do not necessarily reflect the opinion of the author). If repeating the same sentiment
word, there is a decrease in weight for each subsequent repetition.</p>
        <p>The version of SO-CAL adapted for the Russian language, just like the original
one, for each text produces a numerical value corresponding to the degree of
sentiment of the text: a value greater than zero indicates a positive sentiment, a value less
than zero indicates a negative sentiment.</p>
        <p>
          On the training part of the corpora a threshold is determined that separates the texts
into positive and negative ones (for a three-class classification two thresholds are
selected to divide into positive, neutral and negative ones). Next, the sentiment of the
texts of the test part is determined, taking into account the found thresholds.
SentiStrength is a lexicon-based method developed by Mike Thelwall et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. For
texts in English it gives two numerical values: the first score is from –1 for texts that
are not negative, to –5 for extremely negative texts, the second one is from 1 for texts
that are not positive, and up to 5 for extremely positive texts.
        </p>
        <p>The method was originally developed for the English language, we adapted it for
the Russian language by changing the linguistic resources required for the algorithm.
These are a list of sentiment words, a list of modifiers – words that raise or lower the
sentiment score of the following words (for example, bad, a little, very, extremely), a
list of negations (for example, not, never), a list of words that indicate the presence of
a question in a sentence (for example, how, when, why), etc. The replacement of such
language-independent resources as the list of emoticons was not carried out.</p>
        <p>Experiments with SentiStrength were performed with two versions of the datasets.
The first version is raw, unprocessed data. The second version is preprocessed
(lemmatized) data.</p>
        <p>
          RuBERT. In addition to the SO-CAL and SentiStrength methods, we used a model
based on the Russian-language version of BERT – RuBERT [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The multilingual
version of the BERT-base is used as an initialization for RuBERT. The model was
trained on the Russian part of Wikipedia and news articles. In our experiments the
models were fine-tuned based on training data and tested on test data.
        </p>
        <p>In experiments with the RuBERT model two variants of text corpora were used: a
corpus without preprocessing and a corpus in which positive words present in the
combined lexicon were replaced by good, and negative words by bad. This
preprocessing procedure made it possible to test the hypothesis of a potential improvement in
the performance of sentiment analysis based on RuBERT when adding information
from the sentiment lexicon.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Linguistic resources</title>
        <p>The main linguistic resources for solving the problem of sentiment analyses are
sentiment lexicons and text corpora labelled by sentiment.</p>
        <p>
          Sentiment lexicons. The combined sentiment lexicon was formed using nine publicly
available lexicons for the Russian language [6; 12].
1. EmoLex [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Created with crowdsourcing. Words in the lexicon are associated
with positive and negative sentiments and with emotions such as anger,
anticipation, disgust, fear, joy, sadness, surprise, and trust. This lexicon has been translated
into more than 100 languages (including Russian) using Google Translate.
2. Chen-Skiena’s lexicon [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Automatically built for 136 languages, including
Russian, using graph propagation.
3. LinisCrowd [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Created with crowdsourcing. An initially selected list of 7,546
words based on a list of high-frequency adjectives, the lexicon ProductSentiRus
[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], an explanatory dictionary and a translation of the English-language sentiment
lexicon was labelled by annotators. We considered positive and negative words
that received the majority of labels of the corresponding sentiment; contradictory
and neutral words were ignored.
4. RuSentiLex [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. For each word the sentiment (positive, negative, neutral) and the
source (opinion, fact, feeling) are indicated. To create this lexicon, first, lists of
sentiment words were generated based on the RuThes thesaurus, existing sentiment
lexicons, news articles and Twitter, then linguists analyzed the resulting lists to
form a final lexicon. We use the 2017 version and only words and combinations
with positive or negative sentiment.
5. SentiRusColl [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Russian sentiment lexicon of collocations. To create a lexicon a
corpus of reviews for ten domains was used, combinations of candidate words
were automatically selected from it, which were then labelled by three annotators.
        </p>
        <p>
          The lexicon contains the collocations that received the majority of votes.
6. Word Map [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Online thesaurus of words and expressions of the Russian
language. The sentiment lexicon developed within this project contains the Russian
words, supplied with label and strength of sentiment (positive, negative or neutral).
Crowdsourcing was applied to create the lexicon. We used positive and negative
words from the 2019 version of this lexicon.
7. Blinov’s lexicon [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. A manually compiled list of 969 most positive and 1,138
most negative words from the lexicon ProductSentiRus was automatically
expanded with synonyms and antonyms from the Russian Wiktionary.
8. Kotelnikov’s lexicon [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. An automatically selected list of words from five
domains was labelled by four annotators. We took words, on the sentiment of which
at least three out of four annotators agreed.
9. Tutubalina’s lexicon [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. A manually created lexicon based on strictly positive
and negative user reviews about cars has been expanded with synonyms.
        </p>
        <p>Each lexicon has been separately processed as follows:
─ neutral words have been removed;
─ all words have been converted to lower case;
─ words that are both positive and negative in the lexicon have been removed
(including the analyzed words with the spelling "е" and "ё");
─ words containing non-Cyrillic letters in the spelling have been removed;
─ one occurrence of each element has been left.</p>
        <p>The size of lexicon after preprocessing is given in Table 1.
At the next stage a combined sentiment lexicon was obtained from the preprocessed
lexicons. This lexicon includes words that are simultaneously found in at least four
source lexicons: 1,444 negative words (67%) and 712 (33%) positive words, a total of
2,156 words. At the same time only those words were left in which the sentiment was
clearly defined, i.e. there were no controversial cases. A controversial case was
considered when the number of lexicons in which a word was classified as negative
coincided with the number of lexicons in which it was positive. For the rest of the cases
the words were assigned to the prevailing class, i.e. the voting was used. Other values
were investigated for the minimum number of lexicons, but for four the best
performance scores were obtained on the training data.</p>
        <p>We created two versions of the combined lexicon – in the first one positive
elements were assigned a sentiment score of 3, negative ones – a score of –3. This
version of the lexicon is hereinafter referred to as CLex. In the second version of the
lexicon the sentiment score was determined by the number of source lexicons that
contained a given word. The second version of the lexicon is weighted and is further
called WCLex.</p>
        <p>Text Corpora. For experimental research the corpora of the SentiRuEval-2015§
competition were taken. The data consists of labelled reviews of restaurants and cars, as
well as tweets about banks and telecommunications companies. Corpora sizes are
very different: there are fewer reviews than tweets, but the average review length is
almost 10 times longer than the length of a tweet (830 characters versus 85 for
training data and 830 versus 82 for test data). This is due to the maximum possible tweet
length. The organizers of the competition provided training and test data. The training
sample consists of 403 reviews and 9,722 tweets. The size of the test sample is 403
reviews and 8,308 tweets (Table 2).
§ http://www.dialog-21.ru/evaluation/2015/sentiment/.</p>
        <p>Category
Training</p>
        <p>Test
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>Sentiment
Negative
Neutral
Positive
Total
Negative
Neutral
Positive
Total
Three models were tested – the lexicon-based methods SentiStrength (SS) and
SOCAL adapted for the Russian language, as well as the RuBERT deep learning model.</p>
      <p>For lexicon-based methods, two variants of the combined lexicon were used: CLex
and WCLex. For SO-CAL each corpus preprocessed with rnnmorph was evaluated on
both lexicons. For SentiStrength estimates were obtained on the original raw data and
on the preprocessed lemmatized data. For the RuBERT model two variants of text
corpora were used: corpora without preprocessing and corpora with replaces to good
and bad of words from sentiment lexicon.</p>
      <p>Two series of experiments were carried out. In the first series only texts with a
positive and negative sentiment were used, thus a binary classification was carried
out. The second series of experiments was carried out for a three-class classification –
texts with a neutral sentiment were also used.</p>
      <p>Table 3 shows the values of the macro F1-score metric on test data for binary
classification.</p>
      <p>Cars
0.8247
Among the lexicon-based methods, on average, for all experiments, SO-CAL showed
better results by 10 percentage points (pp), while in the classification of reviews, the
results were better for SentiStrength, tweets – for SO-CAL. The use of a weighted
lexicon for both methods sometimes even worsens, but on average not significantly
improves the performance (by about 9 pp on average, and by 28 pp maximum for
tweets about banks when using SentiStrength). In general, among the lexicon-based
methods the best result was shown by SO-CAL when using WCLex.</p>
      <p>The RuBERT model is superior to lexicon-based methods. On average for all
experiments its result exceeds the results of lexicon-based methods by 22 pp (in
comparison with the best lexicon-based method – by 15 pp). The experiment carried out to
replace words with subsequent fine-tuning of the RuBERT model showed a good
result only in one case – for reviews on cars, the quality increased by 2 pp, in other
cases such preprocessing deteriorated the performance.</p>
      <p>The performance of the classification of reviews is on average 32 pp higher than
the performance of tweet classification, the largest difference (by 59 pp) for
SentiStrength on lemmatized data with CLex.</p>
      <p>Table 4 shows the values of the macro F1-score metric on test data for a three-class
classification.
With a three-class classification among the lexicon-based methods on average for all
experiments again by 10 pp SO-CAL performed better, with better results for both
reviews and tweets. The use of a weighted lexicon for both methods has very little
effect on the performance. In general, among the lexicon-based methods the best
result was shown by SO-CAL when using WCLex.</p>
      <p>The best results for the three-class, as well as for the binary classification, are
shown by the RuBERT model. On average for all experiments its results exceeds the
results of lexicon-based methods by 16 pp (in comparison with the best lexicon-based
method – by 10 pp). Replacing words in the original data did not improve the result.
The performance of the classification of reviews and tweets on average differs
slightly (by 3.5 pp).</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>The experiments showed that the RuBERT deep learning model in all cases gets
better results than lexicon-based methods (compared to the best SO-CAL lexicon-based
model with a weighted lexicon – by 15 pp for a binary and 10 pp for a three-class
classification). However, the results of the neural network model are difficult to
interpret, and the lexicon-based methods provide detailed information about the sentiment
words and expressions found in the text. This allows, in particular, analyzing the
errors that these methods made.</p>
      <p>The analysis showed that for reviews the most common causes of errors are, for
example, the following: incorrect search for negation, search for not all sentiment
words, prevalence of vocabulary of the opposite sentiment, a problem of the third
class, when sentiment 0 is used not only for neutral texts, but for contradictory ones
also. There are also two other common types of errors for tweets: lack of knowledge
of the context and incomplete phrases, which are related to the specifics and
limitation of the number of characters in one text message.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The use of the RuBERT deep learning model has a higher performance (by 22 pp for
two-class classification and by 16 pp for three-class classification) compared to the
adapted versions of the SO-CAL and SentiStrength lexicon-based methods, but its
results are difficult to interpret. The ability to analyze errors allows you to identify
ways of improving lexicon-based methods.</p>
      <p>Adding information from the lexicon during preprocessing data for RuBERT led to
an improvement in the result only in one case out of eight.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.: Sentiment</given-names>
          </string-name>
          <string-name>
            <surname>Analysis: Mining Opinions</surname>
          </string-name>
          , Sentiments, and Emotions. Cambridge: Cambridge University Press (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Poria</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hazarika</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mihalcea</surname>
          </string-name>
          , R.:
          <article-title>Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis Research</article-title>
          . In: Computing Research Repository, arXiv:
          <year>2005</year>
          .
          <volume>00357</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Sentiment Analysis: An Overview from Linguistics</article-title>
          .
          <source>In: Annual Review of Linguistics</source>
          ,
          <volume>2</volume>
          ,
          <fpage>325</fpage>
          -
          <lpage>347</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Thelwall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al:
          <article-title>Sentiment strength detection in short informal text</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>61</volume>
          (
          <issue>12</issue>
          ),
          <fpage>2544</fpage>
          -
          <lpage>2558</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Emotion Measurement (Second Edition)</surname>
          </string-name>
          , Chapter 11 -
          <article-title>Sentiment analysis: Automatically detecting valence, emotions, and other affectual states from text</article-title>
          , edited by H.L. Meiselman, Woodhead Publishing,
          <fpage>323</fpage>
          -
          <lpage>379</lpage>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kotelnikov</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          et al:
          <article-title>Modern sentiment lexicons for opinion mining in English and Russian (analytical survey)</article-title>
          .
          <source>Nauchno-tekhnicheskaya informaciya</source>
          ,
          <volume>12</volume>
          ,
          <fpage>16</fpage>
          -
          <lpage>33</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al:
          <article-title>Lexicon-based methods for sentiment analysis</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>37</volume>
          (
          <issue>2</issue>
          ),
          <fpage>267</fpage>
          -
          <lpage>307</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al:
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
          <source>Google AI Language</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kuratov</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arkhipov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language</article-title>
          . MIPT,
          <year>Dialogue 2019</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruder</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Universal Language Model Fine-tuning for Text Classification</article-title>
          .
          <source>In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          , Melbourne, Australia.
          <source>Association for Computational Linguistics</source>
          ,
          <fpage>328</fpage>
          -
          <lpage>339</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alì</surname>
          </string-name>
          , G.:
          <article-title>What's Inside the Black Box? AI Challenges for Lawyers and Researchers</article-title>
          .
          <source>Legal Information Management</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <fpage>2</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kotelnikov</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          et al:
          <article-title>A comparative study of publicly available Russian sentiment lexicons</article-title>
          .
          <source>- Communications in Computer and Information Science: 7th conference on Artificial Intelligence and Natural Language (AINL-2018)</source>
          ,
          <fpage>139</fpage>
          -
          <lpage>151</lpage>
          . Springer (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mohammad</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turney</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          :
          <article-title>Crowdsourcing a word-emotion association lexicon</article-title>
          .
          <source>Computational Intelligence</source>
          ,
          <volume>29</volume>
          (
          <issue>3</issue>
          ),
          <fpage>436</fpage>
          -
          <lpage>465</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skiena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Building Sentiment Lexicons for All Major Languages</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <fpage>383</fpage>
          -
          <lpage>389</lpage>
          ,
          <string-name>
            <surname>Baltimore</surname>
          </string-name>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Koltsova</surname>
            ,
            <given-names>O.Yu.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexeeva</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolcov</surname>
            ,
            <given-names>S.N.:</given-names>
          </string-name>
          <article-title>An Opinion Word Lexicon and a Training Dataset for Russian Sentiment Analysis of Social Media</article-title>
          .
          <source>In: Computational Linguistics and Intellectual Technologies: Papers from the Annual International Conference “Dialogue-2016”</source>
          ,
          <volume>15</volume>
          (
          <issue>22</issue>
          ),
          <fpage>277</fpage>
          -
          <lpage>287</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Chetviorkin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loukachevitch</surname>
          </string-name>
          , N.:
          <article-title>Extraction of Russian Sentiment Lexicon for Product Meta-Domain</article-title>
          .
          <source>In: Proceedings of COLING</source>
          <year>2012</year>
          ,
          <volume>593</volume>
          -
          <fpage>610</fpage>
          , Mumbai (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Loukachevitch</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levchik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Creating a General Russian Sentiment Lexicon</article-title>
          .
          <source>In: Proceedings of Language Resources and Evaluation Conference LREC-2016</source>
          ,
          <fpage>1171</fpage>
          -
          <lpage>1176</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kotelnikova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kotelnikov</surname>
          </string-name>
          , E.:
          <article-title>SentiRusColl: Russian Collocation Lexicon for Sentiment Analysis</article-title>
          .
          <source>In: Artificial Intelligence and Natural Language. AINL 2019. Communications in Computer and Information Science</source>
          ,
          <volume>1119</volume>
          ,
          <fpage>18</fpage>
          -
          <lpage>32</lpage>
          , Springer, Cham (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Kulagin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Russian Word Sentiment Polarity Dictionary: a Publicly Available Dataset</article-title>
          .
          <source>Poster in: Artificial Intelligence and Natural Language. AINL</source>
          <year>2019</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Blinov</surname>
          </string-name>
          , P.D. et al:
          <article-title>Research of lexical approach and machine learning methods for sentiment analysis</article-title>
          .
          <source>In: Computational Linguistics and Intellectual Technologies: Papers from the Annual International Conference “Dialogue-2013”</source>
          ,
          <volume>12</volume>
          (
          <issue>19</issue>
          ).
          <fpage>51</fpage>
          -
          <lpage>61</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kotelnikov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          et al:
          <article-title>Manually Created Sentiment Lexicons: Research and Development</article-title>
          .
          <source>In: Computational Linguistics and Intellectual Technologies: Papers from the Annual International Conference “Dialogue-2016”</source>
          ,
          <volume>15</volume>
          (
          <issue>22</issue>
          ),
          <fpage>300</fpage>
          -
          <lpage>314</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Tutubalina</surname>
            ,
            <given-names>E.V.</given-names>
          </string-name>
          :
          <article-title>Extraction and summarization methods for critical user reviews of a product</article-title>
          . Kazan Federal University, Kazan, Russia (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>