<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Word2Vec Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ekaterina Popova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Spitsyn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research Tomsk Polytechnic University</institution>
          ,
          <addr-line>30, Lenin Avenue, Tomsk, 634050</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article is devoted to modern approaches for sentiment analysis of short Russian texts from social networks using deep neural networks. Sentiment analysis is the process of detecting, extracting, and classifying opinions, sentiments, and attitudes concerning different topics expressed in texts. The importance of this topic is linked to the growth and popularity of social networks, online recommendation services, news portals, and blogs, all of which contain a significant number of people's opinions on a variety of topics. In this paper, we propose machine-learning techniques with BERT and Word2Vec embeddings for tweets sentiment analysis. Two approaches were explored: (a) a method, of word embeddings extraction and using the DNN classifier; (b) refinement of the pre-trained BERT model. As a result, the finetuning BERT outperformed the functional method to solving the problem.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural language processing</kwd>
        <kwd>sentiment analysis</kwd>
        <kwd>deep learning</kwd>
        <kwd>BERT</kwd>
        <kwd>CNN</kwd>
        <kwd>LSTM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>Sentiment analysis is a technique by which one can analyze a piece of text to determine the sentiment behind it. It combines machine learning and natural language processing (NLP) to achieve this. Using basic Sentiment analysis, a program can understand if the sentiment behind a piece of text is positive, negative, or neutral.</title>
      </sec>
      <sec id="sec-1-2">
        <title>The relevance of this topic is associated with the development and growing popularity of social networks, online recommendation services, news portals and blogs, where a large number of people's opinions on various issues are collected.</title>
        <p>

</p>
      </sec>
      <sec id="sec-1-3">
        <title>There are three major types of algorithms used in sentiment analysis:</title>
      </sec>
      <sec id="sec-1-4">
        <title>Rule-based systems automatically perform sentiment analysis based on a set of manually</title>
        <p>crafted rules. This approach to sentiment analysis involves searching for keywords in the text and
matching each of them with a numeric value in a dictionary or associative array.</p>
      </sec>
      <sec id="sec-1-5">
        <title>Automatic methods, contrary to rule-based systems, do not rely on manually crafted rules, but</title>
        <p>on machine learning techniques. A sentiment analysis task is usually modeled as a classification
problem, whereby a classifier is fed a text and returns a category, e.g. positive, negative, or neutral.</p>
      </sec>
      <sec id="sec-1-6">
        <title>Hybrid systems combine the desirable elements of rule-based and automatic techniques into</title>
        <p>one system. One huge benefit of these systems is that results are often more accurate.</p>
      </sec>
      <sec id="sec-1-7">
        <title>The aim of this work is to study modern neural network approaches to automatically determining the sentiment of short texts in Russian using the example of data from the social network Twitter.</title>
        <p>2021 Copyright for this paper by its authors.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Data set</title>
      <sec id="sec-2-1">
        <title>Data analysis from social networks is one of the popular research areas in the field of sentiment analysis. Social media contain huge amount of the sentiment information in the form of tweets, blogs, and updates on the status, posts, etc. Therefore, in this work, messages from the social network Twitter will be used as data.</title>
      </sec>
      <sec id="sec-2-2">
        <title>Twitter is one of the most popular social media platforms in the world, with 330 million monthly active users and 500 million tweets sent each day. By carefully analyzing the sentiment of these tweets – whether they are positive, negative, or neutral, for example – we can learn a lot about how people feel about certain topics.</title>
      </sec>
      <sec id="sec-2-3">
        <title>However, the task is complicated by the fact that Twitter messages contain a large number of slang</title>
        <p>words and misspellings and repeated characters. Maximum length of each tweet in Twitter is 140
characters. Therefore, it is very important to identify correct sentiment of each word.</p>
      </sec>
      <sec id="sec-2-4">
        <title>The corpus of short texts by Yulia Rubtsova (RuTweetCorp) [1], formed based on Russian-language messages from the social network Twitter, was chosen as a training sample this work. The corpus was automatically labeled based on emoticons and contains 114,991 positive and 111,923 negative messages.</title>
        <p>The peculiarity of the dataset is that RuTweetCorp was initially designed for the creation of a
sentiment lexicon, not for sentiment classification. The dataset was collected automatically, i.e. each
text was associated with the sentiment class based on the emoticons it contained. Therefore, even a
simple rule-based approach is able to demonstrate good results. So that, the authors recommend
removing emoticons from the dataset at the preprocessing stage in order to solve the problem of
automatic sentiment analysis. Thus, when analyzing the literature, only those articles were taken into
account in which the preprocessing procedure included the removal of emoticons.</p>
      </sec>
      <sec id="sec-2-5">
        <title>In paper [7] compared logistic regression, XGBoost classifier and Convolutional Neural Network</title>
        <p>
          on RuTweetCorp for binary classification and achieved F1 = 78.1% using a convolutional neural
network. The authors of [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] tested the possibilities of fine-tuning Multilingual BERTBase, RuBERT
and two versions of Multilingual USE for RuTweetCorp. The best results in binary classification were
F1 = 83.69% for RuBERT.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Text pre-prossessing</title>
      <sec id="sec-3-1">
        <title>The text classification results depend on the input text preprocessing quality. In this work all the text were reduced to lower case, punctuation marks were removed, links and usernames were replaced with "URL" and "USER ", respectively, and the word "RT", denoting retweet, is also removed at the preprocessing stage</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Word embedding</title>
      <sec id="sec-4-1">
        <title>The main idea of word embedding is to match each word with some numerical vector of fixed dimension. Vectors are constructed in such a way that words found in similar contexts have similar vector representations. Word embeddings are mostly used as input features for other models built for custom tasks.</title>
      </sec>
      <sec id="sec-4-2">
        <title>An important advantage of word embedding is that no tagged data is required to train them.</title>
      </sec>
      <sec id="sec-4-3">
        <title>Attachments are retrieved from a very large unmarked enclosure. Pretrained word embeddings are</title>
        <p>available online and can be used by researchers around the world for a variety of NLP tasks.</p>
      </sec>
      <sec id="sec-4-4">
        <title>There are two main approaches to generate word embeddings:</title>
        <p> Context-independent (Bag of Words, TF-IDF, Word2Vec, GloVe);
 Context-aware (ELMo, Transformer, BERT, Transformer-XL).</p>
        <p>In this work, word embeddings based on Word2Vec and BERT will be investigated:
 Word2Vec models generate embeddings that are context-independent. There is just one vector
(numeric) representation for each word. Different senses of the word (if any) are combined into one
single vector. Word2Vec embeddings do not take into account the word position in the sentence.
 BERT model generates embeddings that allow us to have multiple (more than one) vector
(numeric) representations for the same word, based on the context in which the word is used. BERT
model explicitly takes as input the position (index) of each word in the sentence before calculating
its embedding. Thus, BERT embeddings are context-dependent.
4.1.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Word2Vec embedding</title>
      <sec id="sec-5-1">
        <title>Word2Vec is a set of algorithms for calculating vector representations of words, developed by a</title>
        <p>
          group of researchers from Google in 2013 [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Word 2Vec implements two main architectures CBOW
(Continuous Bag of Words) and Skip-gram. CBOW predicts a word based on the current context, and
skip-gram, on the contrary, predicts a context based on the current word.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>In this work, the Skip-gram architecture was chosen, which is more suitable for small corpora of</title>
        <p>
          texts (less than one hundred million words), since it better takes into account rare words [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>To train the word2vec model an implementation from the Gensim library was chosen. The model</title>
        <p>was trained with the following hyperparameters:
 Dimension of vector space – 200;
 Scanning window size – 10.</p>
      </sec>
      <sec id="sec-5-4">
        <title>The model was trained on 17.5 million unlabeled Russian-language Twitter posts. The size of the formed dictionary was 712,991 words. Figure 1 shows the visualization of clusters of similar words of the trained Word2Vec model.</title>
        <p>4.2.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>BERT embedding</title>
      <p>
        BERT (Bidirectional Encoder Representations from Transformers) is a neural network developed
by Google researchers in 2018 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which has shown high results on a number of NLP problems. BERT
is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning
on both left and right context. As a result, the pre-trained BERT model can be fine-tuned with just one
additional output layer to create state-of-the-art models for a wide range of NLP tasks. The BERT model
can be used in two ways:
 To generate word attachments, which are further used as input for DNN classifiers;
 To fine-tune a pretrained BERT model for a specific task.
      </p>
      <sec id="sec-6-1">
        <title>In this work, we use the BERT model for the Russian language Conversational RuBERT from</title>
      </sec>
      <sec id="sec-6-2">
        <title>DeepPavlov [5]. The model has 12 stacked transformer encoder layers, with 12 attention heads. The embedding dimension is 768, 180M parameters.</title>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Feature-based approaches</title>
      <sec id="sec-7-1">
        <title>For feature-based approaches, were used pre-trained Word2Vec and BERT models to obtain the</title>
        <p>
          sequence of embeddings for a given messages:
 Word2Vec model: Were used pre-trained Word2Vec model and apply this model to generate
one embedding for each word of comment.
 BERT model: Word-piece tokenization is performed on the comment and then used as input
to a pre-trained BERT model. BERT model provides contextual embedding for the word-pieces [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          The obtained embeddings from both Word2Vec and BERT models are then used as input to a DNN
classifier. Deep learning models (CNN, LSTM, GRU) are used as DNN classifier:
 CNN (Convolutional Neural Network) has traditionally been used in image processing
applications, but has recently become actively applied to various NLP tasks. For example, the article
[
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] demonstrates the effective use of CNN for natural language processing on various test data.
 LSTM (Long short-term memory) is a special kind of recurrent neural network capable of
handling long-term dependencies. LSTM is an advanced RNN [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], a sequential network that allows
information to persist. It is capable of handling the vanishing gradient problem faced by RNN.
 GRU (Gated recurrent units) are a gating mechanism in recurrent neural networks, introduced
in 2014. The GRU [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is like a long short-term memory (LSTM) with a forget gate, but has fewer
parameters than LSTM, as it lacks an output gate.
        </p>
        <p>The configurations of the DNN models that were used in this work are presented below:
 LSTM and GRU: One LSTM (GRU) layer with 128 internal nodes followed by an output
layer containing one neuron for predicting values. It uses a sigmoid activation function to produce a
probabilistic result ranging from 0 to 1, which can be converted to a clear class value.
 СNN (Conv1D): Two CNN layers, with 128 and 64 filters respectively, and the width of the
window for filters is 5, activation function is ReLU, window step size (stride) is 1. The output layer
contains one neuron with a sigmoid activation function.</p>
      </sec>
      <sec id="sec-7-2">
        <title>We use a varying dropout up to 0.2. The models are trained using Adam optimizer with learning rate of 0.001.</title>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>BERT fine-tuning</title>
      <sec id="sec-8-1">
        <title>The BERT pre-trained model can be fine-tuned to a specific task. This consists in the adapting of</title>
        <p>the pre-trained BERT model parameters to a specific task using a small corpus of task specific data.</p>
      </sec>
      <sec id="sec-8-2">
        <title>For the purpose of classification task, a neural network layer is used on top of fine-tuned BERT model.</title>
      </sec>
      <sec id="sec-8-3">
        <title>So, the weights of this layer and the weights of the other layers of the Bert model are trained and fine</title>
        <p>tuned correspondingly using task specific data in order to perform the classification task.</p>
      </sec>
      <sec id="sec-8-4">
        <title>For BERT fine-tuning we used Adam optimizer, with an initial learning rate of 2e–5, batch size of 32 and 3 epochs of training.</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Results and discussion</title>
      <sec id="sec-9-1">
        <title>To assess the quality of the obtained classification results, generally accepted metrics were used:</title>
      </sec>
      <sec id="sec-9-2">
        <title>Recall, Precision, F – measure, Accuracy. The following auxiliary parameters were calculated:</title>
        <p> TP is the number of true positive results;


</p>
      </sec>
      <sec id="sec-9-3">
        <title>TN is the number of true negative results;</title>
      </sec>
      <sec id="sec-9-4">
        <title>FP is the number of false positive results;</title>
      </sec>
      <sec id="sec-9-5">
        <title>FN is the number of false negative results.</title>
      </sec>
      <sec id="sec-9-6">
        <title>Precision can be interpreted as the proportion of objects called positive by the classifier and at the same time really positive, and recall shows what proportion of objects of a positive class from all objects of a positive class the algorithm found. The higher the recall and precision values, the better the classification result will be.</title>
      </sec>
      <sec id="sec-9-7">
        <title>While recall and precision are very important metrics, they will not tell the whole story on their own. One way to summarize them is F – measure, a measure that is the harmonic mean of precision and completeness:</title>
        <p>(1)
(2)
(3)
(4)
 −</p>
      </sec>
      <sec id="sec-9-8">
        <title>This version of calculating the F – measure is also known as the F1 – measure. Since the F1 – measure takes into account precision and completeness, it may be a more appropriate metric for binary classification of unbalanced data.</title>
      </sec>
      <sec id="sec-9-9">
        <title>Accuracy is the proportion of right classified objects in the all classified objects:</title>
      </sec>
      <sec id="sec-9-10">
        <title>The results of classifiers efficiency evaluation are presented in the Table 1. As we see, approaches base on BERT fine-tuning show a significant gap in F1 compared to the feature-based approaches.</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Conclusion</title>
      <p>Accuracy
Precision
75,37
78,01
78,58
78,12
78,98
82,22
75,39
78,49
78,66
78,13
78,97
82,25
Recall
75,34
78,08
78,61
78,10
78,98
82,19
F1-measure
75,35
77,94
78,57
78,11
78,98
82,61</p>
      <sec id="sec-10-1">
        <title>The article compares two approaches to automatically determining the sentiment of messages on</title>
        <p>social networks. These approaches are based on deep learning classifiers and word embeddings.</p>
      </sec>
      <sec id="sec-10-2">
        <title>The combination of feature-based approaches and fine-tuning of pre-trained BERT model were proposed. In feature-based approaches, Word2Vec and BERT embeddings were used as input features to GRU and LSTM classifiers. Further, we have compared these configurations with fine-tuning of pretrained BERT model.</title>
        <p>attachments.</p>
      </sec>
      <sec id="sec-10-3">
        <title>As a result, BERT fine-tuning turned out to be more efficient than the functional approach to solve the problem. In the future, it is planned to explore the possibility of combining Word2Vec and BERT</title>
      </sec>
      <sec id="sec-10-4">
        <title>The reported study was funded by the Russian Foundation for Basics Research under RFBR research</title>
        <p>project No. 18-08-00977 А and supported by Tomsk Polytechnic University Competitive-ness</p>
      </sec>
      <sec id="sec-10-5">
        <title>Enhancement Program.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rubtsova</surname>
          </string-name>
          ,
          <article-title>Constructing a corpus for sentiment classification training</article-title>
          ,
          <source>Softw. Syst</source>
          .
          <volume>109</volume>
          (
          <year>2015</year>
          )
          <fpage>72</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Efficient Estimation of Word Representations in Vector Space</article-title>
          , arXiv:
          <fpage>1301</fpage>
          .3781 (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Rong</surname>
          </string-name>
          , word2vec Parameter Learning Explained, arXiv:
          <fpage>1411</fpage>
          .2738 (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , BERT:
          <article-title>Pretraining of Deep Bidirectional Transformers for Language Understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</article-title>
          , Volume
          <volume>1</volume>
          (Long and Short Papers),
          <year>2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] DeepPavlov's documentation: BERT in DeepPavlov, 2021</article-title>
          . URL: http://docs.deeppavlov.ai/en/master/features/models/bert.html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Convolutional Neural Networks for Sentence Classification</article-title>
          ,
          <source>in: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>1746</fpage>
          -
          <lpage>1751</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zvonarev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bilyi</surname>
          </string-name>
          ,
          <article-title>A comparison of machine learning methods of sentiment analysis based on russian language twitter data</article-title>
          ,
          <source>in: The 11th Majorov International Conference on Software Engineering and Computer Systems</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Smetanin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Komarov</surname>
          </string-name>
          ,
          <article-title>Deep transfer learning baselines for sentiment analysis in Russian</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>58</volume>
          (
          <issue>3</issue>
          ) (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Understanding</surname>
            <given-names>LSTM Networks</given-names>
          </string-name>
          ,
          <year>2015</year>
          . URL: http://colah.github.io/posts/2015-08-UnderstandingLSTMs.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Illustrated Guide to LSTM's and GRU's: A step by step explanation</article-title>
          ,
          <year>2018</year>
          . URL: https://towardsdatascience.com/illustrated-guide
          <article-title>-to-lstms-and-gru-s-a-step-by-stepexplanation44e9eb85bf21</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>BERT</given-names>
            <surname>Word Embeddings</surname>
          </string-name>
          <string-name>
            <surname>Tutorial</surname>
          </string-name>
          ,
          <year>2019</year>
          . URL: https://mccormickml.com/
          <year>2019</year>
          /05/14/BERTword-embeddings-tutorial.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>