<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Moscow Myasnitskaya.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia eklyshinsky@hse.ru</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Keldysh Institute of Applied Mathematics</institution>
          ,
          <addr-line>Miusskaya sq., 4 Moscow, 125047</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Authors of Word2Vec claimed that their technology could solve the word analogy problem using the vector transformation in the introduced vector space. However, the practice demonstrates that it is not always true. In this paper, we investigate several Word2Vec and FastText model trained for the Russian language and find out reasons of such inconsistency. We found out that different types of words are demonstrating different behavior in the semantic space. FastText vectors are tending to find phonological analogies, while Word2Vec vectors are better in finding relations in geographical proper names. However, we found out that just four out of fifteen selected domains are demonstrating accuracy more that 0.8. We also draw a conclusion that in a common case, the task of word analogies could not be solved using a random word pair taken from two investigated categories. Our experiments have demonstrated that in some cases the length of the vectors could differ more than twice. Calculation of an average vector leads to a better solution here since it closer to more vectors.</p>
      </abstract>
      <kwd-group>
        <kwd>Word Embeddings</kwd>
        <kwd>Vector Space</kwd>
        <kwd>Vector Transformation</kwd>
        <kwd>Word Analogies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The basic point for the semantic space of natural language words was the paper [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
published in 2003. It introduced fixed-size vectors (embeddings) generated by a neural
network using statistical information about the word context. This concept was
developed in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where the author demonstrated that such pre-trained vectors can be useful
for solution of different problems of natural language processing. The Word2Vec mode
* Publication financially supported by RFBR grant №18-08-01484
based on the distributive hypothesis was introduced in 2013 [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3-5</xref>
        ]. The Word2Vec
model also uses neural networks reinforced by several new ideas. First of all, the new
approach [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] had less computational complexity compared to the previous systems. The
next article [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] increases its learning rate and accuracy. And the greatest contribution
of the authors was the publication of source codes and pre-trained language models for
free use. However, the authors of these papers processed texts written only in English.
The transfer of these models into inflectional languages meets some problems because
of the large variety of tokens of the same lemma. The statistical distribution of word
forms of the given lemma is not always uniform. As a result, the model needs bigger
corpora to collect good statistics for rare forms. That is why researchers use
lemmatization for inflectional languages. However, some of the problems need information for
individual word forms, what brings us back to the problem of under-tuned models.
      </p>
      <p>
        In order to keep the important lexical information, the FastText model was
introduced in 2017 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The authors of this model keep the following idea. Both prefixes and
postfixes of words carry semantic information as well as word roots. In this case, the
meaning of a word can be composed from the meaning of its parts. Dividing a word
into n-grams, the system collects more information about the same n-gram using
contexts of different words. The authors of [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3-5</xref>
        ] claimed that the new semantic space
allows the vector arithmetic. Their example “Queen = King–man + woman” swiftly
becomes very famous. However, it becomes clear in a short time that such operations do
not always lead us to success. One of the proofs of this concept is the problem of words
analogies. The early experiments demonstrated that another favorite example, countries
and their capitals, does not work correctly for any case – country, capital and pre-trained
language model. The accuracy of this analogy was pretty high but not enough to state
that vector arithmetic works properly. The Word2Vec and FastText models correctly
find the list of semantic neighbors for a given word, what makes it a crucial part of
modern systems of natural language processing. However, the problem of words
analogies does not work as well as it could be.
      </p>
      <p>
        In this paper, we investigate the reasons of such deviations in accuracy. In the
Section 2, we state the problem of word analogies as a vector transformation problem. The
Section 3 gives a short review of existing vector transformation methods for the
problem of word analogies. Sections 4 and 5 describe the used data set and the numerical
evaluation of free language models for the Russian language. The Section 6 analyses
the reasons of low accuracy for some categories of word analogies. The Section 7
concludes the article. In this paper, we do not consider systems of the BERT and ELMO
families [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. The correct investigation of such systems needs a slight correction of
our method and will be conducted in the nearest future. The description of
contextualized words embedding could be found in paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Formal Statement of the Problem of Word Analogies</title>
      <p>In common words, the main question of the problem of word analogies could be stated
as “Is there a word c which relates to the word b as the word a' relates to the word a?”
Answering this question, Word2Vec uses vector representation of the word. Let va' and</p>
      <p>We can reformulate the question for word groups. Let us consider a set of word pairs
(w11:w12), (w21:w22), …, (wN1:wN2) that have the same semantic or lexical relation, and
their corresponding vectors v11, v11, v21, v21, …, vN1, vN1. In this case, the task of word
analogies could be formulated as following: if there is a vector x that makes an affine
transformation of w11, w21, …, wN1 to w12, w22, …, wN2, then x is such that
  2 = 
 ′</p>
      <p>⁡( ′,   1 +  ),</p>
      <p>Let us denote the fact that the word a relates to the word b in the same sense as the
word c relates to the word d by the following equation: (a:b) :: (c:d). For example,
(apple:fruit) :: (cucumber:vegetable), (apple:apples) :: (cucumber:cucumbers), and,
classical, (king:queen) :: (man:woman). In this case, a request to find an analogy looks like
(king:?) :: (man:woman) or (man:woman) :: (king:?).</p>
      <p>=   +   ′ −   .
 ′ =  
 ∉{  ,  ′,  }
⁡( ,   +   ′ −   )
(1)
(2)
(3)
(4)</p>
      <p>Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings 3
va be vectors corresponding to the words a' and a respectively; in this case, the vector
difference va' - va expresses the semantic relation (or in other words, the semantic
difference) between the words a and a'. Thus, in order to find an analogue, we should find
the word x and its corresponding vector y such that y - vb = va' - va, or</p>
      <p>However, the probability of existence of a word having exactly the same vector as
vx is extremely small. That is why Word2Vec finds vector y' that is the closest word to
the vector y:
3</p>
    </sec>
    <sec id="sec-3">
      <title>Review of Affine Transformation Methods for the Problem of</title>
    </sec>
    <sec id="sec-4">
      <title>Word Analogies</title>
      <p>
        As it was mentioned above, the Word2Vec system uses an algorithm that finds the
vector nearest to the calculated one [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]:
 ′ =  
 ∉{  ,  ′,  }
      </p>
      <p>⁡( ,   +   ′ −   ).</p>
      <p>
        Such model is called 3CosAdd. As it was demonstrated in [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref14 ref15">10-15</xref>
        ], the 3CosAdd
method has crucial drawbacks. The accuracy of this method varies depending on the
word set or the trained model [
        <xref ref-type="bibr" rid="ref10 ref11 ref15">10, 11, 15</xref>
        ]. This method is not applicable for such tasks
as analogies of synonyms and antonyms [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. Such variety could be explained by
the drawbacks of trained models and text corpora; however, the paper [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] states that
the reason is the relative ordering of vectors in the model space. Moreover, the
3CosAdd method returns just one vector while in such tasks as “whole-part” there could be
a set of response vectors [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        That is the why researchers have to find other methods to solve the problem of word
analogies. One of them is the Only-b [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] method that takes a nearest neighbor of vb:
The PairDirection [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] mehod supposes that va' - va and vb' - vb are collinear:
 ∉{  ,  ′,  }
 ′ =
      </p>
      <p>⁡( −   ,   ′ −   ).</p>
      <p>
        The authors of 3CosMul method show in the paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that formula (2) is equal to
 ∉{  ,  ′,  }
 ′ = 
(
( ,   ′) − 
( ,   ) + 
( ,   )).
      </p>
      <p>The authors use an example (London:England)::(Baghdad:?). The component cos(v,
vEngland) responses to the similarity of the answer with the country, the component cos(v,
vBaghdad) responses to the cultural and geographical similarity, and the component cos(v,
vLondon) responses to the cultural and geographical difference. The correct answer is
Iraq, but the component cos(v, vBaghdad) dominates all other components and the
resulting answer becomes Mosul. In order to eliminate such domination, the authors
introduce a new formula that keeps balance among the various components:
 ′ = 
 ∉{  ,  ′,  }
 ( ,  ′)

( ,  )+
( ,  ),</p>
      <p>We suppose here that if va and vb in (1) are very close, then va - vb = 0 and y = vb.</p>
      <p>
        The next method is Ignore-a [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] that takes the vector closest to the sum of va' and
vb, i.e., finds the vector located between va' and vb:
(5)
(6)
(7)
(8)
(9)
 ∉{  ,  ′,  }
 ′ =
      </p>
      <p>∉{  ,  ′,  }
 ′ =</p>
      <p>⁡( ,   ′ +   )</p>
      <p>
        As it was shown in [
        <xref ref-type="bibr" rid="ref10 ref16">10, 16</xref>
        ], methods Only-b, Ignore-a, and PairDirection give
unsatisfactory results. In the next part of our article we will compare the results of two
methods (3CosAdd and 3CosAvg) shown on different models trained in Russian.
where ε = 0.001 is a small constant eliminating division by zero.
      </p>
      <p>
        The task of word analogies is very sensitive to the noise in the input data. A word
can be homonymous; this means that it should be presented as two or more separate
vectors representing different meanings of this word. In case of Word2Vec, such a word
will be represented only by a vector that will be a superposition of all its meanings.
Moreover, the resulting vectors of similar entities could express differences in their
occurrence with other words. For example, a dog and a cow are both animals, but dog
is a carnivore and a human’s friend, while a cow is an herbivore and gives milk; thus,
the analogy is not complete here. In order to eliminate such influence, the authors of
the 3CosAvg method [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] introduce a new formula that takes into account not just a
pair of words but whole two groups having the same analogy:
 ′ = 
 ∈ ∖{  }
      </p>
      <p>(  ,   + ∑ =1   ′ − ∑ =1    )</p>
      <p>Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings 5
4</p>
    </sec>
    <sec id="sec-5">
      <title>Used Data Sets</title>
      <p>We used several pre-trained models from the site RusVectōrēs
(http://rusvectores.org/ru/models/): Araneum Upos Skipgram 2018, Ruwikiruscorpora
Upos Skipgram 2018, Ruwikiruscorpora Upos Skipgram 2019, Tayga Upos Skipgram
2019, News Upos Skipgram 2019, Ruscorpora Upos CBOW 2019, Araneum Fasttext
Skipgram 2018, Facebook FastText CBOW 2018 [29, 30]. The first models were
trained using Word2Vec, the two later ones were trained using FastText.</p>
      <p>
        For semantic analogies, we used the Russian versions of Google analogy test set [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
and BATS (The Bigger Analogy Test Set) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. For grammatical analogies we used
morphological dictionary of the Russian language. The list of used categories is
presented in Table 1.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Evaluation</title>
      <p>For the purpose of evaluation, we calculated the accuracy metrics for all category
examples for the given language model. Fig.1 demonstrates the results for the 3CosAdd
method; Fig.2 demonstrates the results for the 3CosAvg method. Dark blue shows
results with higher accuracy, up to 1; light blue shows results closer to zero.</p>
      <p>The best results for the 3CosAdd method were 0.9 for the category Country →
Adjective taught at the Russian Wikipedia. Note that this category demonstrates the best
results for any used language model. It is not surprising that Country categories achieve
better results on Wikipedia text model since there is more information on this topic.</p>
      <p>The worst results are demonstrated at reflective verbs (A14) and verbs with prefixes
(A15). In our experiments, we included not only the initial forms of the verb but a
variety of forms: the past tense, the imperative mood, etc. That is why the resulting
affine transfer vector expresses variety of characteristics, but not the only one, and
failed to find one preferential direction.</p>
      <p>Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings 7
Some words were rare in the learning corpus, and the model failed to represent a
solid vector representation for these words. That is true, for example, for the category
A10 Possessive Adjective → Comparative Adjective.</p>
      <p>The FastText model works better for categories with flexies: A8, A9, A11, A12,
A13. Originally, the FastText model was created to work with character n-grams and
learn grammatical features of the language. So, FastText is better situated to learn the
‘sense’ of prefixes and affixes.</p>
      <p>The results for 3CosAvg are much better. The category A1 Famous capital →
Country demonstrates 100% accuracy on Wikipedia data. However, the results for some
categories are not so impressive. The categories A14 and A15 still demonstrate the same
results by the same reasons. The category A3 Country → currency cannot be solved
better than 37%. We can suppose that the reason is that different countries have the
same name for their currency, but their vectors are different; thus, their transition
vectors should be different as well. Finally, the category A7 Singular → Plural faces to the
variety of Russian homonymous postfixes denoting several tags.
6</p>
    </sec>
    <sec id="sec-7">
      <title>Data Analysis</title>
      <p>In order to find out the reasons of success and fail, we conducted a visual analysis of
the Word2Vec and FastText vectors. First of all, we projected 300-dimensional vectors
into 2D-space using Principal Component Analysis (PCA). Instead of t-SNE and
UMAP, PCA does not create areas with non-linear skewing. As a result, parallel vectors
keep their parallelism. On the other hand, there is a non-zero probability that two
vectors on parallel planes could become parallel on the projection. The later case makes
some distortion in the data, but not as critical as the former one.</p>
      <p>For our experiments, we used two pre-trained language models: Araneum Upos
Skipgram 2018 and Araneum Fasttext Skipgram 2018. They were trained on the same
Araneum corpus. We randomly selected five word pairs for several categories presented
in Fig. 3-8. All vectors are directed from a blue to a red point.</p>
      <p>Fig. 3-8 helps us make the following conclusion. The main reason of low accuracy
in the task of word analogies is the bias between vectors. It is easy to see that Word2Vec
vectors in Fig. 3 are mostly parallel, excluding slightly bias for Рим-Италия
(RomeItaly). The length of the vectors is also almost the same. Averaging among the
beginning and ending points of vectors helps adjusting these small biases in the Word2Vec
model. However, the FastText model is oriented mostly on word parts; that is why its
vectors are almost randomly oriented and have no preferential direction. The same is
true for Fig. 4. However, in Fig. 6 the situation is opposite. The difference between the
starting and finishing points here is the prefix не- (non-, ir-). The FastText model, which
vectors are parallel and have quite the same length, processes such situation better than
Word2Vec. Quite the same is true for the category A13 Verb → Noun with -тель (-er,
-or) (Fig. 8) where the FastText model achieves much better results.</p>
      <p>The worst situation is shown in Fig. 5 and Fig. 7 for prefix при- and reflexive verbs,
which differ from the main form by the suffix -ся or -сь. These parts of word are very
homonymous. So, prefix при- could mean approaching, connection, proximity,
partiality of an action, finalization of an action, etc. According to these meanings, the
corresponding vectors could be perpendicular or antiparallel. This results in very low
accuracy for all used models. We used just initial forms for Fig. 5 and Fig. 7 to eliminate
the influence of grammatical features.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Discussion and Conclusion</title>
      <p>In this article, we found several reasons why the vector transformation does not work
on some categories of word analogies.
1. A used language model should be taught on texts that have enough occurrences of
words for which the task of word analogies is solved. As we can see in Fig. 1 and 2,
categories including countries, their capitals, and other related words are better
analyzed using corpora like Wikipedia, since such corpora have enough information for
inference of logical relations among these words. Moreover, these words are
allocated in the same context in the Wikipedia text; thus, their vectors become more
similar. Other model, which was trained on fiction or news texts, does not have the
same context. That is why categories A1-A5 demonstrate the best results on the
Wikipedia corpus. Note that this conclusion needs to be proved in an independent
investigation.
2. In a common case, the task of word analogies could not be solved using a random
word pair taken from two investigated categories. At least, some words are
homonymous, and their vectors are out of the common systems for words of the selected
category. If these words are chosen as the basis in such methods as Word2Vec, then
such basis will be shifted and word analogies will be found incorrectly. For example,
Moscow and Berlin are respectful representatives of their countries in case of
international or cultural affairs; these cities are used as synonyms of Russian and German
government and culture. However, Bogota and Kampala are rather a capital city than
a government. Moreover, Russian prefixes and suffixes could have different
meanings. Thus, there will be several preferential directions for different prefix or suffix
values. If someone misuses just one word in a pair, he or she will get a wrong
transition vector.</p>
      <p>Averaging of the starting and finishing points helps to eliminate this problem. This
helps more in the case of factual information (about 30% increment for accuracy),
but less for grammatical information (less than 10% for Masculine → Feminine
transition and several percent for other categories in the case of Word2Vec, 1-30%
depending the category for FastText). Moreover, averaging over extremely
homonymous prefixes and suffixes makes the results worse (categories A11, A15 for
FastText over the Araneum corpus).</p>
      <p>Averaging over a word list has some drawbacks. One should have a list of words in
a given category to average their vectors; this is not always possible. On the other
hand, sometime such a list is available, and one could use it to tune vectors in order
Evaluation of Vector Transformations for Russian Word2Vec and FastText Embeddings 11
to predict out-of-list words. But previously he or she should be sure that this list does
not contain several preferred semantic directions.</p>
      <p>This list could be used separately, but not for calculating pairs only. Any group will
have its own center coordinates. If we have a pair a:b, we can use the coordinates of
the nearest centers instead of the vectors of these words.
3. The main idea of an affine transition is that there is one vector that could be added
to the word a to find its analogy b. That means that all vectors for a:b should be
equal, i.e. have approximately equal length and angles of orientation. At least, the
value of the bias between this transition vector and the vector of the correct answer
should be less than half the distance to the nearest neighbor. As we have found out,
it is not always true. In case of homonymous prefixes and suffixes, some vector
groups could be oppositely directed. This means that the word analogies task could
not be solved using only one transition vector. Our experiments have demonstrated
that in some cases the length of the vectors could differ more than twice. Such biases
lead to the situation, when the software module has to generate several outputs, and
the user should find some extra methods to find the correct answer.</p>
      <p>In our future research, we are going to investigate these drawbacks in more detail. One
of the solutions here is to construct an interpretable vector space where each axis
corresponds to a semantic feature. LSA-based methods successfully mark domains
building a hierarchical representation of language semantics. However, these methods are
not able to construct a vector space. In our opinion, incorporation of these two methods
(LSA and vector space) could help achieve reasonable results. On the one hand, the
vector space of Word2Vec-inherited methods allows the vector arithmetic in a
continuous semantic space; on the other hand, the application of a hierarchy of terms to such
a continuum makes it more interpretable.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ducharme</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vincent</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jauvin</surname>
            ,
            <given-names>P. A.</given-names>
          </string-name>
          :
          <article-title>Neural Probabilistic Language Model</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>3</volume>
          ,
          <fpage>1137</fpage>
          -
          <lpage>1155</lpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Collobert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A unified architecture for natural language processing</article-title>
          .
          <source>In: Proceedings of the 25th International Conference on Machine Learning</source>
          , vol.
          <volume>20</volume>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>167</lpage>
          . (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
          </string-name>
          , W.-T.,
          <string-name>
            <surname>Zweig</surname>
          </string-name>
          , G.:
          <article-title>Linguistic Regularities in Continuous Space Word Representations</article-title>
          .
          <source>In: Proc. of HLT-NAACL</source>
          , pp.
          <fpage>746</fpage>
          -
          <lpage>751</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>In: Proc. of International Conference on Learning Representations (ICLR)</source>
          , (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .
          <source>In: Proc. of 27th Annual Conference on Neural Information Processing Systems</source>
          , pp.
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching Word Vectors with Subword Information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>5</volume>
          ,
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>BERT</surname>
          </string-name>
          :
          <article-title>Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . ArXiv:
          <year>1810</year>
          .04805, https://arxiv.org/pdf/
          <year>1810</year>
          .04805.pdf,
          <source>last accessed</source>
          <year>2020</year>
          /07/12.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyyer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gardner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>ArXiv</source>
          :
          <year>1802</year>
          .05365, https://arxiv.org/pdf/
          <year>1802</year>
          .05365.pdf,
          <source>last accessed</source>
          <year>2020</year>
          /07/12.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ethayarajh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>How Contextual are Contextualized Word Representations?Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings</article-title>
          . ArXiv:
          <year>1909</year>
          .00512v1, https://arxiv.org/pdf/
          <year>1909</year>
          .00512.pdf,
          <source>last accessed</source>
          <year>2020</year>
          /07/12.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Linguistic Regularities in Sparse and Explicit Word Representations</article-title>
          .
          <source>In: Proc. of 18th Conf. on Computational Natural Language Learning</source>
          , pp.
          <fpage>171</fpage>
          -
          <lpage>180</lpage>
          . (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Köper</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scheible</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulte</surname>
          </string-name>
          im Walde, S.:
          <article-title>Multilingual reliability and “semantic” structure of continuous word spaces</article-title>
          .
          <source>In: Proc. of 11th International Conference on Computational Semantics</source>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>45</lpage>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Vylomova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rimmel</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohn</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldwin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Take and took, gaggle and goose, book and read: evaluating the utility of vector differences for lexical relation learning</article-title>
          .
          <source>In: Proc. of 54th Annual Meeting of the Association for Computational Linguistics</source>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>1671</fpage>
          -
          <lpage>1682</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Drozd</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>B.:</given-names>
          </string-name>
          <article-title>The (Too Many) Problems of Analogical Reasoning with Word Vectors</article-title>
          .
          <source>In: Proc. of 6th Joint Conference on Lexical and Computational Semantics</source>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>148</lpage>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Newman-Griffis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fosler-Lussier</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Insights into analogy completion from the biomedical domain</article-title>
          .
          <source>BioNLP</source>
          ,
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Drozd</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gladkova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsuoka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Analogy-based Detection of Morphological and Semantic Relations With Word Embeddings: What Works and What Doesn't</article-title>
          .
          <source>In: Proc. of NAACL Student Research Workshop</source>
          , pp.
          <fpage>8</fpage>
          -
          <lpage>15</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Linzen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Issues in evaluating semantic spaces using word analogies</article-title>
          .
          <source>In: Proc. of 1st Workshop on Evaluating Vector-Space Representations for NLP</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>18</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Drozd</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gladkova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsuoka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Word Embeddings, Analogies, and
          <source>Machine Learning: Beyond King-Man+Woman=Queen. In: Proc. of COLING</source>
          <year>2016</year>
          , pp.
          <fpage>3519</fpage>
          -
          <lpage>3530</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Kutuzov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuzmenko</surname>
          </string-name>
          , E.:
          <article-title>WebVectors: A Toolkit for Building Web Interfaces for Vector Semantic Models</article-title>
          .
          <source>In: Analysis of Images, Social Networks and Texts (AIST)</source>
          <year>2016</year>
          , vol.
          <volume>661</volume>
          , pp.
          <fpage>155</fpage>
          -
          <lpage>161</lpage>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Learning Word Vectors for 157 Languages</article-title>
          . In
          <source>: Proc of LREC'2018</source>
          , pp.
          <fpage>3483</fpage>
          -
          <lpage>3487</lpage>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>