<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Techniques Comparison for Natural Language Processing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ender Turing OU</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tallinn</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Estonia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ei}@enderturing.com</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Borys Grinchenko Kyiv University</institution>
          ,
          <addr-line>Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute</institution>
          ,”
          <addr-line>Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>These improvements open many possibilities in solving Natural Language Processing downstream tasks. Such tasks include machine translation, speech recognition, information retrieval, sentiment analysis, summarization, question answering, multilingual dialogue systems development, and many more. Language models are one of the most important components in solving each of the mentioned tasks. This paper is devoted to research and analysis of the most adopted techniques and designs for building and training language models that show a state of the art results. Techniques and components applied in the creation of language models and its parts are observed in this paper, paying attention to neural networks, embedding mechanisms, bidirectionality, encoder and decoder architecture, attention, and self-attention, as well as parallelization through using transformer. As a result, the most promising techniques imply pre-training and fine-tuning of a language model, attention-based neural network as a part of model design, and a complex ensemble of multidimensional embedding to build deep context understanding. The latest offered architectures based on these approaches require a lot of computational power for training language models, and it is a direction of further improvement. Algorithm for choosing right model for relevant business task provided considering current challenges and available architectures.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>NLP</kwd>
        <kwd>Language Model</kwd>
        <kwd>Embedding</kwd>
        <kwd>Recurrent Neural Network</kwd>
        <kwd>RNN</kwd>
        <kwd>Gated Recurrent Unit</kwd>
        <kwd>GRU</kwd>
        <kwd>Long Short-Term Memory</kwd>
        <kwd>LSTM</kwd>
        <kwd>Encoder</kwd>
        <kwd>Decoder</kwd>
        <kwd>Attention</kwd>
        <kwd>Transformer</kwd>
        <kwd>Transfer Learning</kwd>
        <kwd>Deep Learning</kwd>
        <kwd>Neural Network</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Natural Language Processing (NLP) is computer comprehension, analysis,
manipulation, and generation of natural language. NLP covers many different applications like
machine translation, speech recognition, optical character recognition, part of speech
tagging, information retrieval, summarization, question answering, dialog systems
building, and many more. Since the last decade, there have been a great number of
breakthroughs towards making machines understand the language (text or voice) better
that raised enormous interest in this field of scientific research. The highest goal for all
scientists working in the area of NLP is to build such techniques that will allow
computers to comprehend natural language as text or voice at a human level which is not
reached yet. Once achieved, computational systems will be able to understand and
generate accurate human-like language, be it a text or an audible language. This paper
analyses widespread techniques and components in building language models to give a
scientist thorough information for further research and improvement. There have been
a lot of discoveries of language models from computational linguistics scientists, and
those new techniques showed great results on specific tasks. But when it comes to the
broad spectrum of NLP tasks solved by the same language model not many show the
same high results. The recent state of the art architectures (BERT, RoBERTa,
Transformer-XL, XLNet, etc.) leverage the following approaches: contextual embedding,
bidirectionality, encoder-decoder architecture, attention mechanism and transformer,
pretraining modeling and fine-tuning.</p>
      <p>The structure of this paper includes language models architecture observation in
Sect. 2. Word and contextual embedding techniques are described in Sect. 3 and
encoder-decoder—in Sect. 4. Neural Networks applied as a part of encoder and decoder
are observed in Sect. 5. Model choosing algorithm presented in Sect.6. Conclusion in
Sect. 7 provides further promising areas of research in NLP.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Common Approach of Natural Language Processing</title>
    </sec>
    <sec id="sec-3">
      <title>Downstream Tasks</title>
      <p>
        NLP downstream tasks (machine translation, sentiment analysis, question answering,
part of speech tagging, and many more) usually are solved with some different
approaches chosen for a specific task. Generally, it comes to supervised learning on
taskspecific datasets, which is quite consuming in terms of research hours and
computational resources. In addition to these inconveniences, systems that are built with this
approach are very sensitive to task specifications and changes in the data distribution.
Current trends move towards unsupervised universal models and transfer learning as
pre-trained models with further fine-tuning. The techniques and components that are
offered for observation in this article correspond to current trends and groundbreaking
achievements in NLP [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It outlines the architectures of the state of the art pre-training
models [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Everything starts with ready for training or test data. Input data runs through
some techniques to result in word embedding [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] or contextual embedding that could
be described as multidimensional word knowledge embedding. Despite the use of the
term word, readers should not be confused. Word embedding is a form of a vector that
can be based on characters, subwords, words, sentences, or even longer sequences each
of which is called a token. Contextual embedding [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is used as input for an encoder
which forms a context vector and forwards it to a decoder. A decoder in its turn forms
a set of probabilities necessary to figure out an output.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Word and Contextual Embedding Approaches</title>
      <p>
        For a neural network to be able to complete its task there is a necessity to provide a
numerical token representation of input sequence. Word embedding techniques create
vectors out of tokens. Vectors comparison results in tokens semantic similarity.
Embedding techniques such as GloVe and Word to vector explain the concept of modeling
input sequence through representation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The main idea and task are to represent and map words (documents, phrases, context,
a piece of a word, or a character) as a vector of numbers to use probability distributions
or likelihoods of tokens in language corpora to separate semantic similarity categories.
Hence different words with similar meanings will have similar vectors and different by
meaning groups of words should be separable in vector space. The underlying idea that
“a word is characterized by the company it keeps” was popularized by Firth [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Currently, the area of representation of input sequence advances far ahead of initial papers
and new approaches appear. Contextual embedding [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] creates a representation for each
token taking into account its context, meaning getting information of a token usage in
different contexts and encode knowledge that is transferable to some other languages.
4
      </p>
    </sec>
    <sec id="sec-5">
      <title>High-Level Encoder-Decoder Architecture</title>
      <p>The neural network encoder-decoder model significantly improved the performance of
language models. Quite simple Recurrent Neural Network (RNN) architecture to
process input and output sequences of variable lengths was offered. The input sequence of
words [“How,” “are,” “you,” “?”] go through the embedding layer to get numerical
representation, after which numerical representation goes sequentially to RNN. RNN
process input embedding sequentially (from left to right) passing to the next timestamp
RNN hidden state calculated in the current timestamp RNN. After all inputs proceed to
the final time stamp, the final timestamp RNN produces output representing all input
sequences in one hidden state. This part called encoder as the main task is not to
generate predictions but to encode input sequences. After encoder finishes to encode,
hidden state passes to the decoder which task is to decode and generate predictions based
on input hidden state. Decoder process sequentially taking as input to each current
timestamp output activation of previous timestamp RNN and output prediction of
previous timestamp RNN. For the first timestamp, it takes a beginning of sentence token
(BOS) as a prediction of the previous layer. The decoder generates predictions until it
generates end of sentence token (EOS, depending on implementation can be until some
length or different parameter).</p>
      <p>The strongest part of this approach is the ability to train an end-to-end model right
on the source and target data as well as the possibility to handle input and output
sequences of different lengths. Therefore, that it resolves the problem of different lengths
of an input and an output sequence in Neural Machine Translation.</p>
      <p>The encoder-decoder architecture consists of two RNNs or more often Long
ShortTerm Memory (or Gated Recurrent Unit) to avoid the problem of vanishing gradient
covered later in this article. The encoder encodes all input sequences and stores all
information in context or encoder vector (in simplest architecture last hidden state used)
that is input to decoder, which decodes by result predictions.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Neural Networks for Natural Language Processing</title>
    </sec>
    <sec id="sec-7">
      <title>Tasks</title>
      <p>Progress in the application of Neural Networks to NLP tasks brings huge improvements
in both science and business areas.
5.1</p>
      <sec id="sec-7-1">
        <title>Recurrent Neural Networks</title>
        <p>
          RNNs [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] is the main starting point in the deep learning NLP area. Deep neural
networks uncover a second life for RNNs. A strong RNN advantage for the NLP area is
that RNN can store the conditions of all cells that processed language data before
sequentially.
        </p>
        <p>
          The main idea behind RNNs [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is very simple, the network takes an input vector X
and produces an output vector Y. Each RNN cell takes as input current xt and previous
hidden state (activation) ht-1. It learns weights (parameters) Wh, Wx, and bias ba through
the weights learning process. At each iteration of Forward Propagation, nonlinear
activation function g such as tanh (or rarely ReLU) applied to calculate output hidden state
(activation) ht:
ht  g Whht 1  Wx xt  ba .
(1)
If output predictions needed by task then activation function g (or softmax function)
with learned weights Wy and bias by might be applied to current output ht to make
output prediction yt:
        </p>
        <p>yt  g Wyht  by .</p>
        <p>Especially important for NLP areas is that the output vector’s contents are calculated
not only by the one current input but based on the entire history of inputs that network
processed in the past. RNN cells take the output of the previous cell as an input to the
current cell, and the previous cell contains information of its previous cell and so on.
This type of connection is called a recurrent connection.</p>
        <p>
          Despite the wide adoption, RNN possesses significant drawbacks [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Unidirectional
learning leads to the problem that the model cannot rely on information from the later
part of a sequence while working on the beginning of a sequence. For RNNs it is hard
to capture mid and long-term connections/dependencies inside a sequence—this issue
is known as long-range dependencies problem or Gradient Vanishing. This was a
triggering point to search for a solution that resulted in further useful findings like GRU
and LSTM.
5.2
        </p>
      </sec>
      <sec id="sec-7-2">
        <title>Gated Recurrent Unit and Long Short-Term Memory</title>
        <p>In response to medium and long-range dependency problems researchers propose two
architectures, with the core idea of Cell State (kind of residual connection).</p>
        <p>
          The main idea in Gated Recurrent Unit (GRU) [
          <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
          ] is to capture long-term
dependencies by adding Memory Cell (Ct) which in GRU is equal to hidden state (activation)
ht  (1  z)  ht 1  zt  h%t. And each time stamp cell considers rewriting this cell
with Candidate Value h%t tanh(W [rt  ht 1, xt ]), using two Gates described by
equations (Update Gate zt  (Wz [ht 1, xt ]), and Reset Gate rt  (Wr [ht1, xt ]) ) as
shown in Fig. 1. Update gate (z) takes a value between 0 and 1 (most times close to 0
or 1), computed by application of sigmoid activation function to current timestamp
input xt and previous time stamp hidden state (activation) ht-1 with learned weights
(parameters) Wz through the weights learning process. This Update Gate is the main
decision-maker of updating the hidden state as shown in equations. Update Gate decides
how much information from the previous timestamp should be saved for the future.
Reset gate, on the other hand, decides how much information from the previous
timestamp should be removed.
        </p>
        <p>These Update and Reset Gates are the key concepts behind GRU and dealing with
dependencies problems of basic RNNs.
ht-1</p>
        <p>GRU cell
xt
rt</p>
        <p>
Reset
gate
zt
1
Update
gate
tanh
h~t
ht</p>
        <p>
          Another type of architecture that can capture mid-term dependencies, even more
powerfully than GRU, is Long Short-Term Memory (LSTM) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In a difference to GRU
that has two gates, LSTM possesses three gates. An important concept in LSTM is that
Memory Cell (Ct) is not anymore equal to output hidden state (activation) ht. Output
hidden state of the current timestamp in LSTM carry on to the next cell not alone but
with updated Memory Cell value.
        </p>
        <p>LSTM, also, uses two separate gates (Update Gate and a Forget Gate) to update
Memory Cell value, instead of using single Update Gate in GRU (that either keep or
forget previous memory cell value). And instead of using Reset Gate in Candidate
Value, it uses element-wise multiplied with Memory Cell value as shown in Fig. 2.
Input Gate it  (xtU i  ht 1W i ), decides which information crucial to keep and
Forget Gate ft  (xtU f  ht1W f ), decides which and how much information not to keep,
in other words, to forget (which might be intersecting or not with input gate). Input and
Forget gates do this using previous time step hidden state ht-1 and current input xt. Both
gates using sigmoid activation function, which gives possibility in most cases to have
values of gates either close to 0 or 1.</p>
        <p>Usage of separate Update Gate and Forget Gate to calculate Memory Cell value
  =  (  ∗   −1 +   ∗  ̃ ) gives Memory Cell the possibility not only to store new
information in the current time step Memory Cell by using Candidate Memory Cell
( ̃ ), but also the option to keep some amount of information from previous time step
Memory Cell (Ct-1). Output Gate ot  (xtU 0  ht 1W 0 ), at the end uses to calculate
the current time step output hidden state ht  tanh(Ct )  ot , based on updated Memory
Cell value calculated before.</p>
        <p>The state of a cell is straight forward. It flows down the whole unit with minor linear
changes. This is why two proposed architectures were very good at memorizing
longterm dependencies.</p>
        <p>Ct-1
ht-1</p>
        <p>LSTM cell
ft

it

~
Ct
tanh

tanh</p>
        <p>ht
Xt</p>
        <p>Foget
gate</p>
        <p>Input
gate</p>
        <p>Output
gate</p>
        <p>Ct
ht</p>
        <p>These networks are quite consuming for computational resources. Moreover, this type
of architecture cannot be parallelized: hence, it is very expensive to train on a big corpus
of data. And despite the fact it works much better with longer sequences there is a
noticeable loss in sequences with more than 20 words. Retrospectively application of
RNN, GRU, and LSTM was the major stage in modern NLP that significantly affected
future development of the area.
Additionally, there was a significant amount of work on the bidirectionality of RNN to
provide models with the possibility to capture and use information from both earlier
and later in the sequence.</p>
        <p>
          If to express in simple words Bidirectional RNN (BRNN) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is a modification to
RNN, GRU, LSTM consists of two RNNs capturing information simultaneously in
opposite directions and only then making predictions. BRNN has forward recurrent layer
(component) S that takes as input current X and feeds the output to help predict current
output Y forward in time. On the other hand, backward recurrent layer (component) Si
which takes as input current X and feeds the output to help predict current output Y
backward in.
        </p>
        <p>To construct even more powerful models researchers propose to stack units of
RNN/LSTM/GRU. This type of architecture is called Deep RNN. The bottleneck of
BRNN is that it needs the entire sequence of data before making any predictions. Deep
RNN is also much more expensive in computation. All of the networks presented had
problems in neural machine translation as the input and output sequences regularly were
of different lengths because of different language semantics.
5.4</p>
      </sec>
      <sec id="sec-7-3">
        <title>Attention Concept</title>
        <p>
          The focus of researchers was the problem of long sentences (sentence contained more
than 20 words) which cannot be stored effectively in one output vector of
RNN/GRU/LSTM. In [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] demonstrated significant improvement of BLEU score
results using attention mechanisms. As a solution to the problem, Attention Mechanism
was proposed.
        </p>
        <p>The main intuition behind Attention is that humans do not read and memorize whole
long sentences at once, but part by part. And for a decoder, it would be valuable to
know while decoding (for example translation), to which part of the input sequence it
should pay more attention. An idea of attention: at each step, the decoder focuses on
some particular part of the source. Decoder focuses only on particular words at each
step (increased saturation represents more attention), not on the full input sequence.</p>
        <p>Attention mechanism uncovers such possibility to a decoder by Attention Weights
and Context Vector.</p>
        <p>In addition to BRNN (which can also be BGRU, BLSTM) Attention concept utilizes
the idea of alignment scores and attention weights (the amount of attention decoder
should pay while calculating current time step prediction).</p>
        <p>The all-time step hidden states of encoder pass with the last layer hidden state of
encoder to the decoder. Central processing occurs in the decoder. Each time step of
decoder, a set of features (about words and surrounding words) computes and called
Alignment Scores eij  a(si1,h j ) (differences between encoder and decoder hidden
states), which are used to calculate Attention Weights ij 
exp(eij )</p>
        <p>T
kx1ijh j
, by softmax
function.</p>
        <p>Context Vector ci  Tjx1ijh j , is calculated for each time step of Decoding by
combining Attention Weights with the previous decoder outputs to be passed to decoder
RNN
ts 
s1
exp(score(ht , hs ))
S
 exp(score(ht , hs ))
.</p>
        <p>
          Although amazingly, such a simple and generic architecture as bidirectional LSTM
with attention (just a few equations and few tens lines of code) can predict (translate,
classify) with such a great result this architecture admits mistakes and has bottlenecks
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. This type of architecture cannot be parallelized—attention mechanism provided
for sequential RNNs helped solve long-term dependencies issues by using more
appropriate context at each step, but the problem of parallelization of computation raised
even more.
        </p>
        <p>Additionally if to analyze the NLP area not only through the prism of translation
where most time machines just translate sentence by sentence and focus on Natural
Language Understanding area RNNs do not show good results in overall context
understanding and modeling, especially during text generation tasks. This is exactly where
the architecture of the transformer can do better.
5.5</p>
      </sec>
      <sec id="sec-7-4">
        <title>Self-Attention and Transformer</title>
        <p>
          In the paper [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] researchers from Google introduced transformer, a novel neural
network architecture for Language Understanding based on a self-attention mechanism.
High-level architecture of encoder and decoder of the transformer are presented in
Fig. 3 and described in detail below.
        </p>
        <p>The main novelty was that to build a language model there is no need for any
recurrent (RNN) or convolutional (CNN) layers at all. Solely self-attention and feed-forward
layers are enough.</p>
        <p>Encoder</p>
        <p>Feed Forward
Self-Attention</p>
        <p>Decoder</p>
        <p>Feed Forward
Encoder-Decoder Attention</p>
        <p>Self-Attention
Fig. 3. Encoder-decoder architecture applied in the transformer.</p>
        <p>The transformer includes a lot from what was described before: bidirectionality;
encoder-decoder architecture, to support the different length of input-output;
self-attention in addition to attention, researchers develop the idea of attention and presented
self-attention in the transformer; parallelization (at least on Feed-Forward step which
is most expensive) by completely replacing sequential computation (RNNs or
Convolutional based) to Attention-based network.</p>
        <p>The main components and concepts of this architecture will be presented and
described below.</p>
        <p>Encoder. The encoder consists of multiple stacked Self-Attention and Feed Forward
layers with Residual Connections and Positional Encoder. As usual, the embedding
layer is applied in the bottom to convert input sequence to numerical representation.
Feed Forward Network does not have dependencies and thus can be parallelized. This
is an important concept behind the transformer possibility to learn on a truly big amount
of data that LSTM and GRU cannot afford.</p>
        <p>Self-Attention Layers help to understand the model, which parts of the input
sequence (words) to focus on while encoding sequence. The most important novelty is
using three vectors: Query vector, Key vector, and Value vector to create “query,”
“key,” and “value” projection of each word in the input sentence.</p>
        <p>Decoder. The decoder also consists of multiple (equal to the encoder) stacked
SelfAttention, Feed Forward layers with Residual Connections, and additionally
encoderdecoder attention layer in the middle. In comparison to the encoder, Decoder’s
SelfAttention layer differs. The main idea here is Masking Future Positions. In the encoder,
each position can attend to all positions, but in the decoder to prevent leftward
information flow to preserve the auto-regressive property, each position can attend only to
early positions in the output sequence.</p>
        <p>Another important layer of the decoder is Encoder-Decoder Attention layer which
gets outputs of the last Attention layer of the encoder as an input and uses Key and
Value attention vectors to focus on appropriate places in the input sequence.
6</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Comparison Concepts and Architectures Usage</title>
      <p>Embedding. Almost everyone nowadays uses some kind of embedding technique to
tokenize input sequences. If your task is not domain-specific, you will probably end up
using one of the pre-trained embeddings with dimensionality (300–512). And if you
use domain-specific tasks, the choice would be simply based on available
computational resources. Starting from simple Word2Vec and Glove and moving to advanced
Contextual embedding techniques and increased dimensionality can solve your tasks
with high accuracy.</p>
      <p>Architecture. As of today, the transformer is the most powerful architecture, which
can be trained on enormous amounts of training data with billions of parameters. It is
clear that it is impractical to train such a big network from scratch every time, and for
every particular task (even today, it will cost hundreds of thousands of USD and
enormous computational GPU power). Hence, such a big model comes as pre-trained
models, which can then be fine-tuned for various scenarios and tasks. It can be achieved by
an additional layer of neurons on the end that was not trained in the pre-training and
train them as a part of the new model for specific tasks.</p>
      <p>A key advantage of models built using the transformer architecture is that it does not
need to be trained with labeled data, so it can learn using any cleaned raw text. This
provides the possibility to work with very big datasets and leads to even better accuracy.</p>
      <p>The current leaderboard in different NLP competition more and more narrowing to
big tech corporations and not universities. That is because of the availability of
computational power. On the other hand, leaderboard results and practical implementation is
different. Even today, many production-based services use RNN (LSTM and GRU
modifications) with attention e.g., for intent classification in widespread chatbot
frameworks.</p>
      <p>In Fig 4. proposed algorithm for choosing modern approach to solve business NLP
tasks considering domain specification of task and availability of resources. Also, it is
a good idea to compare BLSTM with Attention and Transformer based architectures
results in terms of accuracy, consumption of resources, time to train (hence to improve),
interpretability.
Significant progress was made in the last decade in the NLP area. Application of deep
learning changed rules and uncovered new possibilities with RNNs, BLSTMs, and
Attention. Architectures based on Transformers advances even further and show
state-ofthe-art results on most of the NLP tasks.</p>
      <p>The introduction of deep pre-trained language models in the last couple of year’s
significant shift to transfer learning in NLP. At the same time, many of the latest
approaches are too demanding in computational resources and algorithms for choosing
the right model for the business task presented in current work.</p>
      <p>The authors see demand in the sentence boundary detection task. They will research
the current area in nearest future using the algorithm provided in the current work to
choose models and compare obtained results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cho</surname>
          </string-name>
          , K.,
          <string-name>
            <surname>van Merrienboer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>On the properties of neural machine translation: Encoder-decoder approaches</article-title>
          .
          <source>In SSST-8</source>
          , Eighth Workshop on Syntax, Semantics and Structure in Statistical Translation:
          <fpage>103</fpage>
          -
          <lpage>111</lpage>
          (
          <year>2014</year>
          ). https://doi.org/10.3115/ v1/w14-
          <fpage>4012</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Firth</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          :
          <source>A Synopsis of Linguistic Theory</source>
          ,
          <fpage>1930</fpage>
          -
          <lpage>1955</lpage>
          (
          <year>1957</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Glove: global vectors for word representation</article-title>
          .
          <source>In 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          :
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          ). https://doi.org/1010.3115/v1/d14-
          <fpage>1162</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>In First International Conference on Learning Representations:</source>
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2013</year>
          ). http://arxiv.org/abs/1301.3781
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          <volume>1</volume>
          :
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          (
          <year>2018</year>
          ). https://doi.org/10.18653/v1/n18-
          <fpage>1202</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Rumelhart</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Learning representations by back-propagating errors</article-title>
          .
          <source>Nat</source>
          .
          <volume>323</volume>
          (
          <issue>6088</issue>
          ):
          <fpage>533</fpage>
          -
          <lpage>536</lpage>
          (
          <year>1986</year>
          ). https://doi.org/10.1038/323533a0
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Sequence modeling: recurrent and recursive nets</article-title>
          , Deep Learning:
          <fpage>367</fpage>
          -
          <lpage>415</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Gulcehre, С.,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Empirical evaluation of gated recurrent neural networks on sequence modeling</article-title>
          .
          <source>In NIPS 2014 Workshop on Deep Learning and Representation Learning</source>
          :
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          (
          <year>2014</year>
          ). http://arxiv.org/abs/1412.3555
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural Computation</source>
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          (
          <year>1997</year>
          ). https://doi.org/10.1162/neco.
          <year>1997</year>
          .
          <volume>9</volume>
          .8.
          <fpage>1735</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliwal</surname>
            ,
            <given-names>K.:</given-names>
          </string-name>
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>In IEEE Transactions on Signal Processing</source>
          <volume>45</volume>
          (
          <issue>11</issue>
          ):
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          (
          <year>1997</year>
          ). https://doi.org/10.1109/78.650093
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.:</given-names>
          </string-name>
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
          <source>In International Conference on Learning Representations (ICLR)</source>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          (
          <year>2015</year>
          ). http://arxiv.org/abs/1409.0473
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , et al.:
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          :
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          (
          <year>2017</year>
          ). http://arxiv.org/abs/1706.03762
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>