<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Yuen, S.Y., Chow C. K.: A Genetic Algorithm that Adaptively Mutates and Never Revis-
its. IEEE Transactions on Evolutionary Computation.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/TEVC.2008.2003008</article-id>
      <title-group>
        <article-title>Forecast Method for Natural Language Constructions Based on a Modified Gated Recursive Block</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Cherkasy State Technological University</institution>
          ,
          <addr-line>Cherkasy, Shevchenko blvd., 460, 18006</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2007</year>
      </pub-date>
      <volume>472</volume>
      <issue>2009</issue>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The paper proposes a method for predicting natural language constructions based on a modified gated recursive block. For this, an artificial neural network model was created, a criteria for evaluating the efficiency of the proposed model was selected, two methods for parametric identification of the artificial neural network model were developed is based on the backpropagation through time algorithm and based on simulated annealing particle swarm optimization algorithms. The proposed model and methods for its parametric identification make it possible to more accurately control the share of information coming from the input layer and the hidden layer of the model, increase the parametric identification speed and the prediction probability. The proposed method for predicting natural language constructions can be used in various intelligent natural language processing systems.</p>
      </abstract>
      <kwd-group>
        <kwd>modified gated recursive block</kwd>
        <kwd>prediction of natural language constructions</kwd>
        <kwd>particle swarm optimization</kwd>
        <kwd>simulated annealing</kwd>
        <kwd>parametric identification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Currently, one of the most important problems in the field of natural language
processing is the insufficiently high accuracy of the analysis of alphabetic and/or
phoneme sequences [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. This leads to the fact that natural language processing may be
ineffective. Therefore, the development of methods for predicting natural language
constructions is an important task.
      </p>
      <p>
        As a prediction method, a neural network forecast [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] was chosen, which, when
forecasting natural language constructions, has the following advantages:
─ correlations between factors are studied on existing models;
─ no assumptions regarding the distribution of factors are required ;
─ prior information about factors may be absent;
─ source data can be highly correlated, incomplete or noisy;
─ analysis of systems with a high degree of nonlinearity is possible;
─ fast model development;
─ high adaptability;
─ analysis of systems with a large number of factors is possible;
─ full enumeration of all possible models is not required;
─ analysis of systems with heterogeneous factors is possible.
      </p>
      <p>
        The following recurrent networks are
networks [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4-6</xref>
        ]:
most often
used
as forecast neural
─ Jordon neural network (JNN) [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ];
─ Elman neural network (ENN) or simple recurrent network (SRN) [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ];
─ bidirectional recurrent neural network (BRNN) [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ];
─ long short-term memory(LSTM) [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ];
─ gated recurrent block (GRU) [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ];
─ echo state network (ESN) [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ];
─ liquid state machine (LSM) [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Criterion</title>
      <p>aLolowcaplreoxbtarbeimliutymof getting into - - - - - + +
High learning speed + + + - + + +
Possibility of batch training - - - - - - +
Dynamic control of the share of
information from the input and - - - + + -
hidden layers</p>
      <p>According to table 1, none of the networks meets all the criteria. In this regard, the
creation of training methods that will eliminate these drawbacks is relevant.</p>
      <p>
        To increase the probability of falling into a global extremum and replacing batch
training with multi-agent training, metaheuristic search is often used instead of local
search [
        <xref ref-type="bibr" rid="ref21 ref22 ref23 ref24 ref25">21-25</xref>
        ]. Metaheuristics expands the capabilities of heuristics by combining
heuristic methods based on a high-level strategy [26-30].
      </p>
      <p>
        Modern metaheuristics may have one or more of the following disadvantages:
─ there is only a generalized method structure or the method structure is focused on
solving only a specific problem [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ];
─ iteration numbers are not present when searching for a solution [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ];
─ the method may not converge [31];
─ material potential solutions are unacceptable [32];
─ there is no formalized parameter values search strategy [33];
─ the method is not intended for conditional optimization [34];
─ the method does not possess high accuracy [35].
      </p>
      <p>Thereby, arises the problem of constructing an effective metaheuristic optimization
method.</p>
      <p>Thus, the task to create an effective forecast model for alphabetic and/or phoneme
sequences, which is trained based on effective metaheuristics, is relevant today.</p>
      <p>The purpose of the work is to develop a forecast method for natural language
constructions based on a modified gated recurrent block. To achieve the goal, the
following tasks were set and solved:
1. Create a model of a modified gated recursive block.
2. Select a criteria for evaluating the efficiency of the proposed model.
3. Develop a method for the parametric identification of a model based on local
search.
4. Develop a method for parametric identification of a model based on a multi-agent
metaheuristic search.
5. Conduct a numerical study.
2</p>
      <sec id="sec-2-1">
        <title>Creating a model of a modified gated recursive block</title>
        <p>The paper proposes a modification of GRU by introducing 1 rj n factor for the
weighted sum of the neurons outputs in the input layer, which allows to more
accurately control the share of information coming from the input layer and the hidden
layer.</p>
        <p>The proposed modified gated recurrent unit (MGRU) is a recurrent two-layered
artificial neural network (ANN) with an input layer in , a hidden layer h , an output
layer out . Just as for a regular GRU, each neuron in hidden layer is associated with
reset and update gates is (FIR filters). The structural representation of the MGRU
model is shown in Fig. 1.</p>
        <p>in
…
out
…
h
r
…
…
…
z</p>
        <p>Fig. 1. Structural representation of a modified gated recurrent block (MGRU)
Gates determine how much information to pass. Thereby, the following special cases
are possible. If the share of information passed by the reset gate is close to 0.5, and
the share of information passed by the update gate is close to 0, then we get SRN. If
the share of information passed by the reset gate is close to 0 and the share of
information passed by the update gate is close to 0, then the ANN information is updated
only due to the input (short-term) information. If the share of information passed by
the update gate is close to 1, then the ANN information is not updated. If the share of
information passed by the reset gate is close to 0 and the share of information passed
by the update gate is close to 1, then the ANN information is updated only due to
internal (long-term) information.</p>
        <p>The modified gated recurrent block (MGRU) model is presented in the following
form:
─ calculating the share of information passed by the reset gate</p>
        <p> N0 N1 
rj n  f  i1 iijnr yiin n  i1 ihj r yih n 1  , j 1, N 1 ,
─ calculating the share of information passed by the update gate</p>
        <p> N0 N1 
z j n  f   uinz yiin n   uhz yih n 1  , j 1, N 1 ,</p>
        <p> i1 ij i1 ij 
─ calculating the output signal of the candidate layer</p>
        <p> N(0) N(1) 
y hj (n)  g  (1 rj (n)) i1 wiijnh yin (n  i)  rj (n) i1 wihjh yih (n 1)  , j 1, N 1 ,
─ calculating the output signal of the hidden layer
─ calculating the output signal of the output layer
y hj (n)  z j (n) y hj (n 1)  (1 z j (n)) y hj (n), j 1, N 1 ,</p>
        <p> N1 
y ojut n  f   whout yih n  , j 1, N 2 ,</p>
        <p>ij
 i1 
f  s </p>
        <p>1
1 es ,
g s  tanh s ,
where N 0 is the number of neurons in the input layer;
time n , rj n  0, 1 ;
at a time n , z j n  0, 1 ;</p>
        <p>N 2 is the number of neurons in the output layer;
N 1 is the number of neurons in the hidden layer;
 inr , uiijnz are connection weight from the i th input neuron to the reset gates and
ij
update of the j th hidden neuron;
 hr , uihjz are connection weight from the i th hidden neuron to the reset gates and
ij
update of the j th hidden neuron;
winh is connection weight from the i th input neuron to the j th hidden neuron;
ij
whh is connection weight from the i th hidden neuron to the j th hidden neuron;
ij
whout is connection weight from the i th hidden neuron to the output neuron;
ij
rj n is share of information passed by the reset gate of the j th hidden neuron at a
z j n is share of information passed by the update gate j th of the hidden neuron
yiin n is output of the i th input neuron at time n ;
yiout n is output of the i th output neuron at time n ;
y hj n is output of the j th hidden neuron at time n ;
f  , g  are activation function.</p>
        <p>To evaluate the effectiveness of the proposed model, it is necessary to select
criterion.
3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Selection of criterion for evaluating the effectiveness of the proposed model</title>
        <p>In the paper, to evaluate the parametric identification of the MGRU model, a model
adequacy criterion is chosen, which means the choice of such parameter values
as W   iijnr n, uiijnz n, wiijnh n, ihj r n, uihjz n, wihjh n, wihjout n  , which
deliver a minimum of the mean squared error (the difference between the model
output and the desired output):</p>
        <p>F 
1 P N2 2</p>
        <p>  yopuit  d pi   min .</p>
        <p>PN 2 p1 i1 W
In the paper, to evaluate the functioning of the MGRU model in test mode, a forecast
probability criterion is selected, which means the choice of such parameter values
as W   iijnr n, uiijnz n, wiijnh n, ihj r n, uihjz n, wihjh n, wihjout n  , which
deliver the maximum probability:</p>
        <p>F 
1 P 1,</p>
        <p>  round  yoput  , d p   min ,   a, b  
P p1 W 0,
a  b
a  b
where round  is function that rounds a number to the nearest integer.</p>
        <p>According to the first criterion, the methods of parametric identification of the
MGRU model are proposed in this paper.
4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Creation of method for parametric identification of the</title>
      </sec>
      <sec id="sec-2-4">
        <title>MGRU model based on local search</title>
        <p>In this paper, we first propose a method for parametric identification of the MGRU
model based on the traditional for GRU backpropagation through time (BPTT).</p>
        <p>The proposed method allows you to find a quasi-optimal vector of parameters’
values of the MGRU model and consists of the following blocks.</p>
        <p>Block 1 – Initialization:
─ set the current iteration number n to one;
─ initialization by uniform distribution over the interval 0, 1 or 0.5, 0.5 weights
 inr  n , uinz n , winh  n , i 1, N 0 , j 1, N 1 ,  hr  n , uhz  n ,
ij ij ij ij ij
whh n , i 1, N 1 , j 1, N 1 , whout n , i 1, N 1 , i 1, N 2 , where N 0 is
ij ij
the number of neurons in the input layer, N 2 is the number of neurons in the
output layer, N 1 is number of neurons in the hidden layer.</p>
        <p>Block 2 – Setting the training set  x , d  x  RN0 , d  RN2  ,  1, P , where
x is  th training input vector, d is  th training output vector, P is the power of
the training set. Number of the current pair from the training set   1 .</p>
        <p>Block 3 is Initial calculation of the output signal for the hidden layer hi n 1  0 ,
i 1, N 1 .</p>
        <p>Block 4 is Calculation of the output signal for each layer (forward propagation)
yiin n  xi , rj n  f  srj n ,</p>
        <p>N0 N1
srj n  i1 iijnr n yiin n  i1 ihjr n yih n 1 , j 1, N 1 ,</p>
        <p>z j n  f  s zj n ,

szj n  uinz n yiin n uhz n yih n1, j1, N1 ,</p>
        <p>ij ij
 N(0) N(1) 
shj(n)  (1rj(n)) winh(n)yiin(n)rj(n) wihjh(n)yih(n1) ,</p>
        <p>ij
i1
i1</p>
        <p></p>
        <p>Block 5 – Calculation of ANN error energy
1 N2
2 j1</p>
        <p>En   e2j n, ej n  yjL nd j .</p>
        <p>Block 6 – Setting up synaptic weights based on a generalized delta rule (back
propagation)
whout n1  wihjout n wihjout n , i1, N1 , j1, N2 ,
ij
winh n1  wiijnh n wiijnh n, i1, N0 , j1, N1 ,
ij
whh n1  wihjh n wihjh n , i1, N1 , j1, N1 ,</p>
        <p>ij
uinz n1  uiijnz n uiijnz n , i1, N0 , j1, N1 ,
ij
uhz n1  uihjz n uihjz n, i1, N1 , j1, N1 ,
ij
 inr n1 iijnr n iijnr n , i1, N0 , j1, N1 ,
ij
 hh n1 ihjh n ihjh n, i1, N1 , j1, N1 ,</p>
        <p>ij
where  is a parameter that determines the learning speed (for large  , learning is
faster, but the risk of getting the wrong decision increases), 0  1.</p>
        <p>wihjout n  yih n ojut n ,
wiijnh n  1 rj n yiin n jh n ,
wihjh n  rj n yih n 1 jh n ,
uiijnz n  yiin n jz n ,
uihjz n  yih n 1 jz n ,
 iijnr n  yiin n rj n ,
 ihj r n  yih n 1 rj n ,
 ojut n  f s ojut n y ojut n  d j  ,</p>
        <p> N2 
 jh n  g s hj n 1 z j n   whout n lout n  ,</p>
        <p>jl
 l1 </p>
        <p> N2 
 jz n  f  s zj n y hj n 1  g  s hj n  whout n lout n  ,
jl
 l1 
 N(1) M
 jr (n)  f (srj (n))   wihjh (n) yih (n 1)   winh (n) yiin (n)  hj (n) .</p>
        <p> i1 i1 ij </p>
      </sec>
      <sec id="sec-2-5">
        <title>Creation of a method for parametric identification of the</title>
      </sec>
      <sec id="sec-2-6">
        <title>MGRU model based on multi-agent metaheuristic search</title>
        <p>In this paper, we propose a method for parametric identification of the MGRU model
based on simulated annealing particle swarm optimization (SAPSO).</p>
        <p>The SAPSO method allows you to find a quasi-optimal vector of parameter values
for the MGRU model and consists of the following blocks.</p>
        <p>Block 1 – Initialization:
─ set the current iteration number n to one;
─ set the maximum number of iterations N ;
─ set swarm size K ;
─ set the dimension of the particle position M (corresponds to the number of the</p>
        <p>MGRU model parameters);
─ position initialization xk (corresponds to the parameters vector of the MGRU
model)
xk   xk1,</p>
        <p>, xkM  , xij   x mjax  x mjin U 0, 1  x mjin , k 1, K ,
where x mjin , x mjax are minimum and maximum values, U 0, 1 is function that
provides the calculation of a uniformly distributed random variable on a segment 0, 1 ;
─ initialization of a personal (local) best position xbest
k
xbest  xk , k 1, K ;
k
─ speed initialization  k
─ create an initial swarm of particles
 k   k1,</p>
        <p>, kM  ,  ij  0 , k 1, K ;</p>
        <p>Q   xk , xkbest , k  ;
─ determine the particle from the current population with the best position
(corresponds to the best parameters vector of the MGRU model in target function)
k*  arg km1i,nK F  xk  , x*  xk* .</p>
        <p>Block 2 – Modification of the velocity of each particle using simulated annealing
r1k  r1k1,</p>
        <p>, r1kM  , r1kj U 0, 1 , C 0, 1 , N 0, 1 , k 1, K , j 1, M ,</p>
        <p>, r2kM  , r2kj U 0, 1 , C 0, 1 , N 0, 1 , k 1, K , j 1, M ,
 k  wn k 1 n  xkbest  xk  r1 T  2 n  x*  xk  r2 T , k 1, K ,
1 n   2 n   0 exp 1 T n ,  0   0  0.5  ln 2 ,
wn  w0 exp 1 T n , w0  w0 
,
T n   T n 1 , T 0  T0 ,   N  N11 , T0  N N 1 ,
N
where N 0, 1 is function that provides the calculation of a random variable from
standard normal distribution,</p>
        <p>C 0, 1 is function that provides the calculation of a random variable from
standard Cauchy distribution,</p>
        <p>1  n is parameter controlling the contribution of the component  xkbest  xk  r1 T
to the particle’s velocity at the iteration n ,</p>
        <p> 2  n is parameter controlling the contribution of the component  x*  xk  r2 T
to the particle’s velocity at the iteration n ,</p>
        <p>w n is parameter controlling the contribution of the particle’s velocity at the
iteration n 1 to the particle’s velocity at the iteration n ,
 0 is initial value of 1  n and  2  n parameters,
w0 is initial value of w n parameter,
T  n is annealing temperature at the iteration n ,
T0 is initial annealing temperature,
 is parameter that controls the annealing temperature.</p>
        <p>The simulated annealing introduced in this work allow us to establish the inverse
correlation between the parameters 1  n ,  2  n , w n and the iteration number,
i.e. in the first iterations, the search is global, and in the last iterations, the search
becomes local. In addition, in this work, a direct correlation between the parameters T0 ,
 and the iteration number is established, which allows for automated selection of
these parameters.</p>
        <p>The choice of initial values of  0  0.5  ln 2 and w0  1
2 ln 2
isfies the conditions of a swarm convergence w  1 and w0  1 1  2  1 .
2
Block 3 – Modification of each particle’s position, considering the limitations
is standard and
satxk  xk  k , k 1, K ,
x mjin , xkj  x mjin

xkj  xkj , xkj   x mjin , x mjax  , k 1, K , j 1, M ,
x mjax , xkj  x mjax
 kj ,  kj   x mjin , x mjax 
 kj  
0,  kj x mjin , x mjax 
, k 1, K , j 1, M .</p>
        <p>Block 4 – Determining the personal (local) best position of each particle
If F  xk   F  xkbest  , then xkbest  xk , k 1, K .</p>
        <p>Block 5 – Determining the particle from the current population with the best
position
k*  arg min F  xk  .</p>
        <p>k1, K
Block 6 – Determining the global best position</p>
        <p>If F  xk*   F  x*  , then x*  xk* .</p>
        <p>Block 7 – Stop Condition
If n  N , then increase the iteration number n by 1 and go to block 2.</p>
        <p>The proposed method is intended for implementation through a multi-agent
system.
6</p>
      </sec>
      <sec id="sec-2-7">
        <title>Numerical study</title>
        <p>N0  2 MaxLenLexem  LenCode ,</p>
        <p>The parametric identification of the MGRU model was carried out for 10,000 training
implementations based on the proposed multi-agent metaheuristic search.</p>
        <p>Table 5 presents the forecast probabilities obtained for 1000 test implementations
based on the proposed MGRU model and artificial neural networks’ traditional
models.</p>
        <p>Table 6 presents the number of parameters (link weights) for the proposed MGRU
model and artificial neural networks’ traditional models, which is directly
proportional to the computational complexity of parametric identification.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Criterion</title>
      <p>Forecast probability</p>
      <p>For a complete LSTM, GRU, MGRU, the number of neurons in the hidden layer is
For JNN, ENN (SRN), BRNN, the number of neurons in the hidden layer is
calcuFor LSTM, the number of memory cells is set as M  1 .</p>
      <p>According to tables 5-6, MGRU and complete LSTM give the best forecast
probability results, but MGRU, unlike LSTM, has fewer parameters, i.e. less computational
complexity.</p>
      <p>The increase in accuracy of MGRU prediction was made possible by introducing a
multiplier (1 rj (n)) for the weighted sum of the neurons outputs in the input layer.
This allows to more accurately control the share of information coming from the input
layer and the hidden layer, as well as by using the metaheuristic determination of the
MGRU models parameters.</p>
      <p>The limitations of this work include the MGRUs full connectivity, the requirement
for more parameters than in ENN (SRN), MGRU testing only on trigrams.</p>
      <p>Like the BERT system, the proposed MGRU can work with context-free and
context-sensitive grammars, but unlike BERT, it can be used not only for English.</p>
      <p>The practical contribution of this work consists in the fact that it allows to predict
alphabetic and / or phoneme sequences through an artificial neural network, the
training of which is based on the proposed metaheuristics, which allows to increase the
accuracy of the forecast and can be used as an intermediate stage in the speech
understanding system.
7</p>
      <p>Conclusions
1. To solve the problem of insufficient quality of the natural language sequences
analysis, the corresponding neural network forecast methods were studied. To
increase the efficiency of training neural networks, metaheuristic methods were
studied.
2. The created model of the modified gated recurrent block allows for more precise
control of the share of information coming from the input layer and the hidden
layer, which increases the forecast accuracy.
3. The created method of parametric identification of the MGRU model based on
simulated annealing particle swarm optimization reduces the probability of getting
into local extremum and replaces batch training with multi-agent training, which
increases the forecast probability and the training speed.
4. The proposed method for predicting natural language constructions based on a
modified gated recurrent block can be used in various intelligent natural language
processing systems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Dominey</surname>
            ,
            <given-names>P.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inui</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A Neurolinguistic Model of Grammatical Construction Processing</article-title>
          .
          <source>Journal of Cognitive Neuroscience</source>
          .
          <volume>18</volume>
          (
          <issue>12</issue>
          ),
          <fpage>2088</fpage>
          -
          <lpage>2107</lpage>
          (
          <year>2006</year>
          ). doi:
          <volume>10</volume>
          .1162/jocn.
          <year>2006</year>
          .
          <volume>18</volume>
          .12.2088
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Khairova</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharonova</surname>
          </string-name>
          , N.:
          <article-title>Modeling a Logical Network of Relations of Semantic Items in Superphrasal Unities</article-title>
          .
          <source>In: Proc. of the EWDTS</source>
          . pp.
          <fpage>360</fpage>
          -
          <lpage>365</lpage>
          .
          <string-name>
            <surname>Sevastopol</surname>
          </string-name>
          (
          <year>2011</year>
          ). doi:
          <volume>10</volume>
          .1109/EWDTS.
          <year>2011</year>
          .6116585
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lyubchyk</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bodyansky</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rivtis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Adaptive Harmonic Components Detection and Forecasting in Wave Non-Periodic Time Series using Neural Networks</article-title>
          .
          <source>In: Proc. of the ISCDMCI'2002</source>
          . pp.
          <fpage>433</fpage>
          -
          <lpage>435</lpage>
          .
          <string-name>
            <surname>Evpatoria</surname>
          </string-name>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Du</surname>
          </string-name>
          , K.-L.,
          <string-name>
            <surname>Swamy</surname>
            ,
            <given-names>K.M.S.</given-names>
          </string-name>
          :
          <source>Neural Networks and Statistical Learning</source>
          . Springer-Verlag, London (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4471</fpage>
          -5571-3
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.: Neural</given-names>
          </string-name>
          <string-name>
            <surname>Networks. Pearson Education</surname>
          </string-name>
          , NY (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Sivanandam</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sumathi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deepa</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          :
          <article-title>Introduction to Neural Networks using Matlab 6.0</article-title>
          .
          <string-name>
            <surname>The</surname>
            <given-names>McGraw-Hill</given-names>
          </string-name>
          <string-name>
            <surname>Comp</surname>
          </string-name>
          ., Inc., New Delhi (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Attractor Dynamics and Parallelism in a Connectionist Sequential Machine</article-title>
          .
          <source>In: Proc. of the Ninth Annual Conference of the Cognitive Science Society</source>
          . pp.
          <fpage>531</fpage>
          -
          <lpage>546</lpage>
          . Hillsdale, NJ (
          <year>1986</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rumelhart</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Forward Models:
          <article-title>Supervised Learning with a Distal</article-title>
          .
          <source>Cognitive Science</source>
          .
          <volume>16</volume>
          ,
          <fpage>307</fpage>
          -
          <lpage>354</lpage>
          (
          <year>1992</year>
          ). doi:
          <volume>10</volume>
          .1016/
          <fpage>0364</fpage>
          -
          <lpage>0213</lpage>
          (
          <issue>92</issue>
          )
          <fpage>90036</fpage>
          -T
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vairappan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A Novel Learning Method for Elman Neural Network using Local Search</article-title>
          .
          <source>Neural Information Processing - Letters and Reviews</source>
          .
          <volume>11</volume>
          (
          <issue>8</issue>
          ),
          <fpage>181</fpage>
          -
          <lpage>188</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wiles</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elman</surname>
          </string-name>
          , J.:
          <article-title>Learning to Count without a Counter: a Case Study of Dynamics and Activation Landscapes in Recurrent Networks</article-title>
          .
          <source>In: Proc. of the Seventeenth Annual Conference of the Cognitive Science Society</source>
          . pp.
          <fpage>1200</fpage>
          -
          <lpage>1205</lpage>
          . Cambridge, MA (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliwal</surname>
            ,
            <given-names>K.K.</given-names>
          </string-name>
          :
          <source>Bidirectional Recurrent Neural Networks. IEEE Transactions on Signal Processing</source>
          .
          <volume>45</volume>
          (
          <issue>11</issue>
          ),
          <fpage>2673</fpage>
          -
          <lpage>2681</lpage>
          (
          <year>1997</year>
          ). doi:
          <volume>10</volume>
          .1109/78.650093
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Baldi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frasconi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Soda</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pollastri</surname>
          </string-name>
          , G.:
          <article-title>Exploiting the Past and the Future in Protein Secondary Structure Prediction</article-title>
          .
          <source>Bioinformatics</source>
          .
          <volume>15</volume>
          (
          <issue>11</issue>
          ),
          <fpage>937</fpage>
          -
          <lpage>946</lpage>
          (
          <year>1999</year>
          ). doi:
          <volume>10</volume>
          .1093/bioinformatics/15.11.937
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.: Long</given-names>
          </string-name>
          <string-name>
            <surname>Short-Term Memory</surname>
          </string-name>
          .
          <source>Technical Report FKI-207-95</source>
          , Fakultat fur Informatik, Technische Universitat Munchen. doi:
          <volume>10</volume>
          .1162/neco.
          <year>1997</year>
          .
          <volume>9</volume>
          .8.
          <fpage>1735</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Gers</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Long Short-Term Memory in Recurrent Neural Networks</article-title>
          .
          <source>PhD thesis</source>
          , Ecole Polytechnique Federale de Lausanne.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merrienboer</surname>
            , van
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gulcehre</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bougares</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwenk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation</article-title>
          .
          <source>In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <fpage>1724</fpage>
          -
          <lpage>1734</lpage>
          . Qatar,
          <string-name>
            <surname>Doha</surname>
          </string-name>
          (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>D14</fpage>
          -1179
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Dey</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salem</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          :
          <article-title>Gate-Variants of Gated Recurrent Unit (GRU) Neural Networks</article-title>
          . (
          <year>2017</year>
          ). arXiv:
          <volume>1701</volume>
          .
          <fpage>05923</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Jaeger</surname>
          </string-name>
          , H.:
          <article-title>A Tutorial on Training Recurrent Neural Networks</article-title>
          ,
          <string-name>
            <surname>Covering</surname>
            <given-names>BPPT</given-names>
          </string-name>
          ,
          <article-title>RTRL, EKF and the “Echo State Network” Approach</article-title>
          .
          <source>GMD Report 169</source>
          ,
          <string-name>
            <surname>Fraunhofer</surname>
            <given-names>Institute AIS</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Jaeger</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lukosevicius</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popovici</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siewert</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Optimization and Applications of Echo State Networks with Leakyintegrator Neurons</article-title>
          .
          <source>Neural Networks</source>
          .
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <fpage>335</fpage>
          -
          <lpage>352</lpage>
          (
          <year>2007</year>
          ). doi:
          <volume>10</volume>
          .1016/j.neunet.
          <year>2007</year>
          .
          <volume>04</volume>
          .016
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Maass</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natschläger</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markram</surname>
          </string-name>
          , H.:
          <article-title>Real-Time Computing without Stable States: a New Framework for Neural Computation Based on Perturbations</article-title>
          .
          <source>Neural Computation</source>
          .
          <volume>14</volume>
          (
          <issue>11</issue>
          ),
          <fpage>2531</fpage>
          -
          <lpage>2560</lpage>
          (
          <year>2002</year>
          ).
          <source>doi: 10.1162/089976602760407955</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Jaeger</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maass</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prıncipe</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Special Issue on Echo State Networks and Liquid State Machines</article-title>
          .
          <source>Editorial. Neural Networks</source>
          .
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <fpage>287</fpage>
          -
          <lpage>289</lpage>
          (
          <year>2007</year>
          ). doi:
          <volume>10</volume>
          .4249/scholarpedia.2330
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Talbi</surname>
          </string-name>
          , El-G.:
          <article-title>Metaheuristics: from Design to Implementation</article-title>
          . Wiley &amp; Sons, Hoboken, New Jersey (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Engelbrecht</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          :
          <source>Computational Intelligence: an Introduction</source>
          . Wiley &amp; Sons, Chichester, West
          <string-name>
            <surname>Sussex</surname>
          </string-name>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Zbigniew</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Three Parallel Algorithms for Simulated Annealing</article-title>
          .
          <source>In: Proceedings of the 4th International Conference on Parallel Processing and Applied Mathematics-Revised Papers (PPAM'01)</source>
          . pp.
          <fpage>210</fpage>
          -
          <lpage>217</lpage>
          . Springer-Verlag, Berlin, Heidelberg (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Loshchilov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>CMA-ES with Restarts for Solving CEC 2013 Benchmark Problems</article-title>
          .
          <source>In: Proceedings of the IEEE Congress on Evolutionary Computation (CEC</source>
          '
          <year>2013</year>
          ). pp.
          <fpage>369</fpage>
          -
          <lpage>376</lpage>
          . Cancun,
          <string-name>
            <surname>Mexico</surname>
          </string-name>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Byrne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hemberg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brabazon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Neill</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A Local Search Interface for Interactive Evolutionary Architectural Design</article-title>
          .
          <source>In: Proceedings of the International Conference on Evolutionary and Biologically Inspired Music</source>
          and
          <string-name>
            <surname>Art (Evo-MUSART</surname>
          </string-name>
          '
          <year>2012</year>
          ). pp.
          <fpage>23</fpage>
          -
          <lpage>34</lpage>
          . Springer, Berlin, Heidelberg (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>