<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Intelligent Forecasting in Multi-Criteria Decision-Making</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>v Borys</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuriy Kon</string-name>
          <email>halyna.kondratenko@chmnu.edu.ua</email>
          <email>yuriy.kondratenko@chmnu.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelligent Information Systems Department, Petro Mohyla Black Sea National University</institution>
          ,
          <addr-line>68th Desantnykiv Str., 10, Mykolaiv, 54003</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, the subject of the research is the intelligent forecasting in multi-criteria decision-making, in particular, methods of prediction of time series: traditional models of autoregression, smoothing techniques and machine learning methods with the use of artificial neural networks and deep learning. Time series of Amazon share prices from the official sources over the past years serve as a base for exploration. The purpose of the work is to find out the parameters that influence the efficiency and accuracy of models for analysis and forecast the share prices. Such assessment is complicated and vital because various types of criteria, in particular, sales and profitability, market indexes, exchange rates and general trends, etc., usually influence the decision-making process. The methods of research include the consideration and analysis of prediction methods using specific metrics. The relevance of the topic assumes that the accurate prediction at the financial market may contribute to the financial benefits for companies, government and other players on stock exchanges. A comparative analysis of the considered forecasting methods was conducted. It allows choosing the most appropriate intellectual method for increasing the efficiency of a specific share price prediction. The further development of the subject of research includes ensemble-learning methods for neural networks, feature engineering, collecting a more extensive data set for forecasting.</p>
      </abstract>
      <kwd-group>
        <kwd>time series</kwd>
        <kwd>autoregression</kwd>
        <kwd>smoothing</kwd>
        <kwd>artificial neural networks</kwd>
        <kwd>convolutional neural networks</kwd>
        <kwd>recurrent neural networks</kwd>
        <kwd>decision-making</kwd>
        <kwd>multi-criteria approach</kwd>
        <kwd>stock market</kwd>
        <kwd>share price</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The analysis of the behavior of stock prices is characterized by the ambiguous
behavior of the process, which is usually affected by many factors (trend, seasonality, the
geopolitical situation, etc.). Forecasting is a key point when making investment
decisions. The ability to predict the behavior of stock for making final decisions allows
you to make the best choice which otherwise might be unsuccessful [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        However, not all investors successfully profit from their investment. This is
because the stock price is constantly fluctuating, and at any moment the price may fall
below the price at which it was purchased. Therefore, to predict how the financial
market will behave is one of the most difficult tasks in the economy. In this
prediction, it is necessary to consider various factors, such as physical, psychological,
rational and irrational behavior, and the like. All these aspects lead to the conclusion
that stock prices are very volatile and it’s difficult to predict them with a high degree
of accuracy. However, this task is urgent for the world and the whole international
economy, since the possibility of accurate prediction of the value of the stock is
closely linked to the financial gain of the companies, the government or personal
capital and the formation of the more rational financial behavior [
        <xref ref-type="bibr" rid="ref2 ref3">2-3</xref>
        ].
      </p>
      <p>Accurate predicting the value of assets on the exchange will also help to reduce
investment risk and protect the investment income from the market volatility.
2</p>
      <p>
        Related Works and Problem Statement
A stock market (hereinafter the SM) is an organized market where the securities
owners make the agreements of purchase and sale with the members of the SM as
intermediaries. The prices of these securities are determined by supply and demand, and
the process of sale is governed by the rules and regulations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>There are three different approaches to research and predict the asset prices on the
stock market: technical, technological and fundamental analysis. In this paper, the
combination of the first two approaches was used.</p>
      <p>
        Technical analysis is the study of the dynamics of the main indicators of the market
with the help of graphical methods to predict the future direction of their movement.
A significant number of participants in stock and OTC markets use technical analysis
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. A successful trader O. Elder has spoken figuratively on this subject: "Technical
analysis is related to public opinion polls. It is a combination of science and art. The
scientific part consists in using statistical methods and computers; the creative part is
the interpretation of the data."
      </p>
      <p>Technological analysis originated in the era of the application of computer
technology in business and other fields. The success of its application in solving complex
tasks is largely determined by the possibilities of modern information technologies.
The result of technological research is usually a selection of well-defined alternatives.
Therefore, the origins of this analysis, its methodological concepts lie in those
disciplines that deal with decision-making: Theory of Operations and General
Management Theory.</p>
      <p>The technological analysis includes the ability to analyze, to predict, and to design
decision-making in complex systems of various nature, which are based on the time
series.</p>
      <p>
        A time series can be represented by four components: trend, seasonal variations,
cyclical variations and irregular factors [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The classical approach to the
construction of the time series model is to schedule it into several components, each of which
examines specific methods. A trend is a general systematic linear or non-linear
component that can change over time. The seasonal and cyclical components are
periodically repeating components. These components are not always present simultaneously
in the time series. In our case, there are no seasonal and cyclical components.
3
      </p>
      <p>
        Basic Concepts and Intellectual Forecasting Methods for
Stock Prices Prediction
Since a prediction based on statistical data collected with the same interval is
required, we are dealing with the time series. Therefore, the time series models will be
considered, and those that suit the task of forecasting data will be selected. The most
popular forecasting time series models are autoregressive models, autoregressive
models with a moving average and models derived from them. Thus, we will focus on
these models [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        After selecting the specific methods of solving tasks, it is necessary to directly
conduct a prediction of selected characteristics. Constructing the desired models will
be implemented by their gradual complication, in other words, we will move from
simple to complex. In this paper, models and methods are evaluated using MAE and
MSE metrics [
        <xref ref-type="bibr" rid="ref1 ref7">1, 7</xref>
        ].
      </p>
      <p>MAE measures the average absolute value of errors in a set of predictions for
continuous variables. Assuming that yˆt is the values of the time series forecast in period
t, the metric is given by (1):
MAE is a linear score, which means that all individual differences are resolved to the
same average. MSE is the difference between the forecast and corresponding
observed values squared at each iteration. Because errors rise to the square before they
are averaged, MSE has a relatively high weight to large bias. This means that MSE is
most useful when large errors are particularly undesirable, which is consistent with
the objectives of this work:
(1)
(2)
Both metrics, MAE and MSE, can take values from 0 to ∞ and do not take the
direction of errors into account. The smaller is the value the index takes, the more accurate
is the forecast.</p>
      <p>Smoothing is an important and widespread method of financial market forecasting.
Smoothing methods are used to reduce the influence of random components (random
fluctuations) in time series. They provide an opportunity to obtain more “pure” values
that consist only of deterministic components. Some of the methods are aimed at
recovering some of the components such as trend.</p>
      <p>The authors present 5 smoothing methods typically found when forecasting
financial data: the simple moving average; the weighted moving average; the exponential
MAE   yt  yˆt .</p>
      <p>t
MSE   yt  yˆt 2 .</p>
      <p>
        t
moving average; the double exponential moving average; the triple exponential
moving average [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The basic assumption of these methods is that the fluctuations in past values
represent random deviations from a smooth curve, which can be extrapolated to create a
forecast.</p>
      <p>
        The autoregressive model is another effective tool for understanding and predicting
future values of the time series, which includes devolution of a variable on values of
series in the past. The importance of ARMA models lies in their flexibility and in
their ability to describe almost all features of the stationary time series.
Autoregressive of these models describe how consecutive observations in time affect each other,
while parts of the moving averages capture some possible unobserved upheavals, and
that allows simulating various phenomena that can be observed in a variety of fields
from biology to finance [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ].
      </p>
      <p>
        The main idea of autoregression methods is that future values of the time series
cannot deviate to higher or lower than the previous values of the time series whatever
the reasons that caused those deviations are. The paper has presented such models of
autoregression as simple AR (Autoregression Model), ARMA (Autoregressive
Moving Average Model) and ARIMA (Autoregressive Integrated Moving Average Model)
[
        <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
        ].
      </p>
      <p>
        We’ve also proved that Artificial Neural Networks (ANN) have a significant
advantage in time series forecasting because they are endowed with the capacity to solve
complex problems of forecasting [
        <xref ref-type="bibr" rid="ref10 ref2 ref9">2, 9, 10</xref>
        ].
      </p>
      <p>The output value of the neural network is defined mathematically as:</p>
      <p>q  p 
yt   0   j g 0i   ij yi    t ,</p>
      <p>j1  i1 
where p is the number of the input variables, q is the number of the hidden nodes,
 j and  ij are the weights,  t is the random noise.</p>
      <p>
        As a function of g the following functions can be used [
        <xref ref-type="bibr" rid="ref11 ref14 ref2">2, 11, 14</xref>
        ].
      </p>
      <p>
        Sigmoidal function:
e x  ex .
(3)
(4)
(5)
In some literature, this is called a logistic function. This nonlinear function is one of
the most common activation functions for deep learning [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Hyperbolic tangent:
This function gives the best result for multilayer neural networks compared with the
sigmoidal function. However, the function does not solve the problem of the
vanishing gradient. The main benefit offered by this function is that it is centered relative to
zero, which helps in the process of error backpropagation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Softmax function:</p>
      <p>The Softmax function is another type of activation function used in neural computing.
It is used to calculate the probability distribution of a vector of real numbers and gets
the value range from 0 to 1.</p>
      <p>
        ReLU function:
This function is considered to be the most successful and widely used transfer
function. ReLU shows the best productivity in deep learning compared to sigmoid and
hyperbolic tangent. ReLU is a nearly linear function and therefore preserves the
properties of linear models, which make them easily optimized by methods of gradient
descent. The main advantage of using ReLU in calculus is that this function does not
require computing the exponent or dividing, therefore it guarantees a more rapid
execution [
        <xref ref-type="bibr" rid="ref11 ref13">11, 13</xref>
        ].
      </p>
      <p>
        ANN has been and continues to be actively used in the financial markets; one of
the main advantages of ANN that make them so popular as harbingers of the market is
the natural nonlinearity that allows them to learn nonlinear mapping and correlation
of data. ANN also works on data, can be trained in real-time; they are highly adaptive,
easy to retrain in the event of market fluctuations and, finally, do well with data that
contain a certain number of errors [
        <xref ref-type="bibr" rid="ref14 ref15 ref16 ref17">14-17</xref>
        ]. ANN mechanism implies minimum
participation of the analyst in the model shaping, as far as the learning ability is
characteristic of all neural network models and learning algorithms adapt (adjust) the
weights according to the structure of the data presented for training.
      </p>
      <p>Consider the most common optimization techniques.</p>
      <p>
        Gradient descent [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]:
xt1  xt   f xt  .
(8)
According to this method, the steps proportional to the opposite gradient value are
taken for minimizing functions. The parameter to this method is a descent speed α.
When reaching a large value α will possibly make large steps for finding the
minimum, but there is a risk of skipping the lowest point. At a very low speed of learning,
the algorithm will move confidently in the direction of the negative gradient, but the
implementation, in this case, will take a long time.
      </p>
      <p>
        The method of stochastic gradient descent [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This is a type of gradient descent,
which handles 1 element for learning at each iteration. So, parameters are updated
even after one iteration, which processed only one value of the variable. This allows
optimizing the target function much faster than ordinary gradient descent. But if the
(6)
(7)
number of training data is very large, the number of iterations is therefore sufficiently
large.
      </p>
      <p>
        Mini-batch gradient descent [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. This is the type of gradient descent, which is
faster than batch and stochastic gradient descent. Let there be the m number of input
data, then in one iteration of this method b&lt;m elements will be processed.
Consequently, even if the number of training data is large, it is processed in smaller training
groups at once. Thus, the method works for a large amount of training data by
optimizing an objective function with less number of iterations.
      </p>
      <p>Gradient descent with momentum is the method that helps to accelerate a descent
in the corresponding direction and dampens the oscillations, approaching the local
minimum. This is implemented by adding a part of the update vector from the
previous step to the vector of the current update.</p>
      <p>
        Adam (Adaptive Moment Estimation) [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Adam optimizer is one of the most
popular optimization algorithms of gradient descent since it is computationally
efficient and has very little memory requirements. This method calculates an individual
adaptive learning rate for each parameter from the evaluation of the first and second
moments of the gradients.
      </p>
      <p>In the simulation of isolated time series using ANN models a transformation of the
original data for increasing the number of input neurons and, consequently, increasing
predictive ability is allowed. Modeling of the time series using the ANN mechanism
consists in the formation of ANN as a certain structure, which describes the behavior
of the system under study in points in time, and forecasting is the prediction of the
future behavior of the system in the background.</p>
      <p>To sum it up, this work explains the mathematical foundations of neural networks,
namely multilayer perceptron of Rumelhart, convolutional neural networks and
recurrent neural networks. The correctly chosen input is significant. In this paper, only the
value of the stock price in the past was taken as the initial data.
4</p>
      <p>
        Practical Implementation of the Described Methods and
Techniques
This section describes the architecture of the developed program and the final
comparative results of all models. Compared to the classical mathematical methods, this
work has allowed confirming the viability and the feasibility of using artificial neural
networks for further study of their application in the financial markets [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>For this study, AMAZON stock statistics were taken from the official NASDAQ
(National Association of Securities Dealers Automated Quotation) website for the
past 5 years. Dataset (Fig. 1) is daily information on the price of shares at the
beginning and at the end of the day, the minimum and maximum value of a share during
the day, and the number of shares sold. The above parameters can be considered as
criteria since they have an optimization direction (maximum or minimum). Moreover,
the considered methods and approaches are also used to solve multi-criteria
decisionmaking problems with a more complex structure and the number of criteria.
Therefore, this task of forecasting the stock price with the considered methods and
approaches can be considered in the context of multi-criteria decision-making.</p>
      <p>
        The moving average method was implemented for various values of the N
parameter of the previous time points number that had to be taken into account when building
a forecast (Fig. 2). It was checked that decreasing the N window length model shows
the more accurate result on the test sample, indicating that the latest data is the most
influential to the prediction, i.e. prediction considering the past 5 days is more
accurate than the forecast, which takes the last 20 days into account.
For the exponential moving average (Fig. 3) parameter is the level of the α smoothing
coefficient, which represents the degree of reduction of weighting from 0 to 1. The
lower the level of smoothing is, the more accurate is the forecast, as every previous
reading weighs more [
        <xref ref-type="bibr" rid="ref2 ref3">2-3</xref>
        ].
      </p>
      <p>
        This part of the paper also contains a comparison of the results of forecasting (by
terms of MAE and MSE) stock prices by various types of neural networks (MLP,
CNN, LSTM) with an additional comparison of their models, which have different
structures and parameters. Therefore, a detailed review of the architecture of neural
networks, justification for choosing the number of hidden layers, the number of
neurons in them, is not necessary.
For the data with a distinct trend, which corresponds to the input operation, the
method of double exponential smoothing that is the recursive application of
exponential filter twice, gives the best result. This method has provided an additional
parameter β, which is responsible for the smoothing of the trend (Fig. 4). The combination of
the α and β pair has adjusted the accuracy and the quality of the forecast [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
The results of implementation of all smoothing methods are summarized in Table 1.
      </p>
      <p>
        The autoregressive model [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] is an effective tool for understanding and predicting
future values of the time series, which includes the devolution of a variable on values
of series in the past. Autoregressive of these models describe how consecutive
observations in time affect each other, while the parts of the moving averages capture some
possible unobserved upheavals. For the considered models, the autoregressive Akaike
coefficient allowed the software to choose the model that gives the best forecast.
Therefore, in the case of the dataset, it was the model ARIMA (2,1,1).
The execution results of all three models were summarized in a single comparative
Table 2.
The simplest type of the neural network is a single-layer perceptron network (usually
a multilayer perceptron of Rumelhart), which consists of a single layer of output
nodes, and the inputs provided directly to the outputs via a series of weights.
Universal approximation theorem for neural networks states that each continuous function,
which maps intervals of real numbers to a certain initial interval of real numbers, can
be arbitrarily closely approximated by a multilayer perceptron with only one hidden
layer [
        <xref ref-type="bibr" rid="ref12 ref2">2, 12</xref>
        ]. The main characteristic of the multilayer perceptron is its architecture,
namely the number of hidden layers, nodes and activation functions. Besides, since
the study outcome partially depends on the initialization of variables, each model was
trained separately 5 times, and all the characteristics for a comparative table are
average values. For a more in-depth study of the model 5 architectures were implemented
in the paper, the accuracy of each of them is represented in Table 3.
      </p>
      <p>
        According to the comparative results (Table 3), the dependence of the quality of
the forecast on the number of nodes in the hidden layer, the optimization method and
the activation function is completely traced. The best result was the MLP 20-100-1
model with the Adam optimization algorithm and hyperbolic tangent as an activation
function. Here is a graph of the price of stocks sold predicted by this model (Fig. 5, a),
as well as a graph of the forecast error (Fig. 5, b).
Convolutional neural network [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] is a type of artificial neural network, in which the
pattern of connectivity between neurons is inspired by the organization of the visual
cortex of animals, individual neurons that are arranged in such a way that they
respond to overlapping regions of the field. The main advantage of convolutional neural
networks is that we use the convolutional layers to identify the signs of a network that
allows training of a neural network without complex pre-processing since useful
features are learned during training [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Unlike multilayer perceptrons, convolutional
neural networks have a complex structure, since they need to gauge the number and
type of layers, the number of nodes in each layer, the optimization method, the
activation function and the number of filters. So CNN-20-200-3 will designate 20 knots at
the entrance as the first layer, 200 filters sized 3x3, MaxPooling layer and 2
fullyconnected layers. In this work, 5 different CNN architectures were tested, the results
are shown in Table 4.
      </p>
      <p>The CNN-10-256-3 model with the Adam optimization algorithm and ReLU
activation function showed the best result.</p>
      <p>
        A recurrent neural network (RNN) is an artificial neural network, the neurons of
which transmit feedback signals to each other [24]. The idea of RNN is to use
consistent information. In traditional neural networks, we assume that all the inputs (and the
outputs) are independent of each other. But for many tasks, this is not the best idea.
This type of neural network is well suited to predict the value of the stock, as further
steps may depend on the past.
For the implementation of the forecasting recurrent neural networks, the LSTM model
(Long Short-Term Memory Units) was chosen [25]. LSTMs help to save the bug,
which can be spread through time and layers. Maintaining a more constant error, they
allow the networks to continue exploring for many time steps [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. To analyze the
effect of parameters on the prediction value of the shares 5 variants of the LSTM
architecture summarized in the following Table 5 were implemented.
LSTM-20-3030-30-1 was used to identify neural LSTM networks with initial data of the previous
20 days, the number of nodes 30 on the first, the second and the third layers of LSTM
and one fully connected layer with the resulting output of one element.
      </p>
      <p>Train- Training Test Test
Algo</p>
      <p>ing MSE MAE MSE rithm
MAE
0.0074 1.0564e-04 0.0198 0.0005 adam
The best result was shown by the LSTM model 10-30-30-30-1 with the Adam
optimization algorithm and ReLU activation function, and with the least lag window.</p>
      <p>The performance of classical algorithms and machine-learning algorithms in
section 4 can be summarized in a comparative Table 6.
For investors, capital management is becoming increasingly important today.
Professional investment managers and individual investors tend to have effective tools for
understanding the trends in the stock market, to minimize investment risk and
increase their profits. Some people believe that it is very difficult to predict stock prices.
However, in the real business world, successful traders carry out thousands of
transactions every year.</p>
      <p>Since the very beginning of the financial operations on the stock exchange, people
developed methods to predict the value of assets in the future. With advances in
computing technology and artificial intelligence, the accuracy of the methods is growing
every day.</p>
      <p>In this paper, the authors compared the indicators of forecasting between the neural
network and the classic time series forecast method, namely, the value of the
company's shares on the stock exchange. The analysis showed that the neural network
models described in this study showed a very much better able to accurately predict,
and therefore confirmed the viability and the feasibility of using artificial neural
networks for further study of their application in the financial markets. The predictive
power of the model was influenced by different factors, depending on the architecture
and input data.
24. Hammer, B.: Learning with recurrent neural networks. Springer, London (2000)
25. Manaswi, N.: RNN and LSTM. In: Deep Learning with Applications Using Python, pp.
115–126. Apress, Berkeley, CA (2018). doi: 10.1007/978-1-4842-3516-4_9</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kingdon</surname>
          </string-name>
          , J.:
          <source>Intelligent Systems and Financial Forecasting</source>
          . Springer, London (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Palit</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popovic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Computational Intelligence in Time Series Forecasting.
          <source>Theory and Engineering Applications</source>
          . Springer, London (
          <year>2005</year>
          ). doi:
          <volume>10</volume>
          .1007/1-84628-184-9
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Guerard</surname>
          </string-name>
          , J.:
          <article-title>Introduction to Financial Forecasting in Investment Analysis</article-title>
          . Springer, New York (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4614</fpage>
          -5239-3
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pavlov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilipenko</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <article-title>Krivov'yazyuk, I.: Securities in Ukraine: Students Book</article-title>
          . Condor,
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>(in Ukrainian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Sokhatska</surname>
            ,
            <given-names>O.: Stock</given-names>
          </string-name>
          <string-name>
            <surname>Exchange. Carte Blanche</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ternopil</surname>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>(in Ukrainian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Box</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenkins</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reinsel</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <source>Time Series Analysis: Forecasting &amp; Control</source>
          . John Wiley &amp; Sons, New Jersey (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          : Time Series Analysis: Univariate and
          <string-name>
            <given-names>Multivariate</given-names>
            <surname>Methods</surname>
          </string-name>
          . Pearson, New York (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Sitte</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sitte</surname>
          </string-name>
          , J.:
          <article-title>Neural Network Systems Technology in the Analysis if Financial Time Series</article-title>
          . In: Leondes, C.T. (eds.)
          <source>Intelligent Knowledge-Based Systems</source>
          , pp.
          <fpage>1564</fpage>
          -
          <lpage>1615</lpage>
          . Springer, Boston, MA (
          <year>2005</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4020</fpage>
          -7829-3_
          <fpage>43</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kondratenko</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordienko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Neural Networks for Adaptive Control System of Caterpillar Turn</article-title>
          .
          <source>In: Annals of DAAAM for</source>
          <year>2011</year>
          &amp;
          <article-title>Proceeding of the 22th Int</article-title>
          .
          <source>DAAAM Symp. "Intelligent Manufacturing and Automation"</source>
          , pp.
          <fpage>0305</fpage>
          -
          <lpage>0306</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gomolka</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dudek-Dyduch</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondratenko</surname>
            ,
            <given-names>Y.P.</given-names>
          </string-name>
          :
          <article-title>From homogeneous network to neural nets with fractional derivative mechanism</article-title>
          . In: Rutkowski,
          <string-name>
            <surname>L.</surname>
          </string-name>
          et al. (eds.)
          <source>International Conference on Artificial Intelligence and Soft Computing (ICAISC-2017)</source>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , pp.
          <fpage>52</fpage>
          -
          <lpage>63</lpage>
          , Zakopane, Poland (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nwankpa</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ijomah</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gachagan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marshall</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Activation Functions:
          <article-title>Comparison of Trends in Practice and Research for Deep Learning</article-title>
          . arXiv:
          <year>1811</year>
          .
          <article-title>03378 [cs</article-title>
          .LG] (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The Application of Neural Networks in the Forecasting of Share Prices</article-title>
          . Finance and
          <string-name>
            <given-names>Technology</given-names>
            <surname>Publishing</surname>
          </string-name>
          , New York (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Di</given-names>
            <surname>Persio</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Artificial Neural Networks architectures for stock price prediction: comparisons and applications</article-title>
          .
          <source>International journal of circuits, systems and signal processing</source>
          <volume>31</volume>
          (
          <issue>10</issue>
          ),
          <fpage>404</fpage>
          -
          <lpage>405</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kushneryk</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondratenko</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sidenko</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Intelligent dialogue system based on deep learning technology</article-title>
          .
          <source>In: 15th International Conference on ICT in Education</source>
          , Research, and Industrial Applications:
          <source>PhD Symposium (ICTERI 2019: PhD Symposium)</source>
          , vol.
          <volume>2403</volume>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>62</lpage>
          , Kherson, Ukraine (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sidenko</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Filina</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondratenko</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chabanovskyi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kondratenko</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Eye-tracking technology for the analysis of dynamic data</article-title>
          .
          <source>In: IEEE 9th International Conference on Dependable Systems, Services and Technologies (DESSERT)</source>
          , pp.
          <fpage>479</fpage>
          -
          <lpage>484</lpage>
          , Kiev, Ukraine (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1109/DESSERT.
          <year>2018</year>
          .8409181
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Cryer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chan K</surname>
          </string-name>
          .-S.: Time Series Analysis. With Applications in R. Springer, New York (
          <year>2008</year>
          ). doi:
          <volume>10</volume>
          .1007/978-0-
          <fpage>387</fpage>
          -75959-3
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Beran</surname>
          </string-name>
          , J.:
          <source>Mathematical Foundations of Time Series Analysis. A Concise Introduction</source>
          . Springer, Cham (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -74380-6
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Paper</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Gradient Descent</article-title>
          .
          <source>In: Data Science Fundamentals for Python and MongoDB</source>
          , pp.
          <fpage>97</fpage>
          -
          <lpage>128</lpage>
          . Apress, Berkeley, CA (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4842</fpage>
          -3597-
          <issue>3</issue>
          _
          <fpage>4</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Ketkar</surname>
          </string-name>
          , N.:
          <article-title>Stochastic Gradient Descent</article-title>
          .
          <source>In: Deep Learning with Python</source>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>132</lpage>
          . Apress, Berkeley, CA (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4842</fpage>
          -2766-
          <issue>4</issue>
          _
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Danner</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jelasity</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Fully Distributed Privacy Preserving Mini-batch Gradient Descent Learning</article-title>
          . In: Bessani,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Bouchenak</surname>
          </string-name>
          , S. (eds.)
          <article-title>Distributed Applications and Interoperable Systems</article-title>
          .
          <source>DAIS 2015. Lecture Notes in Computer Science</source>
          , vol.
          <volume>9038</volume>
          , pp.
          <fpage>30</fpage>
          -
          <lpage>44</lpage>
          . Springer, Cham (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -19129-
          <issue>4</issue>
          _
          <fpage>3</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Jiang</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>F</given-names>
          </string-name>
          .:
          <article-title>A Hybrid Algorithm of Adaptive Particle Swarm Optimization Based on Adaptive Moment Estimation Method</article-title>
          . In: Huang,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Bevilacqua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Premaratne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Gupta</surname>
          </string-name>
          , P. (eds.)
          <source>Intelligent Computing Theories and Application. ICIC 2017. Lecture Notes in Computer Science</source>
          , vol.
          <volume>10361</volume>
          , pp.
          <fpage>658</fpage>
          -
          <lpage>667</lpage>
          . Springer, Cham (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Olson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Autoregressive Models</article-title>
          .
          <source>In: Predictive Data Mining Models. Computational Risk Management</source>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>69</lpage>
          . Springer, Singapore (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Ayyadevara</surname>
          </string-name>
          , V.:
          <article-title>Convolutional Neural Network</article-title>
          .
          <source>In: Pro Machine Learning Algorithms</source>
          , pp.
          <fpage>179</fpage>
          -
          <lpage>215</lpage>
          . Apress, Berkeley, CA (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4842</fpage>
          -3564-
          <issue>5</issue>
          _
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>