<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>ORCID:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Visualization of the Epidemics Forecasting Results</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nataliya Shakhovska</string-name>
          <email>nataliya.b.shakhovska@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ihor Darmoriz</string-name>
          <email>ihor.darmoriz.kn.2017@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yarosvav Vyklyuk</string-name>
          <email>yaroslav.vyklyuk@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yurii Kryvenchuk</string-name>
          <email>yurii.p.kryvenchuk@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavlo Pukach</string-name>
          <email>pavlo.p.pukach@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informatics &amp; Data-Driven Medicine</institution>
          ,
          <addr-line>11, 2021, Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Lviv, Ukraine, 79013</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1873</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Modeling and forecasting of time series is one of the most importance for various practical applications. Many things are more or less time-dependent. Its analysis can forecast the future behavior to take some action for better results in the future. Research purpose is to develop a software product that has the ability to forecast the spread of the epidemic in relation to its specific features. The comparison of linear model, Convolution neural network and Recurrent neural network for epidemic forecasting is given. The spread of epidemics occurs over a period of time where anybody can see trends of some features during some time. Because the result is influenced by a large number of factors, and the training took place only on a short history, the results are of high quality because the MAPE error does not exceed 30% with a prediction for all characteristics.</p>
      </abstract>
      <kwd-group>
        <kwd>machine learning</kwd>
        <kwd>forecasting</kwd>
        <kwd>time series</kwd>
        <kwd>epidemic</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>

variability of strains,
method of distribution,</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Epidemics produced by infections and viruses usually come to a first-place amount of the large-scale
disasters and catastrophes that have attended the entire history of humankind, on a par with starvation,
wars, man-made and natural disasters. According to the World Health Organization (WHO), severe
respiratory infections account for 60-70% of the total morbidity of the population, with a tendency to
develop complexities and chronicity of the process. Due to the extreme variability of the pathogen,
acute respiratory infections remain an uncontrolled infection. Another example is coronavirus disease
affected by the new virus SARS-CoV-2 (COVID-19). Nearly 241 million people worldwide have
contracted COVID 19 (https://index.minfin.com.ua/ua/reference/coronavirus/geography/). Of these,
more than 17 million are ill at this moment, and more than 21 million have been cured. More than 4
million people died from the disease. In total, the disease was detected in 203 countries.</p>
      <p>The nature of diseases caused by infections and viruses (even with known treatment prevention
schemes) depends of numerous factors, namely:
parameters of the distribution area: climatic conditions, infrastructure and connections
between towns and inside towns, quality of medical care, the most common life style,
chronic diseases essential in this area, political situation, etc.</p>
      <p>That is why developing simulation models of the spread and character of morbidity and new cases
of various infections and viruses is a problematic scientific task. The main characteristics of this task
are the following:
multicriteria: type of spread (epidemic spread, controlled spread in a mild form of the
disease), initial parameters, the distribution territory,
EMAIL:
(NS);
(YaV),</p>
      <p>2021 Copyright for this paper by its authors.
 time dependence,
 simulation interval,
 variety of input data.</p>
      <p>Therefore, it is necessary to develop a system that is more sensitive to changes in the spread of the
disease to predict its reach in the future and monitor other changes that may occur during an epidemic
(new cases, recovery, mortality, etc).</p>
      <p>The aim of the work is to develop a model and system based on it that would make it possible to
monitor and predict the spread of the epidemic on the basis of various characteristics.</p>
      <p>The main contributions of this paper are the following:
 The new schema of recurrent neural network for COVID-19 infection forecasting is developed.
 To increase the predictive accuracy, the clustering is used on the preprocessing stage. It allows
to reduce the influence of data heterogeneity due to the presence of several locations.</p>
      <p>This paper is organized into several sections. In State of the art section, the methods of times series
analysis are given. In the section #3 “Methods and means”, new schema of recurrent neural network is
proposed. The fourth section presents result of proposed methods and gives data interpretation. The last
section concludes the paper.</p>
    </sec>
    <sec id="sec-3">
      <title>2. State of the Art</title>
      <p>
        After researching time series predictions, we can conclude that the usage of neural networks is a
new application, as over the past century, a large number of linear algorithms have been developed to
analyze and predict time series, including ARMA [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ], ARIMA [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], VAR [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], HWES [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and others.
Although these types have been quite widespread, they may not always be effective enough. Their
disadvantages include the following points:
 They require complete data. Some missing values can actually affect the model. But there are
also ways to deal with missing data.
 They rely on linear relationships. In many traditional models, their assumptions are based on a
linear basis.
 They usually only deal with one-dimensional data. For example, they can analyze a time series
for a single characteristic (such as virus mortality), although when dealing with epidemics, we
analyze several types of data.
 They usually do not work well in the long run.
      </p>
      <p>
        Convolutional neural network (CNN) is a class of deep artificial neural networks that has been
successfully used in the analysis of visual images [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. They are mainly used for work or image analysis,
but in some cases they can be used effectively for time series. An important characteristic for the use
of this network is the correlation of several types of data in the analysis of the series.
      </p>
      <p>
        A recurrent neural network (RNN) is a class of artificial neural networks where connections between
nodes form a directional graph along a time sequence [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. This allows him to demonstrate temporal
dynamic behavior. A well-known RNN is long-term memory or LSTM, and it has the ability to solve
time series problems. LSTM networks eliminate the need for a predefined time window due to the
ability to study long-term correlations in different sequences and are able to accurately model complex
multidimensional sequences. The advantages and disadvantages of each of the models are given in
Table 1.
      </p>
      <p>Therefore, considering the advantages and disadvantages of different methods, we can conclude that
the best choice is RNN, namely LSTM for future data processing.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Materials and Methods</title>
      <p>The main characteristics for the analysis of epidemics are geographical identification, as well as
characteristics that are new values of the following criteria:
 observations,
 confirmation,
 death,
 recovery.</p>
      <p>The feature set is also linked to a specific location, such as the city or region that will be predicted.</p>
      <p>
        A data set was used as input data [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This is a time series with different metrics that can be used
for analysis or prediction. The dataset contains information about the COVID-19 virus in relation to the
cities of Ukraine. In this example, a data set for the Lviv region was used.
      </p>
      <p>The initial data of the developing system should provide the user with an understanding of the
situation regarding the epidemiological main indicators: new diseases, recovery, death, etc. The data is
calculated as a prediction based on the original data and create a forecast for this set of characteristics.</p>
      <p>As output, the user will receive an apology visualization in the form of a graph for a specific data
set, as well as a map with a prediction for a specific region for easier visual perception
The proposed in the paper model is built taking into account three main criteria:
 The number of past days for prediction;
 The number of days to anticipate;
 The number of criteria to consider.</p>
      <p>For our case, one day was taken into account to predict the next day using four characteristics by
analyzing the previous seven days.</p>
      <p>In the beginning, neural network decides what information to remember and what to throw out
of the cell state. This action is performed in the "Forget gate" part. X presents input data, H - the result
of the current stage, t is the step number. In this part, the sigmoid function considers the input data from
ht-1 and xt, and then outputs a number between 0 and 1 for each number from cell Ct-1, where 1 will
mean completely save the state, and 0 - completely forget it.</p>
      <p>ft=sig(W(f)[ht-1,xt]+b(f)).</p>
      <p>The next step is the “Input gate” section to update the cell status. First, the current state Xt and
the previously hidden state ht-1 are passed to a sigmoid function to transform values between 0
(important) and 1 (unimportant). Next, the same information about the hidden and current state will be
transmitted via the tanh function. For network regularization, operator tanh calculates vector Čt in range
from -1 to 1 for the multiplying.</p>
      <p>it=sig(W(i)[ht-1,xt]+b(i)).</p>
      <p>Čt=gt =tanh(W(g)[ht-1,xt]+b(g)).</p>
      <p>When the network has prepared information about the data it receives from the two previous
layers, the next step is to decide to save information from the new state in the cellular state in the "Cell
state". The previous state of cell Ct-1 is multiplied by the forgetting vector ft.</p>
      <p>Ct=ftCt-1+Čtht-1.</p>
      <p>The last step is to determine the values to pass to the next layer. Initially, the values of the
current state and the previous hidden state are passed to the last sigma function. This result is further
multiplied by the new cell state generated from the cell state after transmission through the tanh
function. Based on the final value, the network decides what information the hidden state should carry.
This latent state is used for forecasting. As a result, the new cell state and the new latent state are carried
over to the next time step.</p>
      <p>ot=sig(W(o)xt+U(o)ht-1+b(o)),</p>
      <p>ht=ot+tanh(Ct).</p>
      <p>The structure of the model is given in Fig. 2.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <p>
        Before starting work, all data should be standardized [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], as data is measured at different scales and
a large difference can slow down or even hinder the effective learning process:
      </p>
      <p>At the preprocessing stage the data gaps are found. To remove them, grouping is used. The
distribution before and after grouping is given in Fig.3. Data gaps have narrowed, so grouping data for
specific periods is appropriate. It allows to reduce the influence of data heterogeneity due to the
presence of several locations</p>
      <p>Next, a model will be created to predict the regions, so the next step is to group the data with the
selection of a specific area.
b)
Figure 3. Number of new cases for different periods of time before (a) and after grouping (b)</p>
      <p>The next step is to train the model. At this stage, the training of the previously created model was
carried out with the condition of release, if the accuracy is greater than 92% or all epochs will not be
passed. Each model with the lowest loss result is also stored for validation data. The model training
process can be seen in Figure 4, and the model training history in Figure 5.</p>
      <p>The model accuracy on the training data for each of the parameters is given in Fig. 5 and resuls of
forecasting is demonstrated in Fig. 6. New cases and new obseravations are modelled separetly.
b)
Figure 6. Forecasting on training data: а) – new observations; b) – new cases.</p>
      <p>The graph shows the training data, represented by a blue line, as well as the prediction, represented
by an orange line. From these graphs we can say that the accuracy of prediction relative to the test data
is very high.</p>
      <p>Training process is given in Fig. 7.</p>
      <p>Mean absolute percentage error (MAPE) was used to verify the accuracy of the losses.</p>
      <p>The results are shown in Table 2.</p>
      <p>Table 2
Error for testing dataset</p>
      <p>Measure name</p>
      <p>Number of observations
Number of confirmed cases</p>
      <p>Number of death
Number of recoveries</p>
      <p>Because the result is influenced by a large number of factors, and the training took place only on a
short history, the results are of high quality because the MAPE error does not exceed 30% with a
prediction for all characteristics.</p>
      <p>To work with new data, a ready-made data model and a ready-made MinMaxScaler for further
correct alignment of variables relative to previous data is required. After completing the data
normalization phase, the model is trained and then the function plt_result () is called to graphically
display the accuracy of learning for each parameter, which are in separate columns.</p>
      <p>To display graphic data on the map, you need to perform some pre-processing of data. To do this,
use the function update_df_to_plt (), where an important parameter is "registration_area". This
parameter is responsible for the area that will be displayed later.</p>
      <p>Then the show_map () function is executed to display the data on the map. The result is shown in
Fig. 8. The virus power is marked in different colors.</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>Different approaches to time data analysis were analyzed and the use of an RNN neural network,
namely LSTM, was used. This choice was also justified by comparison with other approaches.
Currently, this option is quite promising for predicting new cases of COVID-19, because it is not limited
to the most rigorous type of problem.</p>
      <p>During the development, a software product was built with a description of the system itself, taking
into account the optimal software. A user guide for working with this product in different situations has
also been described.</p>
      <p>So, as a conclusion, this system performs the task well enough given the number of factors that affect
in one way or another the result and this system is quite relevant for use.</p>
    </sec>
    <sec id="sec-7">
      <title>5. Acknowledgements</title>
      <p>This work is supported by National Foundation of Fundamental research, Ukraine, Project
#103.01.0025.</p>
    </sec>
    <sec id="sec-8">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I. M.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. E.</given-names>
            <surname>Hughes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Leach</surname>
          </string-name>
          , “
          <article-title>Real-time epidemic forecasting for pandemic influenza</article-title>
          ,” Epidemiol. Infect., vol.
          <volume>135</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>372</fpage>
          -
          <lpage>385</lpage>
          , Apr.
          <year>2007</year>
          , doi: 10.1017/S0950268806007084.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tat</surname>
          </string-name>
          Dat et al.,
          <source>“Epidemic Dynamics via Wavelet Theory and Machine Learning</source>
          with Applications to Covid-
          <volume>19</volume>
          ,” Biology, vol.
          <volume>9</volume>
          , no.
          <issue>12</issue>
          , p.
          <fpage>477</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2020</year>
          , doi: 10.3390/biology9120477.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] “Prediction of epidemic trends in COVID-19 with logistic model and machine learning technics,” Chaos Solitons Fractals</article-title>
          , vol.
          <volume>139</volume>
          , p.
          <fpage>110058</fpage>
          ,
          <string-name>
            <surname>Oct</surname>
          </string-name>
          .
          <year>2020</year>
          , doi: 10.1016/j.chaos.
          <year>2020</year>
          .
          <volume>110058</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Gu</surname>
          </string-name>
          , “
          <article-title>Epidemic Model Guided Machine Learning for COVID-19 Forecasts in the United States</article-title>
          ,” Epidemiology, preprint, May
          <year>2020</year>
          . doi:
          <volume>10</volume>
          .1101/
          <year>2020</year>
          .05.24.20111989.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gomes</surname>
          </string-name>
          , R. Castro, “
          <article-title>Wind speed and wind power forecasting using statistical models: autoregressive moving average (ARMA) and artificial neural networks</article-title>
          (ANN)”,
          <source>International Journal of Sustainable Energy Development</source>
          , vol.
          <volume>1</volume>
          , no.
          <issue>1</issue>
          /2, pp
          <fpage>13</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Haider</surname>
          </string-name>
          , Abbas, “
          <article-title>The COVID-19 Impact on Oil Market and Equity Market Link: An Evidence from ARMA-GJR GARCH-M Model”</article-title>
          , Diss. CAPITAL UNIVERSITY,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Benvenuto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Giovanetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Vassallo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Angeletti</surname>
          </string-name>
          , M. Ciccozzi, “
          <article-title>Application of the ARIMA model on the COVID-2019 epidemic dataset”, Data in brief</article-title>
          , vol.
          <volume>29</volume>
          , p.
          <fpage>105340</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Milani</surname>
          </string-name>
          , “
          <article-title>COVID-19 outbreak, social response, and early economic effects: a global VAR analysis of cross-country interdependencies”</article-title>
          ,
          <source>Journal of population economics</source>
          , vol.
          <volume>34</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>223</fpage>
          -
          <lpage>252</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Howell</surname>
          </string-name>
          , “
          <article-title>Battling Burnout at the Frontlines of Health Care Amid COVID-19”, AACN Advanced Critical Care</article-title>
          , vol.
          <volume>32</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>203</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Acharya</surname>
            ,
            <given-names>U. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oh</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagiwara</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>J. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gertych</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp; San Tan, R. “
          <article-title>A deep convolutional neural network model to classify heartbeats”, Computers in biology</article-title>
          and medicine,
          <volume>89</volume>
          ,
          <fpage>389</fpage>
          -
          <lpage>396</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Zaremba</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          “
          <article-title>Recurrent neural network regularization”</article-title>
          ,
          <source>arXiv preprint arXiv:1409.2329</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Donkers</surname>
            , Tim,
            <given-names>Benedikt</given-names>
          </string-name>
          <string-name>
            <surname>Loepp</surname>
            , and
            <given-names>Jürgen</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
          </string-name>
          .
          <article-title>"Sequential user-based recurrent neural network recommendations."</article-title>
          <source>Proceedings of the eleventh ACM conference on recommender systems</source>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Piven</surname>
          </string-name>
          , VasiaPiven/covid19_ua.
          <year>2021</year>
          . Accessed: Apr.
          <volume>26</volume>
          ,
          <year>2021</year>
          . URL:: https://github.com/VasiaPiven/covid19_ua
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Al</surname>
            <given-names>Shorman</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amaal</surname>
            <given-names>R.</given-names>
          </string-name>
          , et al.
          <article-title>"The Influence of Input Data Standardization Methods on the Prediction Accuracy of Genetic Programming Generated Classifiers."</article-title>
          <source>IJCCI</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>