<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Application of Data Assimilation Algorithms Based on Kalman Ensemble Filters for the Lorenz Attractor</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yulia S. Timoshenkova</string-name>
          <email>julia.timoshenkova@urfu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikolay T. Sa ullin</string-name>
          <email>n.t.safiullin@urfu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergey V. Porshnev</string-name>
          <email>s.v.porshnev@urfu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Yeltsin Ural Federal University</institution>
          ,
          <addr-line>Yekaterinburg, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <fpage>82</fpage>
      <lpage>87</lpage>
      <abstract>
        <p>The paper describes application of the Data Assimilation methods based on the Kalman ensemble ltering to forecast the Lorentz system. The quality of the forecast and the in uence of the method for the parameter choise on the result of forecasting the system was assessed. The paper shows the advantages and disadvantages of the chosen methods of Data Assimilation. The recommendations for increasing accuracy of those methods are presented.</p>
      </abstract>
      <kwd-group>
        <kwd>EnKF</kwd>
        <kwd>LEnKF</kwd>
        <kwd>Data Assimilation</kwd>
        <kwd>Lorenz attractor</kwd>
        <kwd>Kalman lter</kwd>
        <kwd>Forecast</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Correction of the forecast based on mathematical models with using new
observations is an actual problem. This approach avoids growth of the forecast
con dence interval as its duration increases.</p>
      <p>One of the instruments for solving this task is the Data Assimilation (DA)
technique.</p>
      <p>There are areas of science, in which the forecast models are built not on a
one-dimensional time series generated by the multidimensional one, but on a
multidimensional underlying process itself.</p>
      <p>The use of methods of assimilation and correction of data ensures correct
prediction and modeling of the complex systems.</p>
      <p>Since the DA methods are based on the mathematical model of the system,
the choice of di erent methods of data assimilation can signi cantly a ect the
nal results of the forecast correction.</p>
      <p>
        There are no recommendations for the choice of a particular family of DA
methods because of their novelty [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Therefore, the problem of comparative
analysis of data assimilation methods is actual. In order to compare those methods
and to test their accuracy one needs a standard suitable system to forecast.
      </p>
      <p>The Lorentz system was designed to construct a simpli ed model of
atmospheric convection for the long-term weather forecasting. Nowadays, the Lorenz
attractor is a standard model for testing the DA methods; thus, it is used as a
model to forecast in this work.</p>
      <p>
        The mathematical model of the Lorenz attractor is described by a system
consisting of three nonlinear ordinary di erential equations that represent a nite
amplitude of convection [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
dx
dt
dx
dt
dx
dt
= (y
      </p>
      <p>x);
= x(r
= xy
y)
bz;
y;
where = =k is the Prandtl number, r = Ra=Rac is the normalized Rayleigh
number (normalized), b = 4=(1 + a2) is the geometric factor x s the convection
intensity, y is the temperature di erence between ascending and descending wind
ows, z is the deviation of the vertical temperature pro le from the linear one.</p>
      <p>
        The peculiarity of this system is that when the parameters are = 10, b =
8/3, and r 24.06, the behavior of the system becomes random. In the phase
space, the attractor has the topology of the tangle of trajectories, in which two
regions can be distinguished [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. At any time point, the solution is in one of these
regions, and the transition of the system state to another region is unpredictable.
      </p>
      <p>The purpose of the work is to carry out a comparative analysis of applying
two DA methods, such as the Ensemble Kalman Filter and Local Ensemble
Transform Kalman Filter to the Lorenz attractor.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Mterials and Methods</title>
      <p>In the Data Assimilation method, the original model is formulated in accordance
with the following system of equations
(1)
(2)
(3)
xk+1 = M (xk) + wk;</p>
      <p>yk = H(xk) + vk;
where M is an operator or transition function that determines the evolution of
the system in time, xk is the known model state at the time point tk, xk+1 is the
model state at the next time point tk+1, yk is the observation vector, H is the
observation operator that describes the multi-dimensional state of the system
based on the observation vector, wk is the model error, vk is the observation
error.</p>
      <p>
        In the DA method, the analysis step is based on the system current state xfk .
The corrected system forecasting state xka for the next time point is based on
xfk . For calculation of xka, only observation vector for the current time point is
used [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>xka = xfk
where K is the transmission coe cient, yk is the observation, e = yk H(xfk ) is
the error between the known observations and values obtained using the model.</p>
      <p>
        Various DA methods minimize error e using di erent techniques described
in this paper. The transmission coe cient is also called the Kalman coe cient.
There is a family of DA methods based on the Kalman lter (KF). One of those
methods is the Ensemble Kalman Filter (EnKF) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The EnKF is a variant of the generalized Kalman lter, in which the
covariance of the prediction errors is estimated using the ensemble of forecasts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The
Kalman ensemble ltration methods are widely used to assimilate observations
in dynamic models.
      </p>
      <p>The EnKF algorithm is implemented by the following sequence of steps.
1. Forming an ensemble of the initial data at the initial time point with the
help of a mathematical model.</p>
      <p>2. Calculation of N predictions for ensemble of the initial data in order to
obtain the values of the observed parameters at the time point for the state
vector obtained on the previous step. Next, the method of successive corrections
for the model is used.</p>
      <p>3. Calculation of the covariance error matrices for the forecast correction.
The step consists of the following items: formation of a covariance error matrix
of the forecast vector; calculation of the Kalman weight operator K; update
of the ensemble model values in accordance with the weight operator and the
observation vector; update of the covariance matrix of the forecast errors in
accordance with the speci ed ensemble; calculation of value of the analyzed
components and the covariance matrix of the analysis errors.</p>
      <p>4. Steps 2 and 3 are executed sequentially for each time point with each new
observation obtained.</p>
      <p>
        The Local Ensemble Filters use localization of the observational error
covariance, which takes into account only those observations that are located in a
certain region around the desired point. The main idea of the Local Ensemble
Transform Kalman Filter (LETKF or LEnKF) is to perform calculations not in
the physical space of the model, but in the space of the ensemble. Typically, the
ensemble space is less than physical one of the model [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>The LEnKF algorithm is implemented by the following sequence of steps.
1. Formation of an ensemble of model data of size N at the initial time point
with the help of a mathematical model.</p>
      <p>2. Calculation of the ensemble forecast of the system state at the next instant
with the help of the model equation and the model state vector at the previous
time point.</p>
      <p>3. Creating a local plane in the state space, and constructing the projection
of each point of the ensemble of the system state onto this plane.</p>
      <p>4. For each state vector in the ensemble obtained on step 2, calculate the
distance from the ensemble average to the current state vector and project this
distance into the low-dimensional state subspace that represents the ensemble
in that region in the best way.
5. Perform the assimilation of data in each of the local low-dimensional
subspaces, obtaining the average analysis value and covariance of the state error in
each local region according to the EnKF.</p>
      <p>6. Obtaining the necessary local analysis of the state ensembles using the
mean value of the local analysis and covariance.</p>
      <p>7. Formation of a new global analysis ensemble using the local analyses
obtained on step 6.</p>
      <p>8. Repeat steps 1 6 until the observations are complete.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>During the study, the standard deviations of the Lorenz attractor were taken
equal to 0.5, 1, and 5, which corresponds to 4:5% 8:8% and 44% of the coordinate
amplitude of the Lorentz system. These values may seem small, but the standard
deviations greater than 50% signi cantly a ect the system and generate very
inaccurate observations for all three components of the system. The deviation
values of the observed states are taken to be from 0.01 to 20, which correspond
to 0:08% to 176% of the Lorentz system coordinate amplitude, or approximately
from 60 dB to {5 dB in the ratio of the real coordinate to the error of its
observation.</p>
      <p>Evaluation of the noise in uence on the result of the algorithms operation
was carried out for several simulation parameters. To analyse the quality of the
algorithm, the mean square error (MSE) parameter is calculated. Results of the
MSE parameter estimations are shown in Fig. 1.</p>
      <p>Fig. 1 shows that the choice of the DA method mainly in uences on the
MSE prognosis. The use of the EnKF gives the best result. For the experiment,
in which the highest values of the observation noise dispersion are taken, the
worst results of the simulation are obtained. This con rms the theory that the
observation error greatly in uences the obtained forecast.</p>
      <p>It should be noted that the results of the experiment are also in uenced by
the fact that all components of the attractor are related to each other; that is,
presence of errors in one of the coordinates a ects the result of the forecast of
each other component.</p>
      <p>Fig. 2 shows the MSE dependence on the number of ensembles for the EnKF
and LEnKF lters. As the number of lter ensembles grows, the resulting value
of the MSE decreases.</p>
      <p>Under the minimum error in the covariance of observations, the e ect of the
observation noise variance on the MSE is minimal, as seen in Fig. 1. However,
with an increase of the observation error variance and model error, the MSE value
becomes approximately equal for all DA methods. This dependence suggests that
with increasing the noise variance to achieve an acceptable MSE, it is necessary
to increase the number of ensembles in the KF.</p>
      <p>Based on the obtained results, it can be concluded that when the MSE of the
observation error is less than 11, the EnKF should be used to obtain a smaller
MSE. When the standard deviation of the observation error is greater than 11,
or more than 100% of the system coordinate amplitude, the best smaller MSE
are given by the ensemble lters. In the case of using the ensemble lters, the
value of the MSE ceases to change signi cantly with the number of ensembles
larger than 100. Thus, when using the EnKF, the optimum number of ensembles
is 100 for predicting the behavior of the Lorenz attractor.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>During the research of the capabilities of under applying the ensemble Kalman
lter methods to perform prediction of the coordinates of the Lorenz attractor,
the following practical results were obtained.</p>
      <p>The increase in variance of the observation error (greater than 70%) signi
cantly a ects the MSE prediction. Accuracy of the prediction falls also with the
increase in the observation error, as the obtained results shown.</p>
      <p>The result of forecast is also signi cantly in uenced by the choice of the lter
type. To obtain the minimum MSE, it is recommended to use the LEnKF. A
small MSE (lower than 10) of the EnKF can be explained by the fact that the
system is modeled several times that averages the error of the modeled chaotic
system.</p>
      <p>A signi cant increase in MSE (up to 2{3 times) occurs when the value of
the standard deviation of system errors increases by larger than 50%. As a
consequence, for high deviation values (higher than 100% of the basic amplitude),
the largest possible number of ensembles in the EnKF and LEnKF should be
used. But after the value of 100, the value of the MSE ceases to change signi
cantly. Thus it is recommended to use the value of 100 ensembles in the EnKF
and LEnKF methods, as a balance between accuracy and speed of the Data
Assimilation.</p>
      <p>Compared to the simple Kalman Filtering, for higher observational errors, the
methods EnKF and LEnKF provide a better accuracy (up to 3 times for small
deviations). For observational errors below 40%, the LEnKF performs better in
accuracy than the EnKF. But with higher observational errors their accuracy is
on the same level. Thus, the choice between those two methods must be made
based on their performance and memory consumption.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Leeuwen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vetra-Carvalho</surname>
          </string-name>
          ,
          <article-title>SANGOMA: Stochastic Assimilation for the Next Generation Ocean Model Applications</article-title>
          , Leeuwen Vetra-Carvalho,
          <article-title>(</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>W. Van den Bossche</surname>
          </string-name>
          ,
          <article-title>Data assimilation toolbox for Matlab</article-title>
          ., (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Evensen</surname>
          </string-name>
          ,
          <article-title>Data assimilation: the ensemble Kalman lter</article-title>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Timoshenkova</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Sa ullin, On the possibility of the forecast correction for inaccurate observations based on data assimilation</article-title>
          ,,
          <source>Dynamics of Systems, Mechanisms and Machines (Dynamics)</source>
          ,
          <volume>1</volume>
          {
          <fpage>4</fpage>
          , (
          <year>2017</year>
          ). DOI:
          <volume>10</volume>
          .1109/Dynamics.
          <year>2017</year>
          .8239519
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. M.
          <article-title>Bocquet and</article-title>
          . des
          <string-name>
            <surname>P. P. CEREA</surname>
          </string-name>
          ,
          <article-title>Introduction to the principles and methods of data assimilation in geosciences</article-title>
          , Notes Cours c.
          <source>Ponts ParisTech</source>
          , (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Houtekamer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. L.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <article-title>Data assimilation using an ensemble Kalman lter technique</article-title>
          ,
          <source>Mon. Weather Rev.</source>
          , Vol.
          <volume>126</volume>
          , No.
          <volume>3</volume>
          ,
          <issue>796</issue>
          {
          <fpage>811</fpage>
          , (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>