<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Aggregation Algorithm vs. Average For Time Series Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Waqas Jamil</string-name>
          <email>jamylwaqas@gmail.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuri Kaliniskan</string-name>
          <email>Yuri.Kalnishkan@rhul.ac.uk</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hamid Bouchachia</string-name>
          <email>abouchachia@bournemouth.ac.uk</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Learning with expert advice as a scheme of on-line learning has been very successfully applied to various learning problems due to its strong theoretical basis. In this paper, for the purpose of times series prediction, we investigate the application of Aggregation Algorithm, which a generalisation of the famous weighted majority algorithm. The results of the experiments done, show that the Aggregation Algorithm performs very well in comparison to average.</p>
      </abstract>
      <kwd-group>
        <kwd>Aggregation Algorithm</kwd>
        <kwd>time-series</kwd>
        <kwd>auto-regressive-movingaverage</kwd>
        <kwd>auto-regressive-integrated-moving-average</kwd>
        <kwd>Fourier transform</kwd>
        <kwd>on-line learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A time series is a set of repeated observations of the same variable, such as stock
return or GNP. A time series consists of cyclic(for example daily uctuations),
seasonal(variation in data due to calender related e ect) or irregular e ects(any
movement other than seasonal or cyclic). Machine Learning methods are now
been used to analyse large data [
        <xref ref-type="bibr" rid="ref3">Ahmed et al., 2010</xref>
        ]. Traditionally time series
auto-regressive moving average (ARMA) and auto-regressive-integrated moving
average(ARIMA) models were designed to work in batch mode, rather than the
on-line mode, however work on on-line ARMA models [Anava et al., 2013] and
ARIMA models [Liu et al., 2016] has been done.
      </p>
      <p>There exist two classes of modelling techniques for time series: statistical
learning and competitive on-line learning [Anava et al., 2013], the former assumes
that the observations are drawn from some unknown distribution. Representative
techniques of this class include the well-known autoregressive moving average
(ARMA) and its alike seen as a standard time-series modelling techniques. The
motivation behind developing competative on-line learning algorithm is that the
statistical(ARMA) models have strong distributional assumptions, due to which
they have asymptotic guarantees [Kuznetsov and Mohri, 2016].</p>
      <p>On-line learning has received a great attention from the machine learning
community. Its origin goes back to the late 1980's and early 1990's with the
advent of the paradigm of prediction with expert advice. The early work appeared
in a number of seminal papers by [Haussler et al., 1994]. On-line learning consists
of learning a sequentially presented set of training data upon arrival, without
re-examining data that has been processed so far. In general on-line learning is
practical for applications where the data set is large and cannot be processed
at once due to memory constraints. Practically an on-line learner receives a
new data instance, along with current hypothesis, checks if the data instance is
covered by the current hypothesis and updates the hypothesis accordingly. The
protocol of on-line learning can be summarized as follows: the learner receives an
observation; the learner makes a decision; the learner receives the ground truth;
learner incurs the loss and updates its hypothesis. The learning process is based
on the minimisation of the loss (regret) which corresponds to the discrepancy
between the loss and the loss of the best expert in hindsight.</p>
      <p>
        Neural network is another approach used in time-series, in particular nancial
and economic time-series. Neural networks are universal function approximators
that can map any non-linear function [
        <xref ref-type="bibr" rid="ref14">White, 1989</xref>
        ], which makes them extremely
attractive when dealing with non-linearity, as linearity in ARMA imposes limits
on their exibility. A hybrid of ARIMA and neural network models has also been
used [D az-Robles et al., 2008] in the past. The basic model is; the target variables
are composed of a linear and non-linear component; it estimates linear portion
using ARIMA; error term consists of non-linear relationship with previous errors,
for which the neural networks are used.
      </p>
      <p>Prediction models have parameters and we are faced with the problem of
selecting the best set of parameters. If we have little information on the predictive
behaviour of parameters, one may want to keep all models and predict using the
average.</p>
      <p>Methods of competitive prediction provide a better alternative to averaging.</p>
      <p>
        The approach of this paper is drawn from the area of competitive on-line
prediction, where the goal is merging predictions of experts. In this paper we
merge predictions of time series models with respect to the square loss.
[DeSantis et al., 1988] for the rst time presented the Bayesian mixing scheme using
log-loss, later [Littlestone and Warmuth, 1989] presented what was known as
weighted majority algorithm and Vovk generalised them which resulted in
Aggregation Algorithm(AA) [
        <xref ref-type="bibr" rid="ref8">Vovk, 1992</xref>
        ] and [
        <xref ref-type="bibr" rid="ref11">Vovk, 1990</xref>
        ]. AA has been proven
in [
        <xref ref-type="bibr" rid="ref12">Vovk, 1995</xref>
        ] to be optimal in some cases.
      </p>
      <p>[Box et al., 2015] were the ones who probably launched auto-regressive
integrated moving-average, the Box-Jenkins methodology eloberated by [Hibon
and Makridakis, 1997]. The idea was extended to state-space representation,
by [Durbin, 2004]. The book by [Hyndman and Athanasopoulos, 2014] captures
broad spectrum of work on time series.</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>We develop an underlying system of obtaining predictions with the objective of
comparing simple average against the AA.</p>
      <p>The system we provide is a hybrid system i.e. the prediction comes from
usual statistical learning approach which we brie y outline in this section along
with the on-line protocol, then we use a competitive on-line learning algorithm
on those predictions, which are outlined in the next section.</p>
      <p>
        We highlight the usefulness of AA in time-series in a similar fashion as done
in [
        <xref ref-type="bibr" rid="ref6">Romanenko, 2015</xref>
        ], the novelty of this paper is in mixing of ARIMA's. We
hope the idea is extended to the on-line time-series set-up, since on-line
timeseries also require parameters selection.
2.1
      </p>
      <sec id="sec-2-1">
        <title>The underlying system</title>
        <p>The observed time series is a realization of a stochastic processes. A stochastic
process is any collection of random variables Xt; t 2 T de ned on a common
probability space , here t denotes the time. So T can be either discrete or
continuous set of series. In our case we only deal with discrete-time series [Berchtold,
1995].</p>
        <p>If a random variable X is indexed to time, we denote time by t, the
observations fXt; t 2 Tg, where T is a time indexed set, for example it may be a
set of integers. The stochastic process is described by a probability distribution
for fXtg, where often elements lack independence. The distribution is usually
characterized using the moments.</p>
        <p>The objective of time-series models is to make predictions, so sometimes it is
also referred as predictive inference. Time-series methods make prediction based
on historical pattern of the data, measurements are taken at successive periods,
such as over day, month, year etc. At the very heart of time series analysis lies
ARMA models, mathematically represented as:</p>
        <p>p
Xt = X
i=1</p>
        <p>q
X
i=1
iXt i +
iWt i + Wt
(1)
We can divide ARMA models into two parts, auto-regressive(AR) part and the
moving-average(MA) part. The challenge is often to determine the correct
order p and q for AR and MA part respectively. Generally the notation used is
ARMA(p,q)1. So AR(p) can be represented by the following equation:
p
Xt = X
i=1
iXt i + Wt
(2)
1 The equations used for time-series ARMA,ARIMA,AR, and MA are adopted from
[Liu et al., 2016]
Similarly for MA(q) the following equation:</p>
        <p>
          q
Xt = X iWt i + Wt
k sin
2 kt
m
+ k cos
where Xt is stationary, 2 Rp, 2 Rq parameters, and Wt is a Gaussian white
noise series with mean 0 and variance 2. The AR models is very similar to the
multiple linear regression models, except that the Xt is regressed on the past
values of Xt [Liu et al., 2016], whereas MA has been used for data
smoothing, without explicitly using the term moving average, they were described as
instantaneous averages by [
          <xref ref-type="bibr" rid="ref15">Yule, 1909</xref>
          ]. Exponential moving average were
formally applied in [Haurlan, 1968] to track stock prices. Exponential smoothing
developed by [Brown, 2004] and [Holt, 2004] is a techniques regularly used now
a days for data smoothing.
        </p>
        <p>In practical situation often a time-series is not a realisation of a stationary
process [Liu et al., 2016], so ARIMA(p; d; q) are used:</p>
        <p>p q
OdXt = X iOdXt i + X iWt i + Wt
i=1
i=1
where OXt = Xt Xt 1, 2 Rp, 2 Rq, and Wt v N (0; 2). It is worth noting
that ARMA(p; q) is a special case of ARIMA(p; 0; q).</p>
        <p>
          Our main experiments uses data from Meteorology, so it is reasonable to
assume that there is some sort of cyclic behaviour or periodicity in it, which
is also justi ed in [Jones and Brelsford, 1967]. In order to tackle periodicity
we make use of the Fourier series. A Fourier series is a speci c type of in nite
mathematical series involving trigonometric functions, the series were introduced
by [baron Fourier, 1831]. Fourier series are used in applied mathematics, and
especially in physics and electronics, to express periodic functions such as those
that comprise communications signals in waveform, however Fourier series truly
began with the profound work of Fourier on heat conduction at the beginning of
the 19th century [
          <xref ref-type="bibr" rid="ref13">Walker, 1988</xref>
          ], Fourier proposed that initial temperatures could
be represented as a series of sin functions. The discrete-time Fourier transform
is a periodic function, often de ned in terms of a Fourier series. [
          <xref ref-type="bibr" rid="ref7">Strang, 1994</xref>
          ]
states \The discrete fourier transform is the most important discrete transform,
used to perform Fourier analysis in many practical applications ". The discrete
Fourier transform takes a time-based pattern, measures every possible cycle,
and returns the overall amplitude, o set, and rotation speed for every cycle that
was found, [
          <xref ref-type="bibr" rid="ref5">Press, 2007</xref>
          ] provides a comprehensive guide on its application.The
data is cyclic or periodic, thus we prefer using Fourier series [Bracewell, 1965]
approach with ARIMA, instead of SARIMA models [Hu et al., 2007], and our
model becomes:
(3)
(4)
(5)
Where Xt is the ARIMA or ARMA model 2. The value of k can be chosen to
adjust the data 3, and c is some constant.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>On-line protocol</title>
        <p>
          We now de ne the notion of game-theoretic frame work, or to be more speci c
we de ne [
          <xref ref-type="bibr" rid="ref10">Vovk and Zhdanov, 2009</xref>
          ] general prediction game. Let be the
prediction space and be the outcome space, on each trial t 2 Z. Also we
assume the number of experts n &lt; 1, then each expert makes a prediction
t 2 ; a learner observes h 1t::: nti; learner makes a t; nature chooses an
outcome !t 2 ; nally each expert and learner loss is calculated using an
appropriate loss function.
        </p>
        <p>The following notation is used
{ outcomes are denoted as !1; !2, they happen sequentially from the outcome
space = [A; B], where A and B are the minimum and maximum(of a
particular data for example).
{ predictions at time t is represented as t, and they belong to prediction space
= fA; Bg.
{ 1; :::; N experts, there are nite number of experts and the experts space
is denoted by .
{ we denote learner by L and its loss at time T as LossT (L) = PT
t=1 ( t; !t),
where t and !t and the expert loss is LossT ( ).</p>
        <p>Following are the set of assumptions or rather a scenario under we work:
{ We do not assume that there is a model generating outcomes
{ Outcomes may be adversarial meaning, for example output of 1 every-time
we predict zero
{ We have access to the experts predictions and expert prediction is
incorporated or consulted before a prediction is given
{ The objective to be as close as possible to the best expert
{ Experts maybe adversarial too
2 An ARIMA is just the di erencing of ARMA model, if ARMA models are not
stationary, but are di erence stationary then we call them ARIMA, where I is the
number of di erence we took to make it stationary [Ghosh, 1976], but forcast
package by Rob J Hyndman provides more exibility.
3 We used (5) in our experiment from the implementation in the package forecast
by Rob J Hyndman, for details [Hyndman and Athanasopoulos, 2014]</p>
        <p>The protocol for on-line prediction described in introduction and here can
be formalised as follows:4
Protocol Prediction With Expert Advise
for t = 1; 2; ::: do</p>
        <p>2 predicts t 2
Learner output t 2
System output !t 2</p>
        <p>2 su ers losses ( t ; !t)</p>
        <p>Learner loss ( t; !t)
end for</p>
        <p>
          There are di erent loss functions that can be used, but for the purposes of
our experiment we use square-loss, the main reasons are as follows:
{ If the outcomes and predictions are bounded, square loss is bounded;
{ bounded square loss is mixable [
          <xref ref-type="bibr" rid="ref9">Vovk, 2001</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Applying AA on Time Series</title>
      <p>The weighted majority algorithm is very restrictive, so we go on a better
algorithm, known as the Aggregating Algorithm(AA). The AA takes two
parameters, a prior probability say W0 and the learning rate &gt; 0. W0 is related to
the weights and the 0 in the subscript makes it the initial weights. Before we
describe the AA, we explain the mechanics of the algorithm i.e. obtaining the
generalised prediction, later that can be easily used for the purposes of the actual
prediction. At every step t experts weight is updated, so intuitively, if an expert
makes a mistake we would reduce its weight, mathematically we represent this
by using our de ned notation:
gt(!) = log</p>
      <p>
        Z
( t;! )Wt 1d
(6)
= WWtt 11((d )) in simple words W
where 2 and W represent the normalised
weights. For mathematical details and optimality of AA see [
        <xref ref-type="bibr" rid="ref9">Vovk, 2001</xref>
        ].
      </p>
      <p>Notice in (6) we use ! not !t, this is because we are at the moment
making prediction. What AA does is uses a substitution function5 which maps the
generalised prediction into .</p>
      <p>
        In certain situations it may not be possible to use a substitution function
that perfectly maps to , we call such situations non-mixable scenario, however
4 The notation are adopted from [
        <xref ref-type="bibr" rid="ref9">Vovk, 2001</xref>
        ] and [
        <xref ref-type="bibr" rid="ref10">Vovk and Zhdanov, 2009</xref>
        ]
5 The substitution function used in this experiment was B+A + g(2B( B) gA(A)) , by
consider2
ing square loss with = = [A; B] and (B 2A)2 , by using [Haussler et al., 1998]
the restriction [ 1; 1] can be removed and then we solve the system of equations.
in the case of square loss it is possible to nd a substitution function that maps
to . Mathematically we say that the loss function is -mixable if there is a
super-prediction = f(x; y)j there is 2 : x ( ; 1); y ( ; 1)g meaning
C( ) = 1, and we can always solve the ( ; 1) gt( 1) and ( ; 1) gt(1).
The loss of AA can not be much larger than the best expert, for a mixable nite
experts game, and uniformly initialising the prior weights of the experts.
      </p>
      <p>
        Loss(AA) = Lossbest( ) +
log N
(7)
where 2 , is the learning rate, and N is the number of experts. This
bound (7) is shown [
        <xref ref-type="bibr" rid="ref12">Vovk, 1995</xref>
        ] to be optimal in a very strong sense i.e. it can't
be improved by any other prediction algorithm.
      </p>
      <p>We now formally present the pseudo-code, where parameters are &gt; 0 and
initial distribution is of q1; q2; ::; qn. This algorithm is also for the cases of
nonmixable game.</p>
      <p>Algorithm Aggregating Algorithm
initialise weights w0n = qn; n = 1; 2; :::N
for t = 1; 2; :: do
notice experts prediction tn, n = 1; 2; :::; N
weights normalisation ptn 1 = PiN=w1tnw1ti 1
solve the system (! 2 ) ( ; !)
output a solution t
notice !t
experts weights update wtn = wtn 1e
end for</p>
      <p>C( ) log PN</p>
      <p>n=1 ptn 1e
( tn;!t); n = 1; 2; :::; N
( tn;!) w.r.t
and
4</p>
    </sec>
    <sec id="sec-4">
      <title>Empirical evaluation</title>
      <p>
        Our experiments uses [
        <xref ref-type="bibr" rid="ref1 ref2">Dai, 2012</xref>
        a] and [
        <xref ref-type="bibr" rid="ref1 ref2">Dai, 2012</xref>
        b] data, we address them as
maximum temperature and minimum temperature data respectively, they have
3650 days of maximum and minimum temperatures in Degrees Celsius
respectively, from year 1981 to 1990. The data can be downloaded and visualised from
the links provided in maximum temperature and minimum temperature data .
The reason behind using this data is its periodicity (the time series model used
in our experiments don't update the coe cients or the value of k or the coe
cients of the ARIMA models) and its size. We performed experiments by tting
18 combinations of ARIMA models (experts for AA) with parameters p = 0; 1; 2,
d = 0; 1, q = 0; 1; 2 see [
        <xref ref-type="bibr" rid="ref4">Pankratz, 2012</xref>
        ], and their respective coe cients on the
rst 365 days, then these models were used to obtain one-step ahead forecast
for each model. For the Fourier part k = 57 for minimum temperature data and
k = 95 for maximum temperature, see (6). The value of k was chosen by doing
cross-validation on the rst 365 days.
      </p>
      <p>To be more explicit, rst 365 days were used to select the coe cients for
the 18 models (experts for AA) and the value of k, then at each step these
models give one step ahead prediction (experts predictions for AA) based on the
previous observations.</p>
      <p>Fig 1 demonstrates that AA follows the best expert, which in both ( A and C
for minimum temperature data) experiments is ARIMA(0; 1; 0) with respective
values of k. In our experiments overall loss of AA is even less than the best
expert. This is because the best expert gives the lowest overall loss, however it
does not give lower loss from the start, there are other experts who are able to
compete with this expert and some outperform the best expert, all of this AA
captures, and is able to outperform the best expert.</p>
      <sec id="sec-4-1">
        <title>Expert</title>
        <p>2.0e+09
1.5e+09
s
s
Lo1.0e+09
5.0e+08
0.0e+00
0
1000
2000</p>
      </sec>
      <sec id="sec-4-2">
        <title>Days</title>
        <p>E1
E2
E3
E4
E5
E6
E7
E8
E9
E10
E11
E12
2000
Days
0
1000
3000
0
1000
3000
2.0e+09
1.5e+09
s
so1.0e+09
L
5.0e+08
0.0e+00</p>
      </sec>
      <sec id="sec-4-3">
        <title>Expert</title>
        <p>0
1000
2000</p>
      </sec>
      <sec id="sec-4-4">
        <title>Days</title>
        <p>E1
E2
E3
E4
E5
E6
E7
E8
E9
E10
E11
E12
3000
E13
E14
E15
E16</p>
      </sec>
      <sec id="sec-4-5">
        <title>Expert</title>
        <p>0
1000
E1
E2
E3
E4
2000</p>
      </sec>
      <sec id="sec-4-6">
        <title>Days</title>
        <p>E5
E6
E7
E8
E9
E10
E11
E12
3000
E13
E14
E15
E16
0
1000
3000
2000</p>
      </sec>
      <sec id="sec-4-7">
        <title>Days</title>
        <p>0
1000</p>
        <p>3000
1904:307 respectively is violated by the average.</p>
        <p>F</p>
        <p>0
−50000
s
s
o
L−100000
−150000
E17
E18
H
2.0e+09
1.5e+09
s
so1.0e+09
L
5.0e+08
0.0e+00</p>
        <p>AA vs Outcomes
Daily Minimum Temperature</p>
        <p>Daily Maximum Temperature
ssLo 074e+−
0
4
0
3
0
2
0
1
0
0
+
e
0
7
0
+
e
2
−
7
0
+
e
6
−
7
0
+
e
8
−
7
0
ssLo 4e+−
0 500 1000 1500 2000 2500 3000</p>
        <p>Days
Fig. 4. The decreasing behaviour of the di erence between square losses of AA and
average suggest that AA outperforms the average.</p>
        <p>0 500 1000 1500 2000 2500 3000</p>
        <p>Days
0 500 1000 1500 2000 2500 3000</p>
        <p>Days
One can compare Fig 2 (E and G for minimum temperature data) with Fig 1
and easily see averaging is not favourable, since the spread between the expert
with maximum loss and the expert with minimum loss is colossal. Fig 3 shows
that AA actually does very well in predicting the outcomes, only the extreme
temperatures are missed mainly.</p>
        <p>Fig 4 is the plot demonstrating the di erence between the AA and the average
(AA minus average) square loss, since average is greater than the AA, hence the
decreasing behaviour in the curve.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We have studied the performance of AA, and average on minimum temperature
and maximum temperature data . We conclude that, best expert ARIMA(0; 1; 0)
with k = 95 for maximum temperature data and k = 54 for minimum
temperature data are outperformed by AA prediction, furthermore AA is signi cantly
better than taking simple average as shown in Table 2 (the square loss of AA
is substantially lower then the average) Fig 1 clearly show that the theoretical
bound by AA has not been violated, due to which AA gives a better result than
the simple average. Experts predictions are used by AA to produce a better
overall result.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>Waqas Jamil and Hamid Bouchachia has been supported by the European
Commission under the Horizon 2020 Grant 687691 related to the project: PROTEUS:
Scalable Online Machine Learning for Predictive Analytics and Real-Time
Interactive Visualization.</p>
      <p>Yuri Kalnishkan has been supported by the Leverhulme Trust through the
grant RPG-2013-047 `Online self-tuning learning algorithms for handling
historical information'.
Anava et al., 2013. Anava, O., Hazan, E., Mannor, S., and Shamir, O. (2013). Online
learning for time series prediction. In COLT, pages 172{184.
baron Fourier, 1831. baron Fourier, J. B. J. (1831). Analyse des equations determinees,
volume 1. Firmin Didot.</p>
      <p>Berchtold, 1995. Berchtold, A. (1995). Autoregressive modelling of markov chains. In</p>
      <p>Statistical Modelling, pages 19{26. Springer.</p>
      <p>Box et al., 2015. Box, G. E., Jenkins, G. M., Reinsel, G. C., and Ljung, G. M. (2015).</p>
      <p>Time series analysis: forecasting and control. John Wiley &amp; Sons.</p>
      <p>Bracewell, 1965. Bracewell, R. (1965). The fourier transform and its applications. New</p>
      <p>York, 5.</p>
      <p>Brown, 2004. Brown, R. G. (2004). Smoothing, forecasting and prediction of discrete
time series. Courier Corporation.</p>
      <p>DeSantis et al., 1988. DeSantis, A., Markowsky, G., and Wegman, M. N. (1988).</p>
      <p>Learning probabilistic prediction functions. In Foundations of Computer Science,
1988., 29th Annual Symposium on, pages 110{119. IEEE.</p>
      <p>D az-Robles et al., 2008. D az-Robles, L. A., Ortega, J. C., Fu, J. S., Reed, G. D.,
Chow, J. C., Watson, J. G., and Moncada-Herrera, J. A. (2008). A hybrid arima and
arti cial neural networks model to forecast particulate matter in urban areas: The
case of temuco, chile. Atmospheric Environment, 42(35):8331{8340.</p>
      <p>Durbin, 2004. Durbin, J. (2004). Introduction to state space time series analysis. State</p>
      <p>Space and unobserved component models: Theory and Applications, pages 3{25.
Ghosh, 1976. Ghosh, N. (1976). Time series analysis and forecasting (the box-jenkins
approach). Journal of the Operational Research Society, 27(3):644{644.
Haurlan, 1968. Haurlan, P. (1968). Measuring Trend Values. Trade Levels.
Haussler et al., 1998. Haussler, D., Kivinen, J., and Warmuth, M. K. (1998).
Sequential prediction of individual sequences under general loss functions. IEEE
Transactions on Information Theory, 44(5):1906{1925.</p>
      <p>Haussler et al., 1994. Haussler, D., Littlestone, N., and Warmuth, M. K. (1994).
Predicting f0, 1g-functions on randomly drawn points. Information and Computation,
115(2):248{292.</p>
      <p>Hibon and Makridakis, 1997. Hibon, M. and Makridakis, S. (1997). Arma models and
the box{jenkins methodology.</p>
      <p>Holt, 2004. Holt, C. C. (2004). Forecasting seasonals and trends by exponentially
weighted moving averages. International journal of forecasting, 20(1):5{10.
Hu et al., 2007. Hu, W., Tong, S., Mengersen, K., and Connell, D. (2007). Weather
variability and the incidence of cryptosporidiosis: comparison of time series poisson
regression and sarima models. Annals of epidemiology, 17(9):679{688.
Hyndman and Athanasopoulos, 2014. Hyndman, R. J. and Athanasopoulos, G.</p>
      <p>(2014). Forecasting: principles and practice. OTexts.</p>
      <p>Jones and Brelsford, 1967. Jones, R. H. and Brelsford, W. M. (1967). Time series with
periodic structure. Biometrika, 54(3-4):403{408.</p>
      <p>Kuznetsov and Mohri, 2016. Kuznetsov, V. and Mohri, M. (2016). Time series
prediction and online learning. In 29th Annual Conference on Learning Theory, pages
1190{1213.</p>
      <p>Littlestone and Warmuth, 1989. Littlestone, N. and Warmuth, M. K. (1989). The
weighted majority algorithm. In Foundations of Computer Science, 1989., 30th
Annual Symposium on, pages 256{261. IEEE.</p>
      <p>Liu et al., 2016. Liu, C., Hoi, S. C., Zhao, P., and Sun, J. (2016). Online arima
algorithms for time series prediction. In Thirtieth AAAI Conference on Arti cial
Intelligence.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Dai</surname>
          </string-name>
          ,
          <year>2012a</year>
          . (
          <year>2012a</year>
          ).
          <article-title>Daily maximum temperatures in Melbourne, Australia</article-title>
          . Australian Bureau of Meteorology. https://datamarket.com/data/set/2323/ daily-maximum
          <article-title>-temperatures-in-</article-title>
          <string-name>
            <surname>melbourne</surname>
          </string-name>
          \-australia
          <string-name>
            <surname>-</surname>
          </string-name>
          1981-1990#!ds=2323&amp; display=line.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Dai</surname>
          </string-name>
          ,
          <year>2012b</year>
          . (
          <year>2012b</year>
          ).
          <article-title>Daily minimum temperatures in Melbourne, Australia</article-title>
          . Australian Bureau of Meteorology. https://datamarket.com/data/set/2324/ daily-minimum
          <article-title>-temperatures-in-</article-title>
          <string-name>
            <surname>melbourne</surname>
          </string-name>
          \-australia
          <string-name>
            <surname>-</surname>
          </string-name>
          1981-1990#!ds=2324&amp; display=line.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          et al.,
          <year>2010</year>
          . Ahmed,
          <string-name>
            <given-names>N. K.</given-names>
            ,
            <surname>Atiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F.</given-names>
            ,
            <surname>Gayar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. E.</given-names>
            , and
            <surname>El-Shishiny</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>An empirical comparison of machine learning models for time series forecasting</article-title>
          .
          <source>Econometric Reviews</source>
          ,
          <volume>29</volume>
          (
          <issue>5-6</issue>
          ):
          <volume>594</volume>
          {
          <fpage>621</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Pankratz</surname>
          </string-name>
          ,
          <year>2012</year>
          . Pankratz,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Forecasting with dynamic regression models</article-title>
          , volume
          <volume>935</volume>
          . John Wiley &amp; Sons.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          Press,
          <year>2007</year>
          . Press,
          <string-name>
            <surname>W. H.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Numerical recipes 3rd edition: The art of scienti c computing</article-title>
          . Cambridge university press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Romanenko</surname>
          </string-name>
          ,
          <year>2015</year>
          . Romanenko,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Aggregation of adaptive forecasting algorithms under asymmetric loss function</article-title>
          .
          <source>In International Symposium on Statistical Learning and Data Sciences</source>
          , pages
          <volume>137</volume>
          {
          <fpage>146</fpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Strang</surname>
          </string-name>
          ,
          <year>1994</year>
          . Strang,
          <string-name>
            <surname>G.</surname>
          </string-name>
          (
          <year>1994</year>
          ). Wavelets. American Scientist,
          <volume>82</volume>
          (
          <issue>3</issue>
          ):
          <volume>250</volume>
          {
          <fpage>255</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Vovk</surname>
          </string-name>
          ,
          <year>1992</year>
          . Vovk,
          <string-name>
            <surname>V.</surname>
          </string-name>
          (
          <year>1992</year>
          ).
          <article-title>Universal forecasting algorithms</article-title>
          .
          <source>Information and Computation</source>
          ,
          <volume>96</volume>
          (
          <issue>2</issue>
          ):
          <volume>245</volume>
          {
          <fpage>277</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Vovk</surname>
          </string-name>
          ,
          <year>2001</year>
          . Vovk,
          <string-name>
            <surname>V.</surname>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Competitive on-line statistics</article-title>
          . International Statistical Review/Revue Internationale de Statistique, pages
          <volume>213</volume>
          {
          <fpage>248</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Vovk</surname>
          </string-name>
          and Zhdanov,
          <year>2009</year>
          . Vovk,
          <string-name>
            <given-names>V.</given-names>
            and
            <surname>Zhdanov</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Prediction with expert advice for the brier game</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>10</volume>
          (Nov):
          <volume>2445</volume>
          {
          <fpage>2471</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Vovk</surname>
          </string-name>
          ,
          <year>1990</year>
          . Vovk,
          <string-name>
            <surname>V. G.</surname>
          </string-name>
          (
          <year>1990</year>
          ).
          <article-title>Aggregating strategies</article-title>
          .
          <source>In Proc. Third Workshop on Computational Learning Theory</source>
          , pages
          <volume>371</volume>
          {
          <fpage>383</fpage>
          . Morgan Kaufmann.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Vovk</surname>
          </string-name>
          ,
          <year>1995</year>
          . Vovk,
          <string-name>
            <surname>V. G.</surname>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>A game of prediction with expert advice</article-title>
          .
          <source>In Proceedings of the eighth annual conference on Computational learning theory</source>
          , pages
          <volume>51</volume>
          {
          <fpage>60</fpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Walker</surname>
          </string-name>
          ,
          <year>1988</year>
          . Walker,
          <string-name>
            <surname>J. S.</surname>
          </string-name>
          (
          <year>1988</year>
          ).
          <article-title>Fourier analysis</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>White</surname>
          </string-name>
          ,
          <year>1989</year>
          . White,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>1989</year>
          ).
          <article-title>Learning in arti cial neural networks: A statistical perspective</article-title>
          .
          <source>Neural computation</source>
          ,
          <volume>1</volume>
          (
          <issue>4</issue>
          ):
          <volume>425</volume>
          {
          <fpage>464</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Yule</surname>
          </string-name>
          ,
          <year>1909</year>
          . Yule,
          <string-name>
            <surname>G. U.</surname>
          </string-name>
          (
          <year>1909</year>
          ).
          <article-title>The applications of the method of correlation to social and economic statistics</article-title>
          .
          <source>Journal of the Royal Statistical Society</source>
          ,
          <volume>72</volume>
          (
          <issue>4</issue>
          ):
          <volume>721</volume>
          {
          <fpage>730</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>