<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LSTM for model-based Anomaly Detection in Cyber-Physical Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benedikt Eiteneuer</string-name>
          <email>benedikt.eiteneuer@hs-owl.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oliver Niggemann</string-name>
          <email>oliver.niggemann@hs-owl.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute Industrial IT, OWL University of Applied Sciences</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Anomaly detection is the task of detecting data which differs from the normal behaviour of a system in a given context. In order to approach this problem, data-driven models can be learned to predict current or future observations. Oftentimes, anomalous behaviour depends on the internal dynamics of the system and looks normal in a static context. To address this problem, the model should also operate depending on state. Long Short-Term Memory (LSTM) neural networks have been shown to be particularly useful to learn time sequences with varying length of temporal dependencies and are therefore an interesting general purpose approach to learn the behaviour of arbitrarily complex Cyber-Physical Systems. In order to perform anomaly detection, we slightly modify the standard norm 2 error to incorporate an estimate of model uncertainty. We analyse the approach on artificial and real data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The diagnosis of time-dependent systems has always been a
focus of research in domains such as industrial applications,
robotics, medical diagnosis and many more [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Typical
applications of diagnosis involve the detection of
anomalous system behaviour, system degradation or sub-optimal
conditions concerning system health, energy consumption
or product quality.
      </p>
      <p>Traditionally, manual models or simulations based on the
system description and its physics have been created by
experts of their domain. However, the ever-increasing
complexity of distributed multi-component Cyber-Physical
Systems (CPS) have made these manual efforts increasingly
difficult. Automated data-driven machine learning applications
fill in the need of model creation to describe the system
behaviour [3]. Because in many such systems a safe and
reliable functioning of its components is a vital requirement, it
is important to detect anomalous or faulty system behaviour
in real time and as early as possible.</p>
      <p>
        In general, a data sample can be considered
anomalous if its generating distribution is significantly different
from the normal behaviour. Usually, such anomalies are
not observed abundantly on real systems because it is too
expensive and sometimes even dangerous to record such
data. Furthermore, even if one is able to record (and
label!) anomalous data properly, it would still not be
sufficiently exhaustive for a classification task, because
anomalies are fundamentally diverse. Instead, one typically
assumes a (sub)set of recorded data to not contain any
anomalies, thereby labelling all training data as normal, called
one-class-classification [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The training data is then used
to learn a model of the normal behaviour. When doing
inference, new data is compared to the expectation and then
classified, see [5] for a recent review of common such
methods.
      </p>
      <p>Many commonly used techniques do not naturally make
use of the time ordering and are therefore unable to detect
certain kind of temporal anomalies which depend on the
systems internal state. Methods that try to adapt commonly
used approaches into the time domain, such as sliding
windows usually only model fixed time temporal dependencies,
while methods analysing spectral properties such as FFT
rely on proper periodic lengths.</p>
      <p>
        Long Short-Term Memory [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (LSTM) networks are a
natural candidate to fill in this gap because they keep a
memory of its past input history and are theoretically able to learn
arbitrary lengths of temporal patterns. These internal,
longterm memory states can be shaped by the learning algorithm
to be useful in order to make present and future predictions
for anomaly detection. LSTM have been used for exactly
this purpose, see e.g. [
        <xref ref-type="bibr" rid="ref8">7, 8, 9</xref>
        ].
      </p>
      <p>However, for the purpose of anomaly detection, it is not
sufficient to compare a model’s predictions to the actual
values. In order to assess how much to trust the model, an
estimate of uncertainty is necessary.</p>
      <p>
        The authors of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] approach this problem by fitting a
multivariate normal distribution, [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ] use variational
inference while [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] use a Monte Carlo dropout scheme to
handle model uncertainty.
      </p>
      <p>In this paper we go a more direct route by learning
explicit uncertainty predictions. We make the following
contributions:</p>
      <p>Modify and motivate the standard loss function to
incorporate estimates of noise inherent to the data. The
problem of anomaly detection can only be tackled if
anomaly scores can be directly related to probabilities
of events given some assumptions.</p>
      <p>Demonstrate the advantages of modelling internal state
dynamics in order to properly monitor dynamic
systems.</p>
      <p>Analyse the learned state representation of LSTM in</p>
    </sec>
    <sec id="sec-2">
      <title>CPS data with respect to anomalous data.</title>
      <p>The Paper is structured as follows: In Section 2, we start
off by describing the general features of CPS that are
important to the problem. We then derive a modified loss function
which enables a model to predict the noise along with the
data. After a brief description of the LSTM model, we
describe the anomaly detection approach by incorporating the
prediction uncertainty into the anomaly classification task.
In Section 3, we analyse the approach with a simple artificial
level control system and a real-world power consumption
time series data set. Discussion and some general remarks
are given in Section 4. We conclude in Section 5.
A CPS is an isolated, dynamic system which interacts
with its environment via signal inputs xt 2 X Rp
and outputs yt 2 Y Rq, where t denotes time, p and
q in- and output dimensions repectively. Amongst
others, a CPS can be comprised of mechanical, electrical
or biological parts.</p>
      <p>Although the underlying system is continuous, all
system in- and output is given in discrete time steps.
At any given point in time, the CPS has an internal
state st 2 S Rn. A state is the minimal description
needed to theoretically be able to predict the outputs yt
as well as the next state, if only the current input xt is
given in addition.</p>
      <p>The system inputs x can indicate discrete or
continuous events triggering a state change. These can be
external or internal influences which have an
environmental, cybernetic (e.g. actor signals), human or even
random origin.</p>
      <p>The system outputs y are observations taken on the
system that depend both on its internal state and its
input. In general, such observations could be given by
some arbitrary non-linear function f (s; x).</p>
      <p>The observation vector y is subject to noise and can
therefore only be determined up to some leftover
covariance, Vt 2 Rq q</p>
      <p>In order to monitor the system behaviour, the vector of
measured observables yt has to be learned for each step in
time t. It can then be compared to the actual value and
information about the system health, such as degradation and
anomalous or untypical behaviour, can be inferred. Because
Vt is not known a priori and its knowledge is important to
the anomaly classification task, its determination is part of
the learning process.</p>
      <p>If complete information of a deterministic CPS is
available in the x and s variables, it is possible to find a mapping
f : (st 1; xt) 7! (yt; Vt). However, the internal state of the
system is not actually measured, but only the vector of
observations y. Therefore, the st variable has to be modelled
along with the data by inference of the system behaviour
in time as well as the inter-dependencies of the xt and yt
variables, see Figure 1.</p>
    </sec>
    <sec id="sec-3">
      <title>Output:</title>
      <p>y 2 Y</p>
    </sec>
    <sec id="sec-4">
      <title>System:</title>
      <p>has state: s 2 S</p>
    </sec>
    <sec id="sec-5">
      <title>Input:</title>
      <p>x 2 X</p>
    </sec>
    <sec id="sec-6">
      <title>Training</title>
    </sec>
    <sec id="sec-7">
      <title>Inference</title>
    </sec>
    <sec id="sec-8">
      <title>Learned Model:</title>
      <p>
        should model state
A simple feed-forward neural net can approximate any
function [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and is therefore a good candidate to find the desired
mapping f , if all information were given along with x and
the system had no internal state. If this is not the case, a
state can be modelled within a neural network by feeding
information about past input through time, which is exactly
what a recurrent architecture does. LSTM’s are particularly
equipped for this task, because a gating mechanism allows
gradients to flow back in time up to arbitrary lengths,
enabling the learning of diverse temporal dependencies. We
would like to remind the reader, that if those temporal
patterns are well known (and maybe not too long) it is not
necessary to learn such states. In cases such as these we could
just group the input up to the desired point back in time and
analyse the aggregate, e.g. with a simple neural network.
The strength of learning internal state representations lies in
no prior knowledge given about possible temporal
dependencies.
      </p>
      <p>The problem is now to find a mapping of the form:
f : (st 1; xt) 7! (st; yt; Vt):
Assumption 1. It is possible to model these states. This
means that all relevant information for the prediction of
outputs yt can be inferred from the dynamic behaviour.</p>
      <p>It is then possible to predict yt up to some leftover
statistical uncertainty. If this were not the case, missing
information could only be compensated for with additional in- or
output variables or by adding more data. If the assumption
can not be met, there will be a number of unknown
internal states that are not modelled properly. System behaviour
that depends on these hidden states can then not properly be
accounted for.</p>
      <p>We denote the model prediction with y^t(st 1; xtj ),
being the model parameter.</p>
      <p>Assumption 2. The leftover covariance of model-to-actual
difference y^t yt is Gaussian.</p>
      <p>This is motivated by the idea of including as much
information as possible in the system state and the ability to
propagate this information into the model outputs, which
y^t</p>
      <p>t
Fully Connected</p>
      <p>Next LSTM Layer
leaves maximum entropy in the prediction difference. Note
that the normal distribution maximises the entropy for a real
valued random variable with specified mean and variance.
The whole process can be understood similar to a
multivariate Gaussian random variable in a Gaussian process with
mean y^t and white noise, meaning: cov(yt; yt0 ) = tt0 Vt
yt is then distributed normally with mean y^t and some
covariance matrix Vt(st 1; xtj ):</p>
      <p>q 1
p (ytjst 1;xt) = (2 ) 2 jVtj 2
1
2
exp
(yt
y^t)T Vt 1 (yt
y^t) :
Assumption 3. All elements of yti 2 yt are conditionally
independent from each other given xt and st 1.</p>
      <p>In a physical system, this corresponds to the fact, that two
sensors might take measurements within the same
environment and therefore are (possibly maximally) correlated, but
given the complete internal state of the system, the
remaining noise on each sensor is independent from each other.
(They are two different sensors.)</p>
      <p>Vt is therefore diagonal and p (ytjst 1; xt) factorises to
the following expression:
p (ytjst 1; xt) =
q
Y
i=0</p>
      <p>1
p2
i
t
exp
" 1
2
yti
i
t
y^ti 2#
;
(2)
where ti are the diagonal elements of Vt and can be
summarised in the vector t.</p>
      <p>Maximising the likelihood function (2) is equivalent to
minimising the negative log likelihood loss function:
Lt = X
i
" yti
i
t
y^ti 2
+ 2 log ti
#
2.3</p>
      <p>
        Long-Short Term Memory Network
We will employ a vanilla LSTM [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] network to find f and
model all of the y^t and t (as its short-term memory state) as
well as the hidden state st (as its long-term memory states)
and use supervised learning with input xt on the targets yt
with stochastic gradient descent to train the weights of the
parametric function.
(1)
(3)
      </p>
      <p>Writing t = exp( t) guarantees &gt; 0 and provides
numerical stability in the learning process. The network
architecture consists of several (1-5) fully recurrent, stacked
LSTM layers with a subsequent fully connected tanh
activation neural network with 2q output channels, one for each
yi and i, see Figure 2. Gradients are truncated in time
after tmax steps, typically between 5 and 50. Optimised is the
sum over Lt containing all terms from the current time t to
t tmax steps into the past. The training can be realised in a
mini-batch process by going through the data sequences in
parallel.</p>
      <p>
        The models used for this work are implemented within
the TENSORFLOW [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] machine learning framework.
      </p>
      <p>In the learning stage the neural network will optimise its
parameters in order to model the domain dependent
functions y^ and as accurate as possible, see Figure 3. Put
differently, the neural net learns the domain mean, taking
for each known state the average over the known
observations falling into that category. The method is therefore self
regularising as long as enough data exists for each domain
to determine sensible mean and standard deviation values.
While this can be said for any sufficiently simple model, the
strength of this approach lies in the learning and grouping of
temporal patterns into states which are forced to be treated
unified within the model.
One advantage of modelling the noise along with the system
output can be understood in the first signal of the Figure 3.
Without learning noisy patterns, one would be far less
sensitive to deviations of order 1 in the regimes with little to zero
noise, because the same signal exhibits a noise level of the
same order of magnitude in another regime. By modelling
the noise, anomaly detection can be performed for every
observed signal independently of different system behaviour
(through time).</p>
      <p>To discriminate between normal and anomalous data, we
can calculate the normalised residuals rt = (y^t yt)= t
which should be distributed normally:
p(rtjst 1; xt) = N (rt; 0; 1);
(4)</p>
      <p>Figure 4 shows the graph of the residuals from Figure 3
The water tank system is a simple and common control
system, see Figure 5. Water is constantly flowing out of a
reservoir according to the root square law which relates the
output flow qout to the water level h: qout = aph. As soon as a
certain lower threshold is reached, a valve is opened which
allows the tank to be refilled, all the while it is constantly
being drained from the bottom. Once the level reaches an
upper threshold, the valve is closed again.</p>
      <p>Input</p>
      <p>qout
Upper threshold
Lower Threshold</p>
      <p>Output
as well as a histogram of its distribution. The histogram of
the first signal (noise with 2 modal variance) resembles a
normal distribution and shows that the LSTM correctly
approximates the different noise levels depending on the
current progress in the period. The second signal has a uniform
noise, which cannot be modelled perfectly. Strictly
speaking, the Assumption 2 is not fullfilled in this case.
However, the variance can still be learned. As is clear for the
purpose of anomaly detection, the only important issue is
that residual outliers should be similarly sparse in the tails
of the distribution. Most problematic cases would be
datadistributions with very long/flat tails which could result in
false positive classifications because the learned variance is
a bad measure for the decision boundary.</p>
      <p>When should samples be considered as anomalous?
Generally, it is possible to combine several samples and/or
features and then analyse the aggregate statistics. A simple
method is to threshold the normal distribution given the test
point rt. Such a threshold can either be learned (as is often
done in the literature [5]) on the training or held out data to
include a specified percentile of the training data (e.g.
95100%) or calculated to include a specified percentile of the
expected distribution.</p>
      <p>If a percentile is calculated from the normal
distribution (e.g. r = 3 corresponds to a 1 2:7 10 3
quantile), this is the expected frequency of points to be
classified as anomalous (False positive rate). Of course,
anomalous points themselves are not expected to be distributed
normally and thus should exceed the threshold value much
more frequently.</p>
      <p>The strength of this approach is that it is not necessary to
learn the threshold with held out data, possibly containing
known anomalies. Instead, it is not necessary at all to have
ever measured an anomaly, which is how it should be.
Approximate statements about the expected frequency of false
positives can be made and thresholds can be chosen
according to ones preferences given the size of the test data and
desired sensitivity.</p>
      <p>In order to monitor system change over time, again the
residual statistics can be used. In a typical degradation
scenario at least one of the observables usually exhibits more
and more anomalous behaviour.
3</p>
      <sec id="sec-8-1">
        <title>Experiments</title>
        <p>In this section, we test the approach on one artificial level
control system and the power demand data set.</p>
        <p>This system has an internal state (the water level) which
we assume to be not measured directly as well as a clear
dependency of current measurements from the past history. A
certain output flux qout;t at any time will result in a larger or
smaller flux in the future depending if the valve was opened
or closed. The LSTM input is a binary signal about the
valve’s state (open/closed), while the output flux is to be
predicted.</p>
        <p>Figure 6 shows the learned representation of the long term
memory state of the simplest LSTM architecture (with only
one layer and one internal state) together with the actual
water level. In this case, a monotonic behaviour suggest that
a meaningful representation of the systems “true" internal
state has been achieved. The water level itself, which can be
considered as an ideal state representation, was not part of
the learning process, but only the process outflow qout:</p>
        <p>
          We test the learned model against a time sequence
containing a period of anomalous behaviour. To generate the
anomalies, at one point a system blockage is introduced
which reduces the output flux by 25%. This kind of
anomalous behaviour is not detectable with state-independent
methods which do not model time in any way, as long as
the flux is still large enough to be in the typical regime.
it is introduced because observations do not match
expectations immediately. Note that the first anomalous points
actually conform to the usual range of observations during
normal operation, and thus could not be detected by a time
independent anomaly detection approach. Figure 8 shows
the Long-Term memory state together with the valve control
signal before, during and after the blockage. It is interesting
to see, that even after the expected time for the valve to open
again passed, the internal state is further reduced, matching
the actual water level.
The Power demand dataset [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] contains a single signal
which is the power consumption of a dutch research
institute recorded every 15 minutes over a total period of one
year. Several characteristics make this dataset interesting as
a toy real world dataset. It contains both short term (daily)
and long term (weekly) temporal patterns characterised by
increased energy consumption during work days. It contains
“anomalies" in the form of holidays with significantly lower
energy consumption (see Figure 9). Additionally, there is
a seasonal shift as the power demand decreases during the
summer, only to increase again at the end of the year.
        </p>
        <p>We use the data from January 2. to March 21. as training
data because the seasonal shift is slow enough and there are
no holidays during this period. We will treat this problem
analogous to a control system, where system input consists
of two Boolean variables marking the beginning of each day
and each weak. (Which is zero all the time except once a
day/week, where it is one). The task is to predict the current
power consumption.</p>
        <p>Figure 9 shows the model predictions ( 1 -band)
together with actual data from training and test set (during
summer). Additionally, two of a total of 10 internal
longterm memory states that are weakly associated with daily
and weekly patterns are shown. About 3/10 states
resembled a daily pattern, 4/10 weekly while the rest could not be
associated with a weekly or daily pattern. The top of Figure
10 shows the residuals of the training set after convergence.
The distribution can hardly be distinguished from a normal
distribution which suggests a successful training stage.</p>
        <p>The bottom of Figure 10 illustrates the anomaly detection
approach. To indicate the summer period we empirically
choose a threshold of 4 standard deviations, while anomalies
should exceed 8 (without further optimisation). All
holidays (days 87,90,120,125,128,139,359,360) have been
detected with two additional false positives at day 365 (which
exhibits a normal day pattern followed by a large power
drop) and day 277 (which exhibits a pattern close to a
working day, but on a Saturday), see Figure 11.</p>
      </sec>
      <sec id="sec-8-2">
        <title>Discussion &amp; Remarks</title>
        <p>
          Another study of anomaly detection with the power demand
data set can be found in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Here, the authors applied
autoregressive LSTM-sequence learning to the task, obtaining
good results (0.94 precision, 0.17 recall and F0:1 = 0:90).
Precision is the fraction of anomalous data amongst all
positively labelled instances, while recall is the fraction of
detected anomalies. The F score is the weighted harmonic
mean of precision and recall, with weighing the
importance of precision over recall. It should be noted that all
observations on a holiday are to be labelled as anomalous,
although most of the day (especially the night cycle) looks
like any other day. This naturally leads to a high false
positive rate for this problem and motivates to focus on a F
score with low , putting a low weight on false positives.
        </p>
        <p>There are a number of differences to our approach. First,
we deliberately abdicate of the use of a validation set
containing anomalies in order to fine-tune thresholds to make
the model sensitive to the right kind of anomalies. Instead,
as is motivated by real world scenarios where no
anomalous behaviour has been observed, the threshold is chosen
more or less arbitrary, by inspecting the distribution of
normalised residuals of the training set (or alternatively of a
small validation set). Second, our approach is not
autoregressive but relies on input related to changes of system
state. This makes the anomaly detection approach more
robust to anomalous data (given also as input to the
autoregressive formalism before being compared to the model
output), but less versatile to different kinds of data, where
such input is not readily available or hard to construct given
high-dimensional time-series data. Using the
aforementioned test-set as well as an 8 anomaly threshold, final
analysis results in a precision of 0.95, a 0.31 recall and
F0:1 = 0:93. Our method is sensitive to the missing power
consumption on a holiday far earlier than the auto-regressive
method simply by being independent of noisy anomalous
model input, misleading the algorithm to assume a late rise
instead of a missing day-pattern. Because data is noisy, the
positively sloped edge at the beginning of the shift occurs
at slightly different times each day, which can be accounted
for by proper uncertainty modelling.</p>
        <p>
          The problem of model uncertainty, in particularly with
regards to neural networks is still an open question of
research, but definitely necessary to quantify in the domain
of anomaly detection. In this work, only the noise which
is inherent in any data is modelled. Systematic uncertainty
because of missing information or insufficient model
complexity is not addressed (see [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] for a general overview of
uncertainty modelling in deep learning approaches).
        </p>
        <p>
          Experiments showed that learning the variance along with
the mean makes learning by gradient descent more difficult,
even when using optimisers such as Adam [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] or
Momentum. It was beneficial to keep the variables frozen in a first
learning stage.
        </p>
        <p>The kind of data modelling described in this approach
requires a distinction of in- and output variables. Only the
outputs are modelled and checked for anomalies. If
anomalies are present in the inputs, the faulty system input should
lead to a miss-prediction of output variables y and therefore
trigger the anomaly classification. However, this is not
guaranteed and needs to be investigated in greater detail. The
Figures 7 and 8 show, that the LSTM is able to
extrapolate beyond the normally observed patterns in a meaning full
way, indicating that even anomalous input could generate a
“correct” (compared to the data) but nevertheless anomalous
(given the context) output.</p>
        <p>For many real world systems it is a problem to clearly
distinguish between the x and y variables. A simple
partition into actor and sensor signals which is often possible in
control theory is not applicable to every system. One should
keep in mind to always include all exogenous variables in x,
because they cannot be predicted by any other variable. In
many CPS, some of these signals are binary in nature, such
as on/off commands. Other possible candidates are
environmental measurements that are not or only barely influenced
by the system that is to be monitored.</p>
        <p>
          A clear disadvantage of requiring a hard separation
between in- and output variables is limited use in systems
where such input can inherently not exist (e.g. heart rate
monitoring or similar systems). Of course it is possible to
adapt the model to include the output variables as input in
something like an encoder-decoder recurrent neural network
[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. However, for the purpose of anomaly detection it is
not guaranteed that every anomalous input is always going
to generate an output significantly different to the observed
value.
        </p>
        <p>Although not discussed in this work, it is generally
possible to use the approach described in this paper for time
series forecast. If the output prediction of observation
vector y is to be estimated into the future one changes the
training procedure for the LSTM to predict ahead of time:
f : (st 1; ::; st+t0 1; xt) 7! (st+t0 ; yt+t0 ) with t0
determining the shift into the future.
5</p>
      </sec>
      <sec id="sec-8-3">
        <title>Conclusions</title>
        <p>We presented a model based anomaly detection for CPS
data based on state dependent LSTM neural networks. In
order to train the LSTM, a modified loss function
intrinsically including the problem of noise prediction is derived.
Equipped with this characteristic, LSTM are able to learn
the temporal patterns inherent in CPS data and can
output reliable predictions of observations made on the system.
Predictions of both signal and noise can then be used to
detect anomalies or system deterioration.</p>
      </sec>
      <sec id="sec-8-4">
        <title>Acknowledgments</title>
        <p>The authors thank Nemanja Hranisavljevic for helpful
discussions and remarks. This work has received funding from
the European Unions Horizon 2020 research and innovation
programme under grant agreement No. 678867.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Zhiwei</given-names>
            <surname>Gao</surname>
          </string-name>
          , Carlo Cecati, and
          <string-name>
            <surname>Steven</surname>
            <given-names>X</given-names>
          </string-name>
          <string-name>
            <surname>Ding</surname>
          </string-name>
          .
          <article-title>A survey of fault diagnosis and fault-tolerant techniquespart i: Fault diagnosis with model-based and signalbased approaches</article-title>
          .
          <source>IEEE Transactions on Industrial Electronics</source>
          ,
          <volume>62</volume>
          (
          <issue>6</issue>
          ):
          <fpage>3757</fpage>
          -
          <lpage>3767</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] [3]
          <string-name>
            <given-names>Oliver</given-names>
            <surname>Niggemann</surname>
          </string-name>
          , Benno Stein, Asmir Vodencarevic, Alexander Maier, and Hans Kleine Büning.
          <article-title>Learning behavior models for hybrid timed systems</article-title>
          .
          <source>In Proceedings of the Twenty-Sixth AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>1083</fpage>
          -
          <lpage>1090</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Oliver</given-names>
            <surname>Niggemann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Volker</given-names>
            <surname>Lohweg</surname>
          </string-name>
          .
          <article-title>On the diagnosis of cyber-physical production systems: state-ofthe-art and research agenda</article-title>
          . Austin, Texas, USA,
          <source>Association for the Advancement of Artificial Intelligence</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Shehroz</surname>
            <given-names>S</given-names>
          </string-name>
          <string-name>
            <surname>Khan and Michael G Madden</surname>
          </string-name>
          .
          <article-title>Oneclass classification: taxonomy of study and review of techniques</article-title>
          .
          <source>The Knowledge Engineering Review</source>
          ,
          <volume>29</volume>
          (
          <issue>3</issue>
          ):
          <fpage>345</fpage>
          -
          <lpage>374</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Marco AF Pimentel</surname>
            , David A Clifton,
            <given-names>Lei</given-names>
          </string-name>
          <string-name>
            <surname>Clifton</surname>
            , and
            <given-names>Lionel</given-names>
          </string-name>
          <string-name>
            <surname>Tarassenko</surname>
          </string-name>
          .
          <article-title>A review of novelty detection</article-title>
          .
          <source>Signal Processing</source>
          ,
          <volume>99</volume>
          :
          <fpage>215</fpage>
          -
          <lpage>249</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <article-title>Long shortterm memory</article-title>
          .
          <source>Neural Computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>Using lstm recurrent neural networks for monitoring the lhc superconducting magnets</article-title>
          .
          <source>Nuclear Instruments and Methods in Physics Research Section A: Accelerators</source>
          , Spectrometers, Detectors and Associated Equipment,
          <volume>867</volume>
          :
          <fpage>40</fpage>
          -
          <lpage>50</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Timothy J O'Shea</surname>
            ,
            <given-names>T Charles</given-names>
          </string-name>
          <string-name>
            <surname>Clancy</surname>
          </string-name>
          , and
          <string-name>
            <surname>Robert W McGwier</surname>
          </string-name>
          .
          <article-title>Recurrent neural radio anomaly detection</article-title>
          .
          <source>arXiv preprint arXiv:1611.00301</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Anvardh</given-names>
            <surname>Nanduri</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lance</given-names>
            <surname>Sherry</surname>
          </string-name>
          .
          <article-title>Anomaly detection in aircraft data using recurrent neural networks (rnn)</article-title>
          .
          <source>In Integrated Communications Navigation and Surveillance (ICNS)</source>
          ,
          <year>2016</year>
          , pages
          <fpage>5C2</fpage>
          -
          <lpage>1</lpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Pankaj</surname>
            <given-names>Malhotra</given-names>
          </string-name>
          , Lovekesh Vig, Gautam Shroff, and
          <article-title>Puneet Agarwal. Long short term memory networks for anomaly detection in time series</article-title>
          .
          <source>In Proceedings of 23. European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN)</source>
          ,
          <year>page</year>
          89. Presses universitaires de Louvain,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Meire</surname>
            <given-names>Fortunato</given-names>
          </string-name>
          , Charles Blundell, and
          <string-name>
            <given-names>Oriol</given-names>
            <surname>Vinyals</surname>
          </string-name>
          .
          <article-title>Bayesian recurrent neural networks</article-title>
          .
          <source>arXiv preprint arXiv:1704.02798</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Maximilian</surname>
            <given-names>Sölch</given-names>
          </string-name>
          , Justin Bayer, Marvin Ludersdorfer, and Patrick van der Smagt.
          <article-title>Variational inference for on-line anomaly detection in high-dimensional time series</article-title>
          .
          <source>arXiv preprint arXiv:1602.07109</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Lingxue</given-names>
            <surname>Zhu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nikolay</given-names>
            <surname>Laptev</surname>
          </string-name>
          .
          <article-title>Deep and confident prediction for time series at uber</article-title>
          .
          <source>In Data Mining Workshops (ICDMW)</source>
          ,
          <year>2017</year>
          IEEE International Conference on, pages
          <fpage>103</fpage>
          -
          <lpage>110</lpage>
          . IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kurt</surname>
            <given-names>Hornik</given-names>
          </string-name>
          , Maxwell Stinchcombe,
          <string-name>
            <given-names>and Halbert</given-names>
            <surname>White</surname>
          </string-name>
          .
          <article-title>Multilayer feedforward networks are universal approximators</article-title>
          .
          <source>Neural networks</source>
          ,
          <volume>2</volume>
          (
          <issue>5</issue>
          ):
          <fpage>359</fpage>
          -
          <lpage>366</lpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Martín</given-names>
            <surname>Abadi</surname>
          </string-name>
          , Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis,
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matthieu</given-names>
            <surname>Devin</surname>
          </string-name>
          , Sanjay Ghemawat, Geoffrey Irving,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Isard</surname>
          </string-name>
          , et al.
          <article-title>Tensorflow: A system for large-scale machine learning</article-title>
          .
          <source>In OSDI</source>
          , volume
          <volume>16</volume>
          , pages
          <fpage>265</fpage>
          -
          <lpage>283</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Yanping</surname>
            <given-names>Chen</given-names>
          </string-name>
          , Eamonn Keogh, Bing Hu, Nurjahan Begum, Anthony Bagnall, Abdullah Mueen, and
          <string-name>
            <given-names>Gustavo</given-names>
            <surname>Batista</surname>
          </string-name>
          .
          <article-title>The ucr time series classification archive</article-title>
          ,
          <year>July 2015</year>
          . www.cs.ucr.edu/~eamonn/time_series_data/.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Yarin</given-names>
            <surname>Gal</surname>
          </string-name>
          .
          <article-title>Uncertainty in deep learning</article-title>
          . University of Cambridge,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Diederik</surname>
            <given-names>P</given-names>
          </string-name>
          <string-name>
            <surname>Kingma and Jimmy Ba</surname>
          </string-name>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412.6980</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>