<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sensor Data Preprocessing, Feature Engineering and Equipment Remaining Lifetime Forecasting for Predictive Maintenance</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Proceedings of the XX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2018)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Evgeniy Latyshev Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>226</fpage>
      <lpage>231</lpage>
      <abstract>
        <p>Analytics based on sensor data is gradually becoming an industry standard in equipment maintenance. However, it involves several challenges, such as sensor data preprocessing, feature engineering and forecasting model development. Due to work in progress, this paper is mainly focused on sensor data preprocessing, which plays a crucial role in predictive maintenance due to the fact, that real-world sensing equipment usually provides data with missing values and a considerable amount of noise. Obviously, poor data quality can render practically useless all the following steps of data analysis. Thus, many missing data imputation, outlier filtering, and noise reduction algorithms were introduced in the literature. Streaming sensor data can be represented in a form of univariate time series. This paper provides an overview of common univariate time series preprocessing steps and the most appropriate methods, with consideration of the field of application. Sensor data from different sources comes in different scales and should be normalized. Thus, the comparison of univariate time series normalization techniques is given. Conventional algorithm quality metrics for each of the preprocessing steps are described. Basic sensor data quality assessment approach is suggested. Moreover, the architecture of a sensor data preprocessing module is proposed. The overview of time series-specific feature engineering techniques is given. The brief enumeration of considered forecasting approaches is provided.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Maintenance costs are a major part of the total operating
costs of any business involving complex equipment.
Conducted surveys of maintenance management
effectiveness indicate that one-third of all maintenance
costs is wasted as the result of unnecessary or improperly
carried out maintenance [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. With the spread of Internet
of Things concept, sensor data can be collected from a
huge amount of devices and equipment. This data can be
used for real-time health monitoring and effective
maintenance. However, this approach to maintenance,
also known as predictive maintenance, involves several
challenges.
      </p>
      <p>
        First of all, often collected data is of poor quality,
which can lead to unreliable analysis and ineffective
maintenance. Consequently, data from sensing
equipment needs to be preprocessed before it can be used
for any analysis. Poor data quality means
noncompliance with requirements on at least one of the data
quality assessment metrics. The root of problems can
vary: connection issues, sensor malfunction, transmitting
hardware failure, data processing server downtime,
software crash, measuring equipment inaccuracy and
many more. Common cases of poor data quality involve
unacceptable amount of missing values, outliers, sudden
spikes etc. Simply ignoring these issues can be critical
due to several reasons. For example, some analysis tools,
including popular machine learning algorithms, can’t
handle missing values. The absence of outlier filtration
can dramatically skew the results. Measuring equipment
standard error can be mistaken for an actual pattern in
data. As a result, time series preprocessing involves
several independent steps: missing data imputation,
noise reduction, and data normalization. After these
steps, data can be evaluated in quality and passed further
for analysis. It is clear, that preprocessing should be done
in near real-time to minimize the delay between data
measurement and decision making. Thus, there is a need
for a fast and scalable independent module, that can
preprocess constantly incoming sensor data. This paper
proposes the design of such a module, keeping in mind
the following integration with the existing architecture of
a predictive maintenance system, introduced in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>Secondly, it can be difficult to distinguish the patterns
and relationships in the initial data. The process of
extracting and generating new characteristics and
features out of the available data, commonly referred to
as feature engineering, has two main objectives. The first
one is to represent the data in such a form, that will make
it easier to establish simple yet strong connections
between the input and the output variables for the
forecasting model, increasing the quality of the forecasts.
The second objective is to pick the most useful features
out of all the available ones, reducing the amount of
computations of the forecasting model.</p>
      <p>Finally, a proper forecasting model is to be chosen
and implemented. There are various approaches to time
series forecasting, from straightforward methods like
naive method to way more sophisticated ones like
longshort term recurrent neural network. The main
complication here is the trade-off between the forecast
quality and the ease of model implementation and
deployment.</p>
      <p>The remaining part of the paper is organized as
follows. The preprocessing module architecture is
described in Section 2. Section 3 reviews missing data
imputation methods. Section 4 is devoted to time series
noise reduction. Section 5 briefly overviews data
normalization techniques. In section 6, some thoughts on
data quality assessment are combined. Section 7 is
devoted to time series feature engineering. The brief
overview of time series forecasting approaches is given
in section 8. Finally, the future directions of presented
work are given in section 9.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Preprocessing Module Architecture</title>
      <p>
        The preprocessing module is a part of the system for
predictive maintenance, deployed to a Hadoop [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]
cluster in a cloud manner. The module is wrapped in
Docker [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] container and runs on a standalone node of
the cluster. One of the key requirements for the module
is the seamless integration into the architecture. The data
is retrieved from Apache Kafka message queue [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
transformed by the preprocessing module and passed in
parallel to OpenTSDB [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and Apache Hive [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for
storage. To satisfy the requirements onto speed and
scalability the transformations are conducted onto
Apache Spark Streaming engine [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        There are many stream data processing frameworks
including but not limited to Apache Storm [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Apache
Flink [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Apache Samza [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Kafka Streams [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Although Spark Streaming has latency issues and sliding
window processing may be tricky due to Spark inherent
batch-based streaming model, is has several advantages
which make Spark Streaming a safer choice.
      </p>
      <p>
        First of all, Spark Streaming is a mature framework with
thorough documentation and huge community. As a
result of long-term popularity, there are plenty of
opensource tools for Spark Streaming, including solutions for
relatively painless integration with Kafka and mentioned
earlier database management systems [
        <xref ref-type="bibr" rid="ref11 ref21 ref28">28, 11, 21</xref>
        ].
Another advantage is the existence of pySpark [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] – an
API for Python, one of the biggest programming
languages at this moment. All other enumerated
frameworks require Scala, Clojure or Java knowledge,
which makes them less accessible.
      </p>
      <p>One of the biggest downsides of Spark Streaming is
performance degradation on sudden bursts of input data.
However, in case of sensor data processing the input data
flow intensity remains nearly the same at all time
intervals, which mitigates the downside.</p>
      <p>The data flow and module components are introduced
below in figure 1. The whole module consists of 4
transformation steps and the data quality assessment
step.</p>
    </sec>
    <sec id="sec-3">
      <title>3 Missing Data Imputation</title>
      <p>Sometimes due to a sensor malfunction, unstable internet
connection or other technical difficulties the data for
some points in time is missing. Simply ignoring those
gaps may be not the best strategy, because it can lead to
a loss of efficiency and unreliable results of the analysis.
Another approach is to try to impute the missing values
based on the available information.</p>
      <sec id="sec-3-1">
        <title>3.1 Methods</title>
        <p>
          The detailed overview of basic imputation methods and
their implementations can be found in imputeTS R
package documentation [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>Some simple methods that are applicable not only to
time series: median imputation, mode imputation, mean
imputation, random imputation. These methods are fast
and very straight-forward, but lack accuracy.</p>
        <p>Simple time series specific methods include LOCF
(last
observation
carried
forward),</p>
        <p>NOCB
(next
observation carried backward), interpolation (linear,
polynomial, Stineman) and moving average (simple,
weighted, exponential). All of them are rather fast and
can work in specific cases, but fall off when there is
seasonality in the data or large missing sub-sequences.</p>
        <p>More sophisticated approaches like Structural Model
&amp;</p>
        <p>
          However,
sensor
data
has
one
unfortunate
characteristic – the gaps of missing data can be too long
for conventional methods to work properly. In this case,
the method proposed in [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] can be the appropriate
choice. The idea of the Dynamic Time Warping Based
Imputation is to find the most similar sub-sequence to the
sub-sequence before the missing values, then complete
the gap by the next sub-sequence of the most similar one.
The result is a very plausible gap imputation with a
drawback of a huge computational cost.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Metrics</title>
        <p>Missing data imputation involves 2 types of quality
metrics based on the pattern of imputation.</p>
        <p>For single value imputations, the metrics coincide
with the ones commonly used in time series forecasting
– RMSE (Root Mean Square Error)and MAPE (Mean
Absolute Percentage Error).</p>
        <p>=
=</p>
        <p>100%
∑ ( ^ −  )2,</p>
        <p>× ∑
  − ^ ,
 
where yi is real value, ŷi is the forecasted value and n
is the number of forecasts.</p>
        <p>However, different metrics are used for long gap
imputation. The most popular of them are similarity and
Dynamic Time Warping distance.</p>
        <p>1
=
×

1 +</p>
        <p>
          DTW calculation algorithm can be found at [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. It is
worth mentioning, that modern implementations often
have adjustments to speed up the calculations (for
1
|  −  ^ |
( ^ ) −  
( ^ )
example, DDTW [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]).
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Noise Reduction</title>
      <p>Similar to missing data points, sensor data is usually
contaminated with noise, which can be mistaken for
actual data pattern, which yet again leads to a loss of
efficiency and unreliable results of the analysis. The task
of noise reduction is to subtract the maximum amount of
noise from the initial data, leaving the maximum amount
of useful signal.</p>
      <sec id="sec-4-1">
        <title>4.1 Methods</title>
        <p>
          According to Chen et al. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] noise reduction methods
can be divided into 2 categories: frequency domain
approaches and time domain approaches.
        </p>
        <p>Frequency domain approaches are based on signal
decomposition into frequency components. The most
common
approaches involve
discrete/fast/short-time
Fourier transform either wavelet transform.</p>
        <p>Most of the time domain approaches are based on
smoothing the signal of each given data point based on
the values of its neighbors.</p>
        <p>
          The comparison of the basic noise reduction methods
can be found in the work of Köhler et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. The
conducted
experiment involves the comparison
of
moving average filter, exponential smoothing filter,
linear Fourier smoothing, nonlinear wavelet shrinkage
and
        </p>
        <p>simple nonlinear noise reduction in different
conditions.</p>
        <p>
          The downside of the approaches listed above is that
they modify almost all the data values, most of which are
initially correct. Song et al. [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] proposed the first
constraint-based approach for cleaning stream data. The
idea is to sanity check the changes of values in time based
on subject area constraints. This method allows to detect
and repair large spike errors in data. The biggest
advantage of this method is the support of online
cleaning over streaming data.
        </p>
        <p>
          However, this method can be used only for large
outlier detection. In some cases, even small errors can be
important and repairing only spike errors is insufficient.
Zhang et al. [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] proposed a novel statistical-based
cleaning by introducing the repairment likelihoods with
respect to
speed
changes.
        </p>
        <sec id="sec-4-1-1">
          <title>Several</title>
          <p>effective
and
computationally efficient heuristics are also introduced
in this work.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 Metrics</title>
        <p>Most of the papers involve RMSE, defined earlier, as a
denoising quality metric. However, there are several less
popular ones, including the Symmetrical Visual Error</p>
        <sec id="sec-4-2-1">
          <title>Measure, proposed in [18].</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 Data Normalization</title>
      <p>Making sure that your data is of uniform scale is key
for many</p>
      <p>methods, including k-NN, linear models,
artificial neural
networks
and
many
more.</p>
      <sec id="sec-5-1">
        <title>Even</title>
        <p>univariate time series data should be normalized because
it might be further used in combination with data of
different scale from other sources.</p>
        <p>The most well-known and widely used are min-max
normalization
and
z-score
normalization.</p>
        <p>Min-max
implies that you know the minimum and the maximum
values in your dataset beforehand, which is often not the
case. Z-score is more robust but performs poorly om
nonstationary time series.</p>
        <p>ŷ    =
ŷ−</p>
        <p>−  ( )

=</p>
        <p>( )−  ( )
 −    ( ),
 
( )
where ŷ is a value after normalization, y is a value
prior normalization and Y is the set of values being
normalized.</p>
        <p>
          Some less popular methods are decimal scaling
normalization, which holds all the drawbacks of
minmax normalization, sigmoid normalization, which is
actively used in neural networks and tanh estimators,
which roughly can be described as a hyperbolic tangent
of the z-score normalization.
Data Quality Assessment (DQA) is the scientific and
statistical evaluation of data to determine if data obtained
from environmental data operations are of the right type,
quality, and quantity to support their intended use [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>
          There is a comprehensive work on time series data
quality assessment done in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which shows that there
are dozens of different metrics that can be used to
measure the quality of data. Obviously, using all of them
is excessive and computationally inefficient, so only a
few are to be chosen. However, there is no common view
on which metrics are better. The simple yet effective
strategy might be to look onto the most popular ones:
• event data loss (gaps in the data);
• values out of range (values out of sane interval for
the domain);
• value spikes (improbable sudden changes);
• wrong timestamps;
• rounded measurement value (not desirable level of
detail);
• signal noise (slightly inaccurate measurements).
        </p>
        <p>The assessment is to be done for both data prior and
after preprocessing to acquire an evaluation of
preprocessing module effectiveness. It is also worth
keeping in mind, that initially clean data is different to
the data, that was made “clean” during preprocessing due
to approximations and inevitable errors of the methods
involved on each step.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7 Feature Engineering</title>
      <p>Feature engineering is, probably, the most peculiar step
of data processing, as it depends on the initial data type,
its origin, quantity, quality, the desired output of the
forecasting model and even the nature of the model itself.
As it was already mentioned, sensor data can be
represented in a form of univariate time series. The
conventional approaches to time series feature
engineering can be divided into 3 categories: timestamp
features, statistical features, and spectral features. The
feature extraction step is usually followed by a
dimensionality reduction step.</p>
      <p>
        It is worth mentioning, that there are automatic time
series feature engineering tools such as tsfresh [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which
achieve decent results with almost no effort required.
      </p>
      <sec id="sec-6-1">
        <title>7.1 Timestamp Features</title>
        <p>The idea of this approach is to extract the features from
the timestamp of each observation. The most commonly
used features are:
• minutes elapsed for the day;
• hour of the day;
• day of the month;
• weekend or not;
• season of the year;
• public holiday or not.</p>
        <p>Talking about sensor data, some examples of useful
timestamp features are:
• time since the last maintenance;
• age of equipment;
• time since the last failure;
• operating time of equipment.</p>
        <p>Using just these features alone for predictions will
likely result in a poor model. However, in combination
with other features, they can boost the quality of
forecasts.</p>
      </sec>
      <sec id="sec-6-2">
        <title>7.2 Statistical Features</title>
        <p>This approach involves sliding through a time series with
the window of a given width and calculating statistics for
each iteration. The most common statistical features are
the mean of the previous few values, the median, the
mode, the minimal value, the maximum value, the
standard deviation and many more. In addition to
calculated statistics, we can also use the lagged values of
a time series as features.</p>
        <p>The biggest challenge of this approach is that the
window can be of any width and there is no general
algorithm to choose it. Usually, the researchers just try
out several values of the width and choose the one that
performs best. However, if there is a seasonal pattern in
data, it is worth making the width of the window not less
than the period of the seasons.</p>
      </sec>
      <sec id="sec-6-3">
        <title>7.3 Spectral Features</title>
        <p>Different variations of Fourier transform and wavelet
transform are used to extract spectral features from a
nonstationary signal. The basic idea behind those
methods is to decompose a given time series into a sum
of several basic functions, providing a different
representation of the initial signal. The biggest drawback
of those methods is that they are relatively
computationally expensive.</p>
      </sec>
      <sec id="sec-6-4">
        <title>7.4 Dimensionality Reduction</title>
        <p>Feature extraction provides many features, some of
which can be useless or strongly correlated with each
other. Excessive features not only add unnecessary
computations but also can decrease the quality of the
model. Thus, several dimensionality reduction methods
were introduced to minimize the number of features, at
the same time keeping the maximum amount of
information. The most common ones are principal
component analysis, independent component analysis
and partial least squares regression.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>8 Remaining Lifetime Forecasting</title>
      <p>There is a variety of methods, that can be used for time
series forecasting. Each of them has advantages and
drawbacks and can be viable in certain circumstances.</p>
      <p>First of all, there are some basic methods such as
average method, naive method, seasonal naive method,
and drift method, which are very simple yet can be
effective when the data pattern is easy.</p>
      <p>Secondly, linear regression models can be used for
forecasting. In the simplest case, the regression model
allows for a linear relationship between the forecast
variable and some predictor variables. The biggest
downside is inherent linearity, while real-world data is
mostly non-linear.</p>
      <p>The most common approach is to use stochastic
models – ARMA, ARIMA, SARIMAX, etc. One of the
biggest drawbacks of those models is that they require
fine-tuning of several hyperparameters, which is
computationally expensive and not intuitive.</p>
      <p>One of the recently popular approaches is to use
decision trees. Random forests and gradient boosting
methods, which are so widely used in machine learning
competitions, can also be used for time series
forecasting.</p>
      <p>Artificial neural networks approach for time series
forecasting gained immense popularity in last few years.
Their modifications – recurrent neural networks (RNNs)
and long short-term memory networks (LSTMs) are
especially effective for this task due to their “memory”
component. Although neural networks tend to be the most
accurate forecasting method when tuned properly and given
enough data, they might be computationally too costly for
the considered conditions.</p>
    </sec>
    <sec id="sec-8">
      <title>9 Conclusion</title>
      <p>In this study, the overview of sensor data preprocessing
steps, methods, and common metrics is held. Some
thoughts on sensor data quality assessment are given.
The architecture of a fast, scalable preprocessing module
is proposed. A brief overview of time series feature
engineering techniques and forecasting methods is given.
The future goals of the ongoing work are to implement
the designed preprocessing module on Spark Streaming
engine, integrate it into the existing predictive
maintenance pipeline, implement the feature engineering
step and to develop a remaining lifetime forecasting
model.</p>
      <p>Acknowledgments. This work is supervised by Dmitriy
Kovalev, Institute of Informatics Problems, Federal
Research Center “Computer Science and Control” of the
Russian Academy of Sciences.</p>
      <p>The research is financially supported by Ministry of
Education and Science of the Russian Federation
(project’s unique identifier RFMEFI60717X0176).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Flink</surname>
          </string-name>
          . https://flink.apache.org/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Kafka</surname>
          </string-name>
          . https://kafka.apache.org/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Kafka Streams Documentation</surname>
          </string-name>
          . https://kafka.apache.org/documentation/streams/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Samza</surname>
          </string-name>
          . http://samza.apache.org/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Storm</surname>
          </string-name>
          . https://storm.apache.org/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , Mithal, Vangala, Brugere, Boriah, Kumar:
          <article-title>A study of time series noise reduction techniques in the context of land cover change detection</article-title>
          .
          <source>NASA Conference on Intelligent Data Understanding</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Christ</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kempa-Liehrb</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feindt</surname>
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Distributed and Parallel Time Series Feature Extraction for Industrial Big Data Applications</article-title>
          .
          <source>ACML Workshop on Learning on Big Data</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Gitzel</surname>
            <given-names>R.</given-names>
          </string-name>
          :
          <source>Data Quality in Time Series Data An Experience Report. CBI Industrial Track</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Guidance for Data Quality Assessment: Practical Methods for Data Analysis</article-title>
          .
          <source>EPA</source>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Harvey</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Forecasting, structural time series models and the Kalman filter</article-title>
          . Cambridge university press (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>[11]Hive on Spark: Getting Started. https://cwiki.apache.org/confluence/display/Hive/H ive+on+Spark%3A+Getting+Started</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Huai</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chauhan</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gates</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hagleitner</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanson</surname>
            <given-names>E.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Malley</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pandey</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>R.</given-names>
          </string-name>
          , Zhang X.:
          <article-title>Major technical advancements in Apache Hive</article-title>
          .
          <source>ACM SIGMOD international conference on management of data</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Keogh</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pazzani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Derivative Dynamic Time Warping</article-title>
          .
          <source>First SIAM International Conference on Data Mining</source>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kovalev</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shanin</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stupnikov</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zakharov</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Data Mining Methods and Techniques for Fault Detection and Predictive Maintenance in Housing and Utility Infrastructure</article-title>
          . Engineering Technologies and Computer Science (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Köhler</surname>
            , Torsten,
            <given-names>Lorenz:</given-names>
          </string-name>
          <article-title>A comparison of denoising methods for one dimensional time series</article-title>
          .
          <source>Zentrum für Technomathematik</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Marron</surname>
            <given-names>J. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsybakov</surname>
            <given-names>A. B.</given-names>
          </string-name>
          :
          <article-title>Visual error criteria for qualitative smoothing</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Merkel</surname>
            <given-names>D.</given-names>
          </string-name>
          : Docker:
          <article-title>Lightweight Linux Containers for Consistent Development and Deployment</article-title>
          . Linux J., vol.
          <source>2014</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Mobley</surname>
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>An Introduction to Predictive Maintenance -</article-title>
          2nd
          <string-name>
            <surname>edition</surname>
          </string-name>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Moritz</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sardá</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bartz-Beielstein</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaefferer</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stork</surname>
            <given-names>J</given-names>
          </string-name>
          :
          <article-title>Comparison of different Methods for Univariate Time Series Imputation in R</article-title>
          .
          <source>CoRR abs/1510</source>
          .03924 (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Nayak</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Misra</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Behera</surname>
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>Impact of Data Normalization on Stock Index Forecasting</article-title>
          .
          <source>International Journal of Computer Information Systems and Industrial Management Applications</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>[21]OpenTSDB 2</source>
          .
          <article-title>3 documentation | HTTP API</article-title>
          . http://opentsdb.net/docs/build/html/api_http/put.ht ml
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Phan</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poisson</surname>
          </string-name>
          Caillault E.,
          <string-name>
            <surname>Lefebvre</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bigand</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Dynamic time warping-based imputation for univariate time series data</article-title>
          .
          <source>Pattern Recognition Letters</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <article-title>pySpark Package Documentation</article-title>
          . http://spark.apache.
          <source>org/docs/2</source>
          .1.0/api/python/pyspa rk.html
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Sakoe</surname>
          </string-name>
          ,
          <article-title>Chiba: Dynamic Programming Algorithm Optimization for Spoken Word Recognition</article-title>
          .
          <source>IEEE Transactions on Acoustics, Speech and Signal Processing</source>
          (
          <year>1978</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Sigoure</surname>
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>OpenTSDB: The distributed, scalable time series database</article-title>
          .
          <source>OSCON</source>
          , vol.
          <volume>11</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Song</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            <given-names>P.</given-names>
          </string-name>
          : SCREEN:
          <article-title>Stream Data Cleaning under Speed Constraints</article-title>
          .
          <source>ACM SIGMOD international conference on management of data</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <article-title>Spark Streaming Programming Guide</article-title>
          . https://spark.apache.org/docs/latest/streamingprogramming-guide.html
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Spark</surname>
            <given-names>Streaming</given-names>
          </string-name>
          +
          <article-title>Kafka Integration Guide</article-title>
          . https://spark.apache.
          <source>org/docs/2</source>
          .2.0/streamingkafka-integration.html
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>White</surname>
            <given-names>T.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Hadoop: The Definitive Guide. O'Reilly Media; Forth Edition</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Zhang</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Sequential Data Cleaning: A Statistical Approach</article-title>
          .
          <source>ACM SIGMOD international conference on management of data</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>