<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Management System for Energy Analytics and its Application to Forecasting</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Bei Chen IBM Research</institution>
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Francesco Fusco IBM Research</institution>
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Jean-Baptiste Fiot IBM Research Ireland jean-</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Mathieu Sinn IBM Research</institution>
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Vincent Lonij IBM Research</institution>
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The e ective management of a power grid with an increasing share of (distributed) renewables and more and more available data, e.g., coming from smart meters, heavily relies on advanced data analytics such as demand and supply forecasting. In this context, data management is one major challenge in electric grids. Large amount of data from multiple heterogeneous sources require transformations, e.g., spatio-temporal alignment or anomaly detection, to serve data analytics tasks and are often applied on di erent views of the data, e.g., on state, substation or feeder level. In this paper, the progress on the development of an energy data management systems for the electricity grid is presented. The design of the system was inspired by the realworld use case of forecasting short-term energy demand in Vermont, using data from a combination of SCADA, smart meters and weather forecasting services. A general data model addressing the aforementioned challenges and aimed at supporting advanced data analytics is introduced. The proposed data model views a time series as an abstract concept that might represent raw measurements or arbitrary operations. The bene ts of the system is demonstrated for the design and live update energy demand forecasts.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The smart grid is the next-generation power system
characterized by the inclusion of highly-distributed intelligent
devices and communication technologies that manage
electricity demand in a sustainable, reliable and economic
manner. Advanced data analytics, such as load classi cation or
forecasting, are essential operations for the optimization of
the power ow in a smart grid. For the successful
application of such operations the reliable processing and
management of the data collected by a smart grid is a key factor.
Due to the increasing number of diverse devices
considerable amount of data is produced by a smart grid, leading to
a trend of emerging big data architectures and the discussion
of technical challenges as well as potential use cases [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>
        Data management is one major challenge in electric grids
as data is incomplete in nature, heterogeneous, di cult to
merge and arrives at di erent rates [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Energy data
arrives from various distributed sources, e.g., supervisory
control and data acquisition (SCADA) systems, smart meters,
renewable generation systems, and di ers signi cantly in
terms of format, resolution and quality. Moreover, data
might represent di erent views of a power system, e.g., load
at feeder level or over a whole substation. Analytics tasks,
such as demand forecasting, also require contextualisation
with additional data sets, such as calendar information and
weather forecasts. The latter might come from di erent
weather services, again arriving at di erent rates and in
different time and location resolutions. The system needs to
be able to consolidate all these diverse data sets and
provide a common view. Analytics are then rarely performed
on raw data, but require transformations of the data such
as the computation of weighted averages or anomaly
detection. The system has to support, store and continuously
re-compute such transformation so that analytics can be
continuously applied.
      </p>
      <p>In this paper, we explicitly address the challenge of data
management in a smart grid, proposing a data management
architecture for the smart grid that is, on the one side, able
to manage such diverse data sets and, on the other side,
supports operations, such as forecasting, that perform various
transformation on the available data. To achieve this, we
introduce a generic data model for the energy domain and
show how this data model can be applied for the use case of
short term energy demand forecasting.</p>
      <p>
        There have been other e orts in the design of smart grid
architectures. For example, Yang et. al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] give a high-level
overview for a smart grid big data management system,
introducing a simple data model, but mainly discussing data
      </p>
      <p>Weather
data
Smart
Meters
SCADA,
ICCP,MV90
FFeeaatutureresseelelecctitoionn
MMooddeelltrtarainininingg
RReefifnineemmeenntsts
MMeetrtircicss&amp;&amp;KKPPIsIs</p>
      <p>MMooddeellssccoorirningg
OOnn-l-ilnineeleleaarnrniningg</p>
      <p>AAnnoommaalylyflfalaggss</p>
      <p>InIntteerrffaacceeAAPPIsIs
SSQQLL</p>
      <p>SSppaarrkk,,HHBBAASSEE
DB2
-Weather feat.
-Power data
-Metadata
-Models
-smaHrtDmFeSters
Data storage</p>
      <p>NetCDF
-Gridded
weather data</p>
      <p>
        External
applications
WWeebbPPoorrttaall
RREESSTTAAPPIsIs
MMaarrkkeett
bbididddiningg
....
distribution and load balancing. Using cloud computing as
a platform for smart grid data management [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
providing time series analytics as a service [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] poses another
trend in this area. A cloud-based demand response
platform is introduced in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which provides a work ow for data
ingestion and scalable demand forecasting on Hadoop. To
handle large amount of data and enable real-time reactions
data streaming techniques have also been applied for the
smart grid context [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. A real-time data management
system was developed within the MIRABEL project with focus
on storing and processing special energy planning objects for
demand-response [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In the speci c context of modeling
energy data, the representation of time series and ex-o ers
within the context of distributed energy markets has been
studied in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], while an ontology for big data in the smart
grids has been considered in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], with particular focus on
o shore wind farms. The proposed data model speci cally
addresses the management of dynamic time series data, as
well as the operations and analytics applied to them. It can
therefore be seen as complementary to the reviewed research
e orts and might serve as a basis for other existing designs
of smart grid architectures.
      </p>
      <p>In the remainder of the paper a general system
architecture supporting data management and analytics tasks for
energy utilities is introduced in section 2. The paper then
focuses on the data model and query operations at the core
of the proposed system in section 3. To demonstrate the
usability of our system, in section 4, the use case of short-term
energy demand forecasting is discussed and concrete
examples on the stored data and query operations are given. Final
conclusions and future possibilities are discussed in section
5.</p>
    </sec>
    <sec id="sec-2">
      <title>2. SYSTEM ARCHITECTURE</title>
      <p>A data management system was developed, conceptually
shown in Fig. 1, in order to support general data analytics
and services for energy utilities by overcoming the challenges
involved in dealing with the ingestion of data from multiple,
heterogeneous sources.</p>
      <p>
        The system was developed to support the speci c use case
of short-term energy forecasting for a number of distribution
utilities in Vermont, United States. In particular, the
systems provides predictions of hourly energy consumption and
distributed solar photovoltaic (PV) generation up to 2 days
ahead at various aggregation levels. Time series of energy
demand are derived from a combination of thousands of
active power measurement points available from SCADA and
interval energy data from MV-90. Where required, energy
demand or distributed generation is obtained by aggregating
data from AMI up to the desired level of the grid (feeder,
substation). IBM Deep Thunder [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] was used to obtain
weather predictions up to 72 hours ahead. Deep Thunder
produces weather forecasts on a grid with 1 km resolution
and 10 minute time steps. Forecasts are updated twice per
day and each run produces about 300 gigabytes of data.
Internally, the data are stored in a combination of relational
database management systems (RDBMS) for electrical asset
data and SCADA/MV-90, Hadoop distributed le system
(HDFS) for the AMI data, network common data format
(NetCDF) for the weather gridded data.
      </p>
      <p>As shown in Fig. 1, the data management system supports
the tasks of: data ingestion, curation and spatio-temporal
alignment of energy and weather data; training forecasting
models, which requires retrieving raw data and designing
input features (covariates); retrieving data of covariates and
scoring forecasting models at runtime; interfaces to client
applications (e.g. web portal, market bidding services).
3.</p>
    </sec>
    <sec id="sec-3">
      <title>THE DATA MODEL</title>
      <p>In order to manage highly heterogeneous sets of data and
support a variety of analytical operations typical of energy
and utilities, as the ones detailed in section 2, a data model
was developed. The main objective of the data model was
providing a transparent, high-level interface to the users and
client applications, as well maintaining consistency, integrity
and traceability between the various data sources.</p>
      <p>Figure 2 shows a conceptual structure of the proposed
data model. All dynamic data in an electricity grid can be
represented as a time series. Analogue, operations on the
data applied by analytics also produce time series, such as
lags or rolling-window forecasts. Consequently, at the core of
the proposed data model is the representation of time series,
including timestamp/value pairs and abstract operations, as
described in section 3.1. The data model also contextualises
the time series with respect to a physical quantity (e.g.
energy demand, power, temperature) and entity (e.g.
substation, service territory), through the concept of signals, as
discussed in section 3.2. Finally, the representation of
analytical models, which links the output time series of complex
operations, such as energy forecasts, to one or more several
input time series, is then detailed in section 3.3.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Time series</title>
      <p>At the core of the data model is the abstract concept of
a TimeSeries, which in general can represent materialized
data or abstract operations.</p>
      <p>Some examples of speci c time series which can be used
to represent raw observation data are
TimeSeriesUnstructured, with values of type double, and
TimeSeriesCategoricalIndexed, with values coming from a set of labels
represented as a CategoricalIndex. The values of a TimeSeries
are represented through the concept of
TimeSeriesMaterialization entity, which points to a TimeSeriesStore where
the unique pairs timestamp and values,
TimeSeriesMaterializedValues, are stored. Note that the
TimeSeriesStore could be one or more tables in the database itself but
could also be a di erent system, for example an Hadoop
Distributed File System (HDFS) in the case of very large data
sets such as the smart meter data.</p>
      <p>Typically the raw time series data are not immediately
applicable for analytics purposes and several operations are
required for cleaning and spatio-temporal alignment. In the
proposed data model, the concept of TimeSeries is also used
to support and trace such operations, which can be
calculated on the y where possible or in batch processes using
the materialization concept. Some examples are:
TimeSeriesWeightedSum: Time series whose value xt,
at time t, is the weighted sum of the value of n input
txitm=e sejnr=ie1swyjtjy,tj.j T=yp1;ica::l :unse, acatstehseofsawmeieghttimede,sunma msaerlye
spatial aggregations (interpolation of high-resolution
weather data, extraction of PV generation in a spatial
region, etc.) or applications of power ow equations to
derive the electrical load at a substation.</p>
      <p>TimeSeriesLagged: Time series whose value xt, at
time t is given as the value of a reference time series
yt at time t h, namely xt = yt h. Lagged time series
are critical for representing delayed dynamical e ects
in statistical models. Note that the combination of
lags and weighted sums can be used to represent quite
complex time series models.</p>
      <p>TimeSeriesWindowed: Represents a time series with
values xt, at time t, resulting from an operation
applied to the values of an input time series y at times
within a window 2 [t ; t]. Use cases for such an
operator are time interpolation and integration, when
raw data come at irregular sampling intervals or
moving averaging is required (e.g. SCADA gives
instantaneous power but we are interested in modelling
energy). An equally important use case is the
computation of summary statistics of high-resolution data, such
as daily minimum/maximum of temperature or energy
consumption, which can be quite useful in build
forecasting models, as demonstrated in section 4.</p>
      <sec id="sec-4-1">
        <title>Another important type of TimeSeries is the</title>
        <p>TimeSeriesPriority, internally composed of a sorted
map of TimeSeries indexed according to a priority
order. For a given timestamp, by default, the
TimeSeriesPriority returns the value from the rst
TimeSeries in the map where a value is available.
Alternative behaviors can be implemented, where the value
from the TimeSeries at a speci c index is returned, or
from the rst TimeSeries starting at the index above
a certain threshold.</p>
        <p>The main use case of the TimeSeriesPriority is
dealing with rolling-horizon forecasts, which was the
primary application behind the development of the
proposed data model, as mentioned in section 2.
Rollingwindow forecasts, e.g. from weather forecasts
produced by numerical models, are multi-step ahead
forecasts that are refreshed at regular intervals. Such
quantities cannot be represented as a one-dimensional time
series because there is a one-to-many relation between
timestamps and values. An option could be to
overwrite the values with the latest available forecasts, at
the loss of potentially valuable information (most
recent forecasts are not necessarily more accurate) and
traceability of operations between live and batch
calculations. By using the TimeSeriesPriority,
rollingwindow forecasts are represented as multiple
TimeSeries indexed by a quantity proportional to the
forecasting horizon h, speci cally where each TimeSeries
is x^(t + hjt), with h xed. By default, the most
recent forecast is returned, but one could select the value
from forecasts at least 24-hour ahead, or many other
alternative behaviors could be implemented.</p>
        <p>An alternative use case of the TimeSeriesPriority is
the fallback mechanism between multiple forecasting
models applied to the same energy signal. If the
models are prioritised based on some accuracy measure, the
TimeSeriesPriority allows to deal with temporary
issues in one model (e.g. data anomalies or missing
inputs) and to transparently fall back to the next
available model output. Such mechanism will be
demonstrated in section 4.</p>
        <p>Finally, the TimeSeriesFlag is also de ned, as a
mechanism to associate ags to particular data points of a time
series, in order to deal with anomalies or data quality issues.
Flagging time series points prevents them from being used,
for example, as input to analytical models and increase the
robustness of the live system, for example in the case of
faulty autoregressive model features.</p>
        <p>The proposed data model is quite general and can be
easily extended with other fundamental types of TimeSeries.
The listed examples already form the basis for quite a rich
grammar able to drive powerful features from the raw data.
3.2</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Signals and Entities</title>
      <p>One of the main objective of the proposed data model is
to provide context to the existing data with respect to
highlevel physical entities and types of quantities of interest in
the speci c domain of application. As a result,as also shown
in Fig. 2, two main dimensions are utilised in the system for
identifying one or more TimeSeries:</p>
      <p>Entity: represents a physical entity of interest. In the
context of energy utilities, for example, we have
administrative or geographical entities such as State,
DistributionUtility, County, Town, and grid assets such
as DistributionSubstation, DistributionFeeder,
NetworkBus, NetworkBranch, ServicePoint.</p>
      <p>SignalType: represents the type of a physical
quantity for which it can be expected to have observation
data or for which estimated time series data are
expected to be required. Some examples are
TEMPERATURE, ENERGY_DEMAND, ENERGY_RESIDUAL_DEMAND,
ENERGY_GENERATION_PV, ACTIVE_POWER. Signal types are
the mechanism for cataloguing time series according to
some high-level human-understandable meaning, and
to maintain consistency between data of the same type,
for example with respect to UnitMeasure.</p>
      <p>The two dimensions of Entity and SignalType are
gathered within the concept of Signal, which is a required
property of a TimeSeries. The proposed constructs allow the
data and more abstract time series available within the
system to be navigated with very high-level queries of the type:</p>
      <sec id="sec-5-1">
        <title>SELECT FROM TIME SERIES TS</title>
        <p>INNER JOIN SIGNALS S ON S . ID=TS . SIGNAL
INNER JOIN SIGNAL TYPES ST ON ST . ID=S . STYPE
INNER JOIN ENTITIES E ON E . ID=S . ENTITY
WHERE E .NAME= ' S u b s t a t i o n n a m e '</p>
        <p>AND ST .NAME= 'ENERGY RESIDUAL DEMAND ' .
3.3</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Models</title>
      <p>Another important component of the proposed data model
was designed in order to represent the output of statistical
models, which relies on yet another type of time series, the
TimeSeriesModelled. The details of the model are
represented through the following entities:
ModelClass: Speci es the structure of the model, in
terms of requires set of inputs, as pairs of SignalType
and variable name, and the SignalType of the output.
Model: A realization of a ModelClass, with speci c
values for the parameters and trained to model a speci c
Signal (pair SignalType/Entity). The parameters
are stored using an XML le following the principles
of the Predictive Model Markup Language1 (PMML).
ModelInstance: An instantiation of a Model, where
each required input, a CovariateInstance, is linked
to a speci c TimeSeries.</p>
      <p>Such a representation of the analytical models is powerful
in supporting the o ine tasks of model design and training
of the data scientist: the relevant data can be transparently
extracted and aligned by relying on queries of the type given
in section 3.2; derived features can be designed by using
the abstract fundamental operation described in section 3.1.
Moreover, when dealing with thousands of statistical
models, the system makes it quite easy to navigate between
models to verify performance and retrain them where needed.</p>
    </sec>
    <sec id="sec-7">
      <title>4. SYSTEM DEMONSTRATION</title>
      <p>In this section, the application of our system to short-term
forecasting of electricity demand is demonstrated. Due to
con dentiality reasons, electricity demand data of Vermont
from January 1st, 2012, till August 31st, 2015, was obtained
from the website of ISO New England2. Those data came in
hourly format, with the measurements describing energy
usage (in MWh) over the previous hour. When ingesting those
data into our database, the time stamps were converted to
Eastern Standard Time (EST). Three classes of covariates
were used in the forecasting model: weather data, calendar
variables, and lagged demand values. Next, the work ow
that was used for designing and training forecasting models
is described, and the con guration of the system to apply
forecasting models in a \live" environment is demonstrated.
4.1</p>
    </sec>
    <sec id="sec-8">
      <title>Modelling</title>
      <p>
        For the modeling and forecasting of electricity demand,
a popular class of non-linear regression models, which
represent the e ect of covariates in an additive fashion was used:
Generalized Additive Models (GAMs). For more background
on GAMs and their application to electricity demand data,
we refer to [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Note that, in principal, the proposed data
model would support any other class of regression or classi
cation methods, as it only describes the inputs and outputs
of the models, but not the exact form of the functional
relation. The mgcv package in R was used for training GAMs
(see [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]), as part of the following general work- ow:
1. A training data frame is compiled by querying
historical electricity demand data and the covariates aligned
with it. For the Vermont data, the choice of the
covariates had already been de ned and implemented
through abstract time series in the data model. Since
the relation between the model and covariates is also
represented in our data model, the compilation of the
training data frame could be done fully automatically.
1http://dmg.org/
2http://iso-ne.com
Signal 6792, ENERGY DEMAND MEAN, ISO-NE Vermont
iterative fashion - new model covariates are derived. Some
of the key operations supported by our data model are, e.g.,
Ts 43775, TS CATEGORICAL UNSTRUCTURED, hour
Ts 43774, TS CATEGORICAL INDEXED, dayType
Ts 43777, TS CATEGORICAL UNSTRUCTURED, timeOfYear
Ts 43776, TS CATEGORICAL INDEXED, season
Ts 46463, TS PRIORITY, dryBulbTemperature
      </p>
      <p>Ts 46464, TS MEASURED, dryBulbTemperature h = 1hr
Ts 46464, TS MEASURED, dryBulbTemperature h = 2hr
. . . . . . . . . . . .</p>
      <p>Ts 46464, TS MEASURED, dryBulbTemperature h = 72hr
Ts 46063, TS PRIORITY, irradiance</p>
      <p>Ts 46064, TS MEASURED, irradiance h = 1hr
. . . . . . . . . . . .
+ Ts 46159, TS LAGGED, dryBulbTemperature.lag1
. . . . . . . . . . . .</p>
      <p>Ts 361410, TS LAGGED, DEMAND.lag.24</p>
      <p>
        Ts 360977, TS MEASURED, Vermont ISO-NE demand
+ Ts 361411, TS LAGGED, DEMAND.lag.36
+ Ts 361412, TS LAGGED, DEMAND.lag.48
+ Ts 361413, TS MODELLED, Output of S ISO NE Vermont mean demand A1 v0.pmml
2. Two models were produced: a model using
autoregressive features (S_ISO_NE_Vermont_mean_demand_A5_v0
in Fig. 3), speci cally the demand at various lags;
a more robust model without autoregressive features
(S_ISO_NE_Vermont_mean_demand_A1_v0 in Fig. 3),
which could be used in case of data anomalies. The two
models are represented through a
TimeSeriesPriority, as described in section 3.1, which allows the
implementation of a fallback mechanism where the preferred
model (the one with autoregressive features) fails to
produce a value because of data anomalies.
3. The models were trained on a speci ed period of time,
in-sample and out-of-sample statistics were calculated
to assess its accuracy, and nally exported into PMML
format, required for model scoring in the live system
described in the following section.
4. Besides the GAM models for the conditional mean,
which served as forecast of electricity demand, a model
for the conditional variance was also trained and
exported, which served as forecast of the associated
uncertainty. Using the methodology in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], this model
was obtained by tting a GAM to the squared model
residuals in the training data.
      </p>
      <p>Figure 3 shows a conceptual structure of the forecasting
models for the mean of energy residual demand, and their
links to the input covariates. The one short-term energy
forecast time series that a user would see as an output is
internally represented by two statistical models applied to
around 500 time series each. Note how each input coming
from a Deep Thunder forecast is internally represented as a
TimeSeriesPriority composed of 72 time series.</p>
      <p>Note that, in cases where the user wants to design new
forecasting model from scratch, the work- ow is more
involved. Typically, the user would start by querying
historical energy data and raw inputs from which - often in an
4.2</p>
    </sec>
    <sec id="sec-9">
      <title>Live system operations</title>
      <p>In the live system context, the following work- ow was
used for applying the forecasting models to live data:
1. Adapters for automatically ingesting weather forecasts
from IBM Deep Thunder, extracting spatio-temporal
weather features as described in Section 2, and
inserting them into the database were developed.
2. Adapters for automatically ingesting live data feeds
from SCADA, MV90 and AMI were also developed.
Besides being able to display to the user the latest
actual measurements, this also helps improving
forecasting accuracy, e.g., by using the current electricity
demand for predicting the demand in 24 hours from
now. Data from SCADA systems is available typically
with very low latency (seconds or few minutes); in the
case of Vermont, the data from ISO New England
becomes available only after a couple of days. However,
the scenario where data would be available in real-time
was emulated and a 12-,24-,36-hours lagged demand
variable were included in the forecasting models. In
order to avoid that anomalous values distorts the
output of forecasting models, a lter was implemented
for agging such values, such that the system would
avoid scoring the corresponding models. In this case,
through the fallback mechanism based on the
TimeSeriesPriority, the system would fall back to the
model without lagged variables.
3. Upon the availability of new weather forecasts, a list of
timestamps over the 72-hour forecasting horizon was
given as input to the model scoring engine of the
system. For the implementation of the engine, IBM
InfoSphere Streams was used. Basically - for all models
registered in the database - the required covariates for
the given timestamps are retrieved, the GAM model
speci ed in PMML format is applied, and the
forecasts are written back to the database. If covariates
are missing for a particular model and timestamp, then
a log message is generated and no forecast is produced.</p>
      <p>Figure 4(a) shows the forecasts for August 27th-29th, 2015,
based on models that were trained with data from January
1st, 2012 till July 31st, 2015. The graph also displays
uncertainty bands obtained from the conditional variance
forecasts. Note that the forecasts of the conditional mean are
based on a model which uses 12-,24-,36-hours lagged
demand values. The graph also shows, in a dotted line the
less recent forecasts, &gt; 24-hours ahead, which are also
produced by the system and can easily be retrieved using the
estimation horizon and the concept of TimeSeriesPriority.
Figure 4(b) illustrates the fall-back mechanism in the case
of missing inputs: the solid line shows the forecasts from a
model with 24-hours lagged demand values; the dashed line
corresponds to the forecasts from a fall-back model without
lagged values. In the case where real-time demand
information is missing or anomalous (and hence agged at data
ingestion data), our system would automatically return the
output of the fall-back model, rather than not providing any
forecasts for those instances at all. If real-time information
is available, it will return the forecasts from the model with
lagged demand values, which are more accurate in general.</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION AND FUTURE WORK</title>
      <p>A data and analytics management system for energy
utilities was described. In particular, focus was put on the design
of the core data model required for providing a transparent,
high-level interface to the users of the data and the client
applications. The data model was also designed for
maintaining consistency, integrity and traceability between the
various complex data sources relevant to energy utilities,
particularly energy and weather data.</p>
      <p>An implementation of the proposed data model was
demonstrated in the context of a short-term energy forecasting
system, particularly in support of the model
training/deployment tasks and in the system live operations. Further
applications and extensions can be considered along the direction
of analytical model management, automation of model
(re)training and support for model feature design. Further
research will also go in the direction of a more formal study of
the time series grammar and its potential in support many
other use cases. The Big Data aspect of the data was not
discussed, but the de nition of a hybrid architecture where
data are stored in a mix between traditional RDBMS and
HDFS is also scope for further study.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Aiello</surname>
          </string-name>
          and
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Pagani</surname>
          </string-name>
          , \
          <article-title>The Smart Grid's Data Generating Potentials,"</article-title>
          <source>Proceedings of the Federated Conference on Computer Science and Information Systems</source>
          , vol.
          <volume>2</volume>
          , pp.
          <volume>9</volume>
          {
          <issue>16</issue>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Diamantoulakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. M.</given-names>
            <surname>Kapinas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Karagiannidis</surname>
          </string-name>
          , \
          <article-title>Big Data Analytics for Dynamic Energy Management in Smart Grids,"</article-title>
          <source>Big Data Research</source>
          , vol.
          <volume>2</volume>
          , no.
          <issue>3</issue>
          , pp.
          <volume>94</volume>
          {
          <issue>101</issue>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , R. Sherick,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ieee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Loparo</surname>
          </string-name>
          , \
          <article-title>Big Data Analytics in Power Distribution Systems,"</article-title>
          <source>in Proceedings of the Innovative Smart Grid Technologies Conference</source>
          , Washington, DC, USA,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Ma</surname>
          </string-name>
          , P. X. Cheng, and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          , \
          <article-title>The Design and Implementation of Smart Grid High Volume Data Management Platform Architecture,"</article-title>
          <source>in Proc. of the Innovative Smart Grid Technologies Conf</source>
          ., Washington, DC, USA,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rusitschka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eger</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Gerdes</surname>
          </string-name>
          , \
          <article-title>Smart Grid Data Cloud: A Model for Utilizing Cloud Computing in the Smart Grid Domain,"</article-title>
          <source>First IEEE Int. Conf. on Smart Grid Communications</source>
          , pp.
          <volume>483</volume>
          {
          <issue>488</issue>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y. C.</given-names>
            <surname>Xiaomin Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sheng</given-names>
            <surname>Huang</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Browny</surname>
          </string-name>
          , \
          <article-title>TSAaaS: Time Series Analytics as a Service on IoT,"</article-title>
          <source>in Proc. of the IEEE Int. Conf. on Web Services (ICWS)</source>
          , Alaska, USA,
          <year>2014</year>
          , pp.
          <volume>249</volume>
          {
          <fpage>256</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Simmhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Aman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Kumbhare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stevens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V. K.</given-names>
            <surname>Prasanna</surname>
          </string-name>
          , \
          <article-title>Cloud-Based Software Platform for Big Data Analytics in Smart Grids,"</article-title>
          <source>Computing in Science and Engineering</source>
          , vol.
          <volume>15</volume>
          , no.
          <issue>4</issue>
          , pp.
          <volume>38</volume>
          {
          <issue>47</issue>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Couceiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferrando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Manzano</surname>
          </string-name>
          , and L. Lafuente, \
          <article-title>Stream analytics for utilities. Predicting power supply and demand in a smart grid,"</article-title>
          <source>Proceedings of the 3rd International Workshop on Cognitive Information Processing (CIP)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>U.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kaulakiene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Khalefa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lehner</surname>
          </string-name>
          , T. B.
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Siksnys</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Thomsen</surname>
          </string-name>
          , \
          <article-title>Real-Time Business Intelligence in the MIRABEL Smart Grid System," in Enabling Real-Time Business Intelligence, ser</article-title>
          .
          <source>Lecture Notes in Business Information Processing</source>
          . Springer Berlin Heidelberg,
          <year>2013</year>
          , vol.
          <volume>154</volume>
          , pp.
          <volume>1</volume>
          {
          <fpage>22</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Siksnys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Thomsen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T. B.</given-names>
            <surname>Pedersen</surname>
          </string-name>
          , \MIRABEL DW:
          <article-title>Managing Complex Energy Data in a Smart Grid," in Data Warehousing and Knowledge Discovery, ser</article-title>
          .
          <source>Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg,
          <year>2012</year>
          , vol.
          <volume>7448</volume>
          , pp.
          <volume>443</volume>
          {
          <fpage>457</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nunavath</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Prinz</surname>
          </string-name>
          , \
          <article-title>Big Data Metadata Management in Smart Grids," in Big Data and Internet of Things: A Roadmap for Smart Environments, ser</article-title>
          .
          <source>Studies in Computational Intelligence</source>
          . Springer International Publishing,
          <year>2014</year>
          , vol.
          <volume>546</volume>
          , pp.
          <volume>189</volume>
          {
          <fpage>214</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Treinish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Praino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Novakovskaia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Drexel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Derech</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Hertell</surname>
          </string-name>
          , \
          <article-title>Operational Evaluation of a Meso-Scale Weather and Outage Prediction Service for Electric Utility Operations,"</article-title>
          <source>in First Conference on Weather, Climate, and the New Energy Economy</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sinn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Goude</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pompey</surname>
          </string-name>
          , \
          <article-title>Adaptive Learning of Smoothing Functions : Application to Electricity Load Forecasting,"</article-title>
          <source>in Proceedings of the Neural Information Processing Systems (NIPS) Conference</source>
          , Lake Tahoe, Nevada, USA,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wood</surname>
          </string-name>
          , Generalized Additive Models:
          <article-title>An Introduction with R</article-title>
          . CRC Press,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>T. K. Wijaya</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sinn</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , \
          <article-title>Forecasting Uncertainty in Electricity Demand,"</article-title>
          <source>AAAI-15 Workshop on Computational Sustainability</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>