<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Energy Efficiency Platform Characterization for Heterogeneous Multicore Architectures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hergys Rexha</string-name>
          <email>hrexha@abo</email>
          <email>hrexha@abo.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Se´bastien Lafond</string-name>
          <email>slafond@abo</email>
          <email>slafond@abo.fi</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Science and Engineering, A ̊ bo Akademi University</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Science and Engineering, A ̊bo Akademi University</institution>
          ,
          <addr-line>Turku</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Runtime estimation of power dissipation and performance is crucial in every computing platform. In mobile systems, a special focus is set on energy efficiency in order to achieve the longest possible battery life and at the same time adhering to performance requirements. Powered by heterogeneous SoC's, mobile systems are called to reach an energy efficient state of execution, with a runtime system or scheduler that requires knowledge on the current performance and power dissipation. Today, highly heterogeneous architectures provide many actuators to reach better efficiency, the effect of which is usually unknown at runtime. In this paper, we propose a fast approach to build an energy efficiency model based on hardware performance counters. Our approach obviates the need for power sensors present at the chip level and deals with high numbers of execution modes. In building the energy efficiency model we account for the change in temperature which, as we show, has an impact on the optimal energy efficiency choice. The proposed approach reduces significantly the time to characterize the energy efficiency of a Multiprocessor System-on-Chip (MPSoC) and includes the environment temperature as a variable in determining the energy efficiency. Index Terms-MPSoC, energy efficiency models, platform configuration point, PMC, power models</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        The past years have seen rapid development in the amount
of data produced, processed and exchanged through
computing systems, ranging from high-end server farms to simple
household devices, and the trend of technology seems to fuel
even more this direction. Based on electricity usage ascribed
to Information and Communication Technology (ICT), it is
predicted that by the end of 2030 this sector will use as much
as 51% of global electricity production [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Following this
scenario, by the year 2030, the only ICT industry will be
responsible for up to 23% of the globally released greenhouse
gas emissions [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. A 2016 report [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] says that the US
datacenters held 350 million terabytes of data in 2015, and by
2020 they will require 100TWh of electricity to operate. This
is the equivalent of 7 nuclear power stations like Olkiluoto 3
in Finland. There is also an increase of datacenters capacity
in Europe, with London, Frankfurt, Paris, and Amsterdam
which grew their electricity consumption by 200MW in 2017.
Countries like Ireland and Denmark in Europe are becoming
a data base for the world’s biggest tech companies and by the
next 5 years promise to increase the power consumption by
1TW [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The emergence of the Internet of Things (IoT)
with devices operating at the edge of the network, poses
a new challenge to the Cloud to provide efficient service
provisioning. IoT devices are low powered devices and their
usage promises to decrease the overall power consumption
by increasing energy efficiency, but their number could be
overwhelming with the consequence of having a ”rebound
effect” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Cisco predicts that by the year 2020 in the world
will be 50 billion IoT devices, which is an order of magnitude
bigger than the number of smartphones and tablets working
today. So in this scenario, using the cloud services offered by
large datacenters to receive the data generated by IoT devices
will not be a sustainable solution in terms of cost, latency, and
environmental impact [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Recently the idea of edge devices
that provide the computation and storage closer to the source
of data has been formulated under the term of Edge or Fog
computing [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. As an edge device example, we can mention
smartphones, as intermediates between body sensors and the
cloud services, gateways as intermediates for smart homes, or
nano data centers that manage the caching or processing of
video contents. By using these edge devices in the proximity
of data sources, we could have as an end result in a reduction
of energy consumption w.r.t. implementing the logic in the
cloud, and at the same time keeping latency requirements of
certain applications [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        Therefore one key requirement of such computing
systems is undoubtedly energy efficiency. Basically, this means
that systems should minimize their energy consumption to
complete the required task and achieve a satisfying energy
proportionality [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. One of the largest consumers of energy
in computing environments is the CPU [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which requires
special attention especially in the multicore era. Today mobile
devices are using the same CPU as traditional gateways or
cloudlets in Edge Computing. The need to achieve energy
efficiency in today’s MPSoC is stringent, especially for mobile
devices that operate on battery, and that is a clear scenario
where the end user wants a better experience and longer
battery life.
      </p>
      <p>
        Workload variability makes the control of energy
expenditure especially difficult in mobile CPUs. Mobile devices
are not the only which require energy efficient solutions,
but also cloud providers need to lower the energy cost of
computations and cooling [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Today large scale computing
facilities are using energy as a resource to be scheduled and
charge according to the energy consumption [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
Heterogeneity shows a promise to increase the energy efficiency levels
achieved in MPSoC, hence several paths have been followed
by research and industry. For example, exploring heterogeneity
inside the CPU chip by using multiple technologies with
different power and performance characteristics or using cores
that alternatively behave as out-of-order computing elements
or as in-order cores [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Probably one of the most popular
and researched types of heterogeneity is the one provided by
different computing cores integrated into the same physical
chip. This type of heterogeneity is the one where computing
cores share the same Instruction Set Architecture (ISA) but
have different microarchitectures. However, an intelligent use
of these power and performance tradeoffs proves to be not
a simple challenge [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Being able to predict the optimal
choice between a number of hardware actuators such as the
number of cores, type of core and operating performance
point, or Dynamic Voltage and Frequency Scaling (DVFS), is
a difficult task that must be handled well in order to achieve
energy efficiency.
      </p>
      <p>
        With asymmetric multiprocessing (AMP) architecture there
is a better way to respond to the diversity of applications
present in the mobile environment. We have compute-intensive
applications which need to produce results in real time and
must use fast cores in order to meet the deadlines. On the
other side, background processes that may be memory bound
require little computation and are more suitable to run on
simple cores that achieve better levels of energy efficiency.
Even within a single application, we have different “windows
of activity” which may require varying levels of computing
intensity, e.g. reading, scrolling, responding through different
messages inside a social media application. Recently industry
has moved towards increasing the level of heterogeneity found
inside a single chip. From examples such as ARM big.LITTLE
with two types of cores, to Mediatek tri-cluster MPSoC [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
which promise to increase performance and reduce power
dissipation. DynamIQ from ARM [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] advances the concept
of big.LITTLE by providing better flexibility in the cluster
organization and frequency setting.
      </p>
      <p>
        High levels of heterogeneity present in recently
embedded architectures produce an increase in the design space
exploration to find an efficient use of platform actuators. By
increasing the number and type of cores and the number of
voltages and frequency levels for each computing element,
there is an increasing number of operating points on which
the platform may perform. In this scenario making the right
choice for execution could have a tremendous impact on
energy efficiency. Temperature also has a major effect on the
power dissipation of today’s systems [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which makes it an
important factor to account for in order to make the optimal
energy efficient choice.
      </p>
      <p>To manage efficiently the workload scenarios faced by
mobile devices, edge devices in IoT, or nano data centers,
there is a need to continuously monitor power data in order to
fB
fs fs
fs fs
Configurationx
fB fB
fB fB
Configurationt
fB fB
fB
Configurationy
fB fB fs fs</p>
      <p>fs fs
Configurationu
Configurationz
fs fs</p>
      <p>fs
fs fs
Configurationv
Big Cores</p>
      <p>Small Cores</p>
      <p>MPSoC
Fig. 1. Examples of possible platform configuration points in a multicore
architecture
choose the optimal power and performance trade-off.
Unfortunately, most of the hardware platforms today are not equipped
with power sensors, which significantly complicates
energyefficient management of the system settings.</p>
      <p>
        This paper follows our previous work which experimentally
builds an energy efficiency model based on platform
configuration points, for ARM big.LITTLE architecture [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. As
platform configuration point we denoted the set of platform
actuators such as number, type of core, core performance level
or DVFS and core utilization level. The model is derived by
testing all the possible configuration points of the platform.
Following the recent trend in platform complexity, this
approach is difficult to apply in the case of the combinatorial
explosion in the number of configuration points. The goal of this
paper is to explore new approaches in providing knowledge
of the platform energy efficiency to a runtime system based
on the concept of platform configuration points. We redefine
the set of parameters in the configuration point by removing
utilization level from the aforementioned description. Meaning
of the notion of platform configuration point is demonstrated
with several examples (from x to v) in a multicore platform
(Figure 1). In our energy efficiency model, we account for
the environment temperature variable, which provides valuable
information for the correct accounting of the CPU dissipated
power. Knowing the large impact that static power has on the
energy efficiency achieved in today’s CPUs the second purpose
of this work is to build thermally aware energy efficiency
models.
      </p>
      <p>The contributions of this paper are the following:
we propose an approach to characterize the energy
efficiency of a hardware platform based on the notion of
configuration points.
we include environment temperature in the energy
efficiency model and show the impact this variable has on
the relative efficiency of the points from the model.</p>
    </sec>
    <sec id="sec-2">
      <title>II. RELATED WORK</title>
      <p>less energy consumption.</p>
      <p>
        Exploring the usage of platform actuators for energy
management was studied by different research works. The authors
in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] all propose the creation of a runtime
system which is able to manage the scheduling and mapping
of threads dynamically with the objective of maximizing the
energy efficiency of MPSoC. In [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] a load balancer schedules
the workload in periodic time frames called epochs, wherein
each, a set of actions are performed to set the threads in
the appropriate core type. The platform considered is highly
heterogeneous with 4 types of core and in each epoch the load
balancer estimates the performance and power of every thread
in each core type. This information is used by the internal
algorithm to decide where to map the threads. Similarly,
in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is proposed a runtime scheme which is used to
schedule dynamically workloads in a MPSoC. The approach
is based on the sense-decide-act policy and operates on
an aggressive heterogeneous environment. It uses regression
models for estimating performance and power of threads in
different core type and also the contribution of a thread in
a total load of a core. An evolutionary algorithm is used
to decide in each term the scheduling of the threads. The
authors in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] propose a run-time task allocation approach
called SPARTA which categorizes task in computing bound or
memory bound and a heuristic that selects the configuration
that achieves the requested throughput with the minimal power
consumption. In these works is not considered the possibility
of DVFS as a mechanism to reduce power consumption and
also the hardware counters used for estimating performance
are not easily found in real hardware platforms. Sensors
for estimating the power consumption of different mapping
decisions are not available in many of today’s platforms.
Finding the optimal configuration for executing workloads in
a data-center in order to achieve better energy efficiency is
the goal presented in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Authors present a programming
and execution platform called Empya that uses hardware and
software techniques to determine the best trade-off between
performance and energy consumption. The run-time system
continuously monitors application performance and energy
consumption through Running Average Power Limit (RAPL)
registers. As actuators, the system operates on the number of
threads to use and the power cap on the CPU. In contrast with
this, our work focuses on heterogeneous platforms where for
achieving energy efficiency we use actuators such as number,
type of core and DVFS point. In [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] authors target again
High-Performance Computing applications running on a single
node with the goal of reducing the energy consumption by
choosing the right configuration, which is composed of the
number of cores and DVFS level. The work is based on
the application-agnostic power model and the performance
model of the application is obtained with a supervised learning
method of regression. Frequency, number of cores and input
size are used in the regression model. The methodology is
clear and straightforward, but there is no mention of the
performance requirement which is the value we trade off for
      </p>
    </sec>
    <sec id="sec-3">
      <title>III. CMOS POWER DISSIPATION</title>
      <p>CMOS technology has been mostly used in MPSoCs due
to the fact that has quite good noise immunity and low heat
production while the device is in operation mode. Power in
these circuits can be divided into two categories: dynamic
power and static power. Dynamic power is created by the
circuit activity (transistor switching) and is dependent on the
usage scenario, clock rates, and I/O activity. Switching power
is dissipated during the transistor changing from 0 to 1 and
vice versa, the dynamic power is defined as:</p>
      <p>Pdynamic =</p>
      <p>C</p>
      <p>VD2D
fclk
(1)
where C is the load capacitance, VDD is the source voltage,
is the activity factor and f is the operating frequency.
Static power is dissipated due to the leakage currents on
the transistors while they are in the “OFF” mode. The are
several sources of the leakage current which are strongly
influenced by the chip temperature. The dynamic part of the
power dissipated from the chip is modeled by two terms in
Equation 2, as a dynamic activity which relates to the active
running workloads and the background activity that represents
the system processes that run on the background. In Equation 3
the dynamic power is modeled by a single term due to the low
power dissipated by background processes in the A7 cluster.
Static power is modeled by the third term in Equation 2 and
is dependent on temperature and the supply voltage. For the
A7 cluster, there is no temperature sensor to monitor, hence
the static part is modeled together with the dynamic power
dissipation of background activity.</p>
    </sec>
    <sec id="sec-4">
      <title>IV. PROPOSED APPROACH</title>
      <p>Today embedded systems face a multitude of working
scenarios that range from burst in high performance requests,
to low power operation modes, going through the need to
provide sustainable performance in thermally constrained
situations. To do an efficient managing of such a number of use
cases the runtime scheduling manager need to have refreshed
information about the effect of changing different actuators
on the running applications. Thus there is a need for an
energy efficiency model which is based on the current runtime
power data. The envisioned system diagram is shown in Figure
2, where our work in this paper is focused in providing
the platform configuration points database for helping the
scheduler decisions in reaching the optimal efficiency level
of the running applications.</p>
      <p>
        The work in this paper is based on power models for
mobile CPUs based on hardware program counters (HPC). The
methodology for building such models is adopted from [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ],
which presents a statistical method for identifying and using
hardware counters. Their analyses propose the usage of
counters which show a high correlation to power and have also the
smallest multicollinearity. The authors in [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] show that this
brings high model stability with an average error of 3,8%.
      </p>
      <p>We start by building power models for two popular ARM
v7a architecture CPU’s, which are ARM Cortex-A7 and ARM
Cortex-A15. The micro-architecture limits the number of
events which can be sampled at once: 6 counters for A15
and 4 counters for A7 plus the cycle counter. The goal is
to search for those events which have the highest
correlation with power dissipation and at the same time show the
smallest intercorrelation with each other. To have high model
stability the predictors should be chosen to keep low levels
of multicollinearity in multivariate models. First, is measured
the correlation of all available events with the power, then
the counters are divided into clusters which include events
with high intercorrelation. Then, from each cluster is selected
the event which has more impact on the power dissipation
but keeping a low Variance Inflation Factor (VIF). The total
amount of events for the A7 is 40 and for the A15 in 120,
among these are selected 7 for the A15 and 5 for the A7.
The events used in the models are general and can be found
on most core types used in mobile systems. For each core
type, the events are listed on Table I. The power for A15
and A7 is divided in dynamic and static, plus the background
power which is related to the operating system activities.
(2)
(3)
The modelled formula for the power dissipation is showed
in Equation 2 and 3,</p>
      <p>{z
dynamic activity
nEnVD2Dfclk) + bVD2Dfclk + f (VDD; T )</p>
      <p>|BG d{yznamic} | st{aztic }
}
}
|</p>
      <p>{z
dynamic activity
nEnVD2Dfclk) + f (VDD; fclk)
static and {BzG dynamic
| }
where N is the number of events selected, n is the weight
given to certain event, En is the number of events per second
divided by the frequency (fclk) in MHz, VDD is the operating
voltage and T is the temperature of the core.</p>
      <p>The power model for the A15 has a thermal compensation
term for calculating the static power and background dissipated
power when the system is idling (Equation 2). In the power
model for A7 the static and background power are included in
the second term of Equation 3. This is related to the absence
of a thermal monitoring sensor in the A7 cluster. We have
calculated four sets of model coefficients for the parameters
in each cluster, representing the power with a different number
of cores for each CPU type. The model parameters for each
core type are given in Tables II and III. In the tables, it is
shown the event rate divided by the frequency in MHz, the
weight given to each coefficient and the statistical significance.
In some model terms, f and V are respectively the operating
frequency and voltage of each cluster (Table IV). The event
rates are divided by the operating frequency in order to avoid
correlation with it in the first term of power equations. The
power models need to be obtained only once by running on
the target platform a set of embedded representative workloads
which we call platform characterization set. After obtaining
the power model we compute the energy efficiency table
which provides a sort of database of all the possible platform
configuration points and the resulting performance, power
and energy efficiency values. By having this information the
runtime system is able to make decisions about the mapping of
a certain application with regard of the performance. If there
is a change in the environment temperature above a certain
threshold, then the power dissipation can be recomputed and
the table is redefined for the new thermal level.</p>
      <p>These models are build by running the characterization
workload set in each of the operating points of both CPUs.
The set contains workloads that test different levels of the
microarchitecture and memory subsystem. In part is composed
of real applications from the embedded domain, and for the
other part synthetic benchmarks designed to stress specific
parts of the CPU. Having the power models and by measuring
the performance in terms on instructions per second (IPS) we
can obtain an energy efficiency model of the platform. The
model is presented as a table that lists all the platform
configuration points with the energy efficiency levels achieved in
terms of instructions per Joule, performance point (instructions
per second) and the power dissipation (W). The table is used to
decide the optimal configuration point for an application that
has defined performance requirements. Once an application
is submitted into the system or is resumed by the scheduler,
the runtime system can sample the hardware counters in a
single frequency level and scans the table to find the optimal
configuration point, to run the application, in terms of energy
efficiency. In this work, we consider multi-threaded
applications, which matches our methodology of achieving optimal
levels of energy efficiency by using configuration points that
possibly use several cores. In the case where the performance
requirement of the application changes, the control logic of
the runtime system can select another configuration point that
provides the requested performance level and has a high level
of energy efficiency. When the temperature of the environment
changes above a certain threshold, the power model can be
used to recompute the energy efficiency table in accordance
with the new temperature conditions. A temperature increase
in the outside environment produces an increased level of static
power in the CPU, which affects the relative efficiencies of the
configurations inside the energy efficiency table. The runtime
system can continuously monitor the power usage of the
running application in order to not exceed the Thermal Design
Power (TDP) of the CPU. By sampling the performance
counters of each running application the power model shows
the power dissipation at runtime of the running applications,
thus the runtime system can make a decision of reducing the
power dissipation of certain applications by choosing another
configuration point from the system.</p>
      <p>The runtime system inputs temperature variations inside the
model and can recompute the energy efficiency table by taking
into account the new level of static power. The new table
needs to be searched for configuration points that satisfy the
performance request with the highest level of efficiency. A
basic schematic of the proposed approach is given in Figure 2.</p>
    </sec>
    <sec id="sec-5">
      <title>V. EXPERIMENTAL SETUP</title>
      <p>To evaluate our approach we used an ODROID XU3
development board from HARDKERNEL. The application
processor implements the ARM big.LITTLE architecture with
two clusters composed of 4 cores each. The big cluster consists
of a high-performance Cortex-A15 quad-core block, and a
low power Cortex-A7 quad-core CPU. The board description
is complete with a Mali-T628 GPU and 2GB LPDDR3 of
memory. The board contains 4 current sensors that offer
the possibility to measure power dissipation in four
different domains: big cluster (A15), LITTLE cluster (A7), GPU
and memory. Besides this, the board contains 4 temperature
sensors for the cores in the big cluster and one temperature
sensor for the GPU. The characteristics of the hardware can
be found in Table IV.</p>
      <p>
        To build the power model we used a set of benchmarks
from different application domains. We call the training set as
the platform characterization workloads. In the platform
characterization set we include a sequence of 76 workloads which
consists of a collection of synthetic and real world applications
from Roy Longbottom [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], PARSEC [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], CoremarkPro [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
ParMiBench [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and Multibench [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A full list of the used
workloads is in Table V.
      </p>
      <p>The choice of the workload set is based on the idea of
allinclusiveness of applications that characterize the embedded
systems domain.</p>
      <p>Experiments were conducted in different environments to
account for the outside temperature change in the SoC power
dissipation. The goal here is to evaluate the change in the
energy efficiency table in accordance with temperature. For the
first environment, the board fan is running with 100% speed
with the system located in a highly refrigerated environment.
In the second case, the board is working with the fan disabled
in a normal outside temperature to account for a high
temperature outside the environment. In the third case, the board
is working with the fan always on in a normal environment,
to justify the middle case. In Table, VI on Section V we will
show the result of the energy efficiency table computed in
different environments.</p>
    </sec>
    <sec id="sec-6">
      <title>VI. RESULTS</title>
      <p>
        By using the power and performance models defined
previously we are able to derive an energy efficiency model which
is based on platform configuration points. In Figure 3 we show
the efficiency of all configuration points from the model. Each
point describes a single configuration that provides a certain
level of performance in terms of instructions per second and
energy efficiency. By going towards high levels of performance
we notice a decrease in the density of the points. This means
that fewer options for achieving good energy efficiency levels.
The list of configurations is organized as an energy efficiency
table that lists all possible configuration points with their
efficiency and performance. An example of the table derived
from the workloads in the training set of the power model is
shown in Table VII. By searching inside the table we find
several sets of configuration points that provide the same
performance but with different energy efficiency levels, some
of the sets are shown in Figure 4. First usage of the table
would be the one for choosing the optimal configuration point
based on a certain requirement for the performance level. As it
is shown by Figure 4, it is possible to gain in terms of energy
efficiency if we make the right choice for the configuration
point. As a second objective of our work, we wanted to test
the effect of temperature on the relative energy efficiency of
configuration points in the model. For testing thermal effects
on the efficiency model, we choose to run a testing application
with the system located in different environments. We run
Basicmath application from the ParMiBench suite [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In
environment 1, the system running in a highly refrigerated
environment (we call it “cold” case). In Environment 2, the
system is running without a fan with an outside temperature
of 25 C (we call it “hot” case). Environment 3, consists of
the system running on a 25 C outside temperature with the
fan always on at 100% speed (we call it “middle” case).
We noticed the relative order of configuration points changes
between the environments and so does the energy efficiency
levels achieved.
      </p>
      <p>The top rows of the energy efficiency table for different
temperature environments are shown in Table VI. Different
temperature levels produce different order of configuration
points and efficiency levels achieved. This shows that there
is a need to change the platform configuration point when the
temperature changes significantly, in order to keep the high
levels of energy efficiency.</p>
      <p>In Figure 5 we show a possible runtime scenario. We are
running Basicmath test application with a required level of
performance such as e.g. 1,61E+9 inst/s in a system with
a temperature t1, according to the model the optimal
configuration point for this performance level is composed by
2a7@400MHz + 4a15@500MHz. In the case, the temperature
increases to t2, then the efficiency of that configuration point
decreases and thus we need to reconfigure with the new table
that shows that we should execute the application by using the
following configuration 4a7@700MHz + 4a15@200MHz.
Another example is shown with the performance requirement of
3,27E+9 inst/s, where again there is a need for reconfiguration
in order to keep high levels of energy efficiency.</p>
      <p>The change in the environment temperature of the system
(from “cold” to “hot”) produces large differences in the energy
efficiency levels that the model defines as an optimal
configuration point for the required performance. By looking at the
first 100 highly energy efficient configurations in the energy
efficiency table, we find few test cases, whereby changing the
configuration point when the system temperature changes the
gain in terms of energy efficiency is up to 33%. By searching
for new target reconfiguration points we account for the same
performance or 5% bigger. An interesting observation can be
noticed in Figure 3 where all points are plotted in the energy
efficiency and performance graph. If we take the points from
the upper outer layer of the scatter plot we have a situation
like in Figure 6. Those points show the configurations with the
optimal energy efficiency for a certain level of performance
at a defined temperature. Or otherwise, we can think of the
graph as the result of scanning the model from the lowest
performance point and keeping only those points which have
higher performance and the highest possible level of energy
efficiency. As a further validation of our approach, we measure
in percentage the difference between the predicted power
dissipation and the measured power in configuration points
with high levels of energy efficiency. The results are shown
in Figure 7, where we notice the highest error is 2,82%. We
measure the model errors in configurations that provide the
highest levels of energy efficiency for different performance
levels. These are more intriguing configuration points, which
give the best of the platform’s energy efficiency. Knowing that
most of the time these points will be used as configuration
options, having a low error rate from the model is very useful.</p>
    </sec>
    <sec id="sec-7">
      <title>VII. CONCLUSION</title>
      <p>In this work, we present an approach for building an energy
efficiency model which is based on platform configuration
points. The target of the approach are heterogeneous platforms
which are continuously increasing the depth of heterogeneity.
The model is based on hardware performance counters which
are widely available in today’s CPU architectures. The set
of workloads for building the model is representative of the
embedded domain which has shown to be more critical to the
energy efficient application execution. But also, the training
set, in inclusive of the IoT world. The novelty of this approach
compared to previous works is that it doesn’t necessarily
need power sensors for measuring the power dissipation in
each configuration point, but by sampling the counters on
one configuration point we can characterize the efficiency of
other configuration points. From all the points in the model,
we show that less than 1% of them (see points in Figure 7)
represent the highest levels of energy efficiency possible, in
all the performance spectrum offered by the platform. Also,
we include the environment temperature as a variable for
defining the need for application reconfiguration. As we show
by the tests if the temperature changes, by reconfiguring the
application execution we can gain up to 33% in terms of
energy efficiency.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] Arm DynamIQ Technology for the next era of compute - Processors blog - Processors - Arm Community</article-title>
          . https://community.arm.com/processors/b/blog/posts/arm
          <article-title>-dynamiqtechnology-for-the-next-era-of-compute.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>CPU</given-names>
            <surname>Benchmark - CoreMark-PRO - EEMBC Embedded Microprocessor Benchmark Consortium</surname>
          </string-name>
          . https://www.eembc.org/coremarkpro/index.php.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>EEMBC - MultiBench - Multicore Benchmark</surname>
          </string-name>
          . https://www.eembc.org/multibench/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Roy</given-names>
            <surname>Longbottom's PC Benchmark Collection - Free PC</surname>
          </string-name>
          Benchmarks. http://www.roylongbottom.org.uk/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Anders</surname>
            <given-names>S. G.</given-names>
          </string-name>
          <string-name>
            <surname>Andrae</surname>
            and
            <given-names>Tomas</given-names>
          </string-name>
          <string-name>
            <surname>Edler</surname>
          </string-name>
          .
          <article-title>On global electricity usage of communication technology: Trends to 2030</article-title>
          . Challenges,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>117</fpage>
          -
          <lpage>157</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashouri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Davidsson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Spalazzese</surname>
          </string-name>
          .
          <article-title>Cloud, edge, or both? towards decision support for designing iot applications</article-title>
          .
          <source>In 2018 Fifth International Conference on Internet of Things: Systems, Management and Security</source>
          , pages
          <fpage>155</fpage>
          -
          <lpage>162</lpage>
          ,
          <year>Oct 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bienia</surname>
          </string-name>
          .
          <article-title>Benchmarking Modern Multiprocessors</article-title>
          .
          <source>PhD thesis</source>
          , Princeton University,
          <year>January 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W. L.</given-names>
            <surname>Bircher</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. K.</given-names>
            <surname>John. Complete System</surname>
          </string-name>
          <article-title>Power Estimation Using Processor Performance Events</article-title>
          .
          <source>IEEE Transactions on Computers</source>
          ,
          <volume>61</volume>
          (
          <issue>4</issue>
          ):
          <fpage>563</fpage>
          -
          <lpage>577</lpage>
          ,
          <year>April 2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Peter</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Corcoran</surname>
          </string-name>
          .
          <article-title>Third time is the charm - why the world just might be ready for the internet of things this time around</article-title>
          .
          <source>CoRR, abs/1704.00384</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Bryan</surname>
            <given-names>Donyanavard</given-names>
          </string-name>
          , Tiago Mu¨ck, Santanu Sarma, and
          <string-name>
            <given-names>Nikil</given-names>
            <surname>Dutt</surname>
          </string-name>
          . SPARTA:
          <article-title>Runtime Task Allocation for Energy Efficient Heterogeneous Many-cores</article-title>
          .
          <source>In Proceedings of the Eleventh IEEE/ACM/IFIP International Conference on Hardware/Software Codesign and System Synthesis</source>
          ,
          <source>CODES '16</source>
          , pages
          <fpage>27</fpage>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>10</fpage>
          , New York, NY, USA,
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Eibel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Do</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Meissner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Distler</surname>
          </string-name>
          . Empya:
          <article-title>Saving Energy in the Face of Varying Workloads</article-title>
          .
          <source>In 2018 IEEE International Conference on Cloud Engineering (IC2E)</source>
          , pages
          <fpage>134</fpage>
          -
          <lpage>140</lpage>
          ,
          <year>April 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>EIRGRID</surname>
          </string-name>
          .
          <article-title>All-island generation capacity statement 2017-</article-title>
          <year>2026</year>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S. M. Z.</given-names>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Grahn. ParMiBench - An</surname>
          </string-name>
          Open-Source
          <source>Benchmark for Embedded Multiprocessor Systems. IEEE Computer Architecture Letters</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):
          <fpage>45</fpage>
          -
          <lpage>48</lpage>
          ,
          <year>February 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>V.</given-names>
            <surname>Jimenez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cazorla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gioiosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kursun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Isci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Buyuktosunoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bose</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Valero</surname>
          </string-name>
          .
          <article-title>Energy-Aware Accounting and Billing in Large-Scale Computing Facilities</article-title>
          .
          <source>IEEE Micro</source>
          ,
          <volume>31</volume>
          (
          <issue>3</issue>
          ):
          <fpage>60</fpage>
          -
          <lpage>71</lpage>
          , May
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Skadron</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Chung</surname>
          </string-name>
          . Predictive
          <string-name>
            <surname>Temperature-Aware</surname>
            <given-names>DVFS</given-names>
          </string-name>
          .
          <source>IEEE Transactions on Computers</source>
          ,
          <volume>59</volume>
          (
          <issue>1</issue>
          ):
          <fpage>127</fpage>
          -
          <lpage>133</lpage>
          ,
          <year>January 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tsai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gururajarao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lagerquist</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Son</surname>
          </string-name>
          , G. Gammie,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Thippana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          , U. Fu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dosluoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dia</surname>
          </string-name>
          , G. Gurumurthy,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hsieh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tzeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>U.</given-names>
            <surname>Ko</surname>
          </string-name>
          .
          <volume>3</volume>
          .
          <article-title>4 A 10nm FinFET 2.8GHz tri-gear deca-core CPU complex with optimized power-delivery network for mobile SoC performance</article-title>
          .
          <source>In 2017 IEEE International SolidState Circuits Conference (ISSCC)</source>
          , pages
          <fpage>56</fpage>
          -
          <lpage>57</lpage>
          ,
          <year>February 2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Jozef</surname>
            <given-names>Mocnej</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            <given-names>Misˇkuf</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peter Papcun</surname>
          </string-name>
          , and Iveta Zolotova´.
          <source>Impact of Edge Computing Paradigm on Energy Consumption in IoT. IFACPapersOnLine</source>
          ,
          <volume>51</volume>
          (
          <issue>6</issue>
          ):
          <fpage>162</fpage>
          -
          <lpage>167</lpage>
          ,
          <year>2018</year>
          .
          <source>15th IFAC Conference on Programmable Devices and Embedded Systems PDeS 2018 Citation Key: MOCNEJ2018162.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Tiago</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>¨ck, Santanu Sarma, and Nikil Dutt</article-title>
          .
          <article-title>Run-DMC: Runtime Dynamic Heterogeneous Multicore Performance and Power Estimation for Energy Efficiency</article-title>
          .
          <source>In Proceedings of the 10th International Conference on Hardware/Software Codesign and System Synthesis</source>
          ,
          <source>CODES '15</source>
          , pages
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          , Piscataway, NJ, USA,
          <year>2015</year>
          . IEEE Press.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>V.</given-names>
            <surname>Petrucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Laurenzano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Doherty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , D. Mosse´,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mars</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Tang.</surname>
          </string-name>
          Octopus-Man:
          <article-title>QoS-driven task management for heterogeneous multicores in warehouse-scale computers</article-title>
          .
          <source>In 2015 IEEE 21st International Symposium on High Performance Computer Architecture (HPCA)</source>
          , pages
          <fpage>246</fpage>
          -
          <lpage>258</lpage>
          ,
          <year>February 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ramapantulu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Loghin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y. M.</given-names>
            <surname>Teo</surname>
          </string-name>
          .
          <article-title>On Energy Proportionality and Time-Energy Performance of Heterogeneous Clusters</article-title>
          .
          <source>In 2016 IEEE International Conference on Cluster Computing (CLUSTER)</source>
          , pages
          <fpage>221</fpage>
          -
          <lpage>230</lpage>
          ,
          <year>September 2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Hergys</surname>
            <given-names>Rexha</given-names>
          </string-name>
          , Simon Holmbacka, and
          <string-name>
            <given-names>Sebastien</given-names>
            <surname>Lafond</surname>
          </string-name>
          .
          <article-title>Core Level Utilization for Achieving Energy Efficiency in Heterogeneous Systems</article-title>
          .
          <source>In 2017 25th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)</source>
          , pages
          <fpage>401</fpage>
          -
          <lpage>407</lpage>
          , St. Petersburg, Russia,
          <year>2017</year>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>V.</given-names>
            <surname>Saripalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mishra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Datta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          .
          <article-title>Exploiting Heterogeneity for Energy Efficiency in Chip Multiprocessors</article-title>
          .
          <source>IEEE Journal on Emerging and Selected Topics in Circuits and Systems</source>
          ,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>109</fpage>
          -
          <lpage>119</lpage>
          ,
          <year>June 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Santanu</given-names>
            <surname>Sarma</surname>
          </string-name>
          , T. Muck,
          <string-name>
            <surname>Luis A. D. Bathen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Dutt</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. Nicolau.</surname>
          </string-name>
          <article-title>SmartBalance: A Sensing-driven Linux Load Balancer for Energy Efficiency of Heterogeneous MPSoCs</article-title>
          .
          <source>In Proceedings of the 52Nd Annual Design Automation Conference</source>
          ,
          <source>DAC '15</source>
          , pages
          <fpage>109</fpage>
          :
          <fpage>1</fpage>
          -
          <lpage>109</lpage>
          :
          <fpage>6</fpage>
          , New York, NY, USA,
          <year>2015</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Arman</surname>
            <given-names>Shehabi</given-names>
          </string-name>
          , Sarah Josephine Smith,
          <string-name>
            <given-names>Dale A</given-names>
            . Sartor, Richard E. Brown, Magnus Herrlin, Jonathan G. Koomey,
            <surname>Eric R. Masanet</surname>
          </string-name>
          , Nathaniel Horner, Ineˆs Lima Azevedo, and
          <string-name>
            <given-names>William</given-names>
            <surname>Lintner</surname>
          </string-name>
          .
          <article-title>United states data center energy usage report</article-title>
          .
          <source>Technical report</source>
          , 06/
          <year>2016</year>
          2016.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>W.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <article-title>Edge computing: Vision and challenges</article-title>
          .
          <source>IEEE Internet of Things Journal</source>
          ,
          <volume>3</volume>
          (
          <issue>5</issue>
          ):
          <fpage>637</fpage>
          -
          <lpage>646</lpage>
          ,
          <year>Oct 2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Vitor</surname>
            <given-names>R. G.</given-names>
          </string-name>
          <string-name>
            <surname>Silva</surname>
            , Alex Furtunato, Kyriakos Georgiou, Kerstin Eder, and
            <given-names>Samuel</given-names>
          </string-name>
          <string-name>
            <surname>Xavier-de-Souza</surname>
          </string-name>
          .
          <article-title>Energy-Optimal Configurations for SingleNode HPC Applications</article-title>
          . arXiv:
          <year>1805</year>
          .00998 [cs], May
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>M. J. Walker</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Diestelhorst</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Hansson</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. K. Das</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>B. M.</given-names>
          </string-name>
          <string-name>
            <surname>Al-Hashimi</surname>
            , and
            <given-names>G. V.</given-names>
          </string-name>
          <string-name>
            <surname>Merrett</surname>
          </string-name>
          .
          <article-title>Accurate and Stable Run-Time Power Modeling for Mobile and Embedded CPUs</article-title>
          .
          <source>IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems</source>
          ,
          <volume>36</volume>
          (
          <issue>1</issue>
          ):
          <fpage>106</fpage>
          -
          <lpage>119</lpage>
          ,
          <year>January 2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>