<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluating espresso cofee quality by means of time-series feature engineering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniele Apiletti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eliana Pastor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riccardo Callà</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Baralis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Control and Computer Engineering</institution>
          ,
          <addr-line>Politecnico di Torino</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Espresso quality attracts the interest of many stakeholders: from consumers to local business activities, from cofee-machine vendors to international cofee industries. So far, it has been mostly addressed by means of human experts, electronic noses, and chemical approaches. The current work, instead, proposes a datadriven analysis exploiting time-series feature engineering. We analyze a real-world dataset of espresso brewing by professional cofee-making machines. The novelty of the proposed work is provided by the focus on the brewing time series, from which we propose to engineer features able to improve previous data-driven metrics determining the quality of the espresso. Thanks to the exploitation of the proposed features, better quality-evaluation predictions are achieved with respect to previous data-driven approaches that relied solely on metrics describing each brewing as a whole (e.g., average flow, total amount of water). Yet, the engineered features are simple to compute and add a very limited workload to the cofee-machine sensor-data collection device, hence being suitable for large-scale IoT installations on-board of professional cofee machines, such as those typically installed in consumer-oriented business activities, shops, and workplaces. To the best of the authors' knowledge, this is the first attempt to perform a data-driven analysis of real-world espresso-brewing time series. Presented results yield to three-fold improvements in classification accuracy of high-quality espresso cofees with respect to current data-driven approaches (from 30% to 100%), exploiting simple threshold-based quality evaluations, defined in the newly proposed feature space.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Espresso is an almost syrupy beverage generated by a machine,
typically using a motor-driven pump, forcing pressurized hot
water through finely ground cofee. Each espresso shot in a bar
can generate one or two cups of cofee, being called, respectively,
single or double, and requiring proportional amounts of ground
cofee.</p>
      <p>
        Drinking espresso cofee is a ritual rooted in the pleasure of
its taste. In some countries, such as Italy, where 97% of adults
drink espresso daily [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], espresso quality is a main driver for
consumers’ habits and a primary focus of cofee industries.
      </p>
      <p>
        In 2018, each Italian had 2.2 daily espresso cups on average,
i.e., 6 kg yearly, in one of the 150 thousand bars, with each bar
using 1.2 kg of ground cofee daily to serve almost 200 cofees on
average, and most of them were espresso, representing
approximately one third of a medium bar turnover [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        According to common knowledge and online sources [
        <xref ref-type="bibr" rid="ref12 ref18">12, 18</xref>
        ],
such as the Italian Espresso National Institute, a perfect espresso
depends on diferent variables: (i) the cofee blend, (ii) the grinder
settings, i.e., the weight of cofee grounds and how fine it is
ground; (iii) the espresso machine, with professional machine
makers improving such technology over and over to promise the
perfect espresso all the time; (iv) the barista, i.e., the
human-inthe-loop preparing the espresso in the bar, from blend choice,
to manual grinder settings, and to proper usage of the cofee
machine and its brewing procedure.
      </p>
      <p>
        In the current work, among the diferent quality-influencing
variables, we focus on (i) cofee ground size, (ii) ground amount,
and (iii) water pressure. Regarding the quality-evaluation
variables, we exploit the following common metrics as selected by
domain experts and related works: (i) total extraction time, (ii)
the total volume of cofee in cup, and (iii) the derived average
lfow of the extraction [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        The ideal portion [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] of ground cofee for each cup is declared
to be 7 ± 0.5 g, while the water pressure should be 9 ± 1 bar, the
extraction time 25 ± 5 s, and the volume in cup 25 ± 5 ml.
      </p>
      <p>The cofee ground derives from the process of cofee grinding
from cofee beans. Small changes in the grind size can drastically
afect the taste and the quality of the brewed espresso. In general,
if the cofee is ground too coarse, the espresso can be
underextracted and less flavorful. On the other hand, too fine ground
may result in an over-extracted and bitter cofee. The amount
of ground itself impacts on quality, resulting in a too watery
or bitter cofee. Water pressure must be set to brew the right
cofee amount in a proper time, thus leading to the right flow
rate determining an intense flavour.</p>
      <p>The novelty of the proposed work is provided by the
exploitation of the brewing time series, from which we propose to
engineer features able to improve the standard data-driven metrics
determining the quality of the espresso, i.e., extraction time,
volume, and flow (as the ratio of volume and time). The proposed
features are applied on a real-world dataset where we show that
they can provide better quality-evaluation predictions, by
allowing to reduce the false positives, i.e., apparently good cofees,
without any loss in true positives.</p>
      <p>Since the engineered features are simple to compute and add a
very limited workload to the cofee-machine sensor-data
collection device, they are also suitable for large-scale IoT installations
on-board of professional cofee machines, such as those typically
installed in consumer-oriented business activities, shops, and
workplaces.</p>
      <p>Presented results uncover insights into the espresso quality
evaluation, its relationships with the main quality variables,
leading to positive impacts on both cofee consumers and
cofeemaking industries, respectively enjoying and providing more
pleasure in drinking higher-quality espresso cofee.</p>
      <p>The rest of the paper is structured as follows. Section 2
discusses related works, Section 3 describes the dataset and the
experimental design, Section 4 introduces the time-series feature
engineering algorithm, and Section 5 presents experimental results.
Finally, Section 6 draws conclusions and outlines future works.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>Espresso quality assessment is traditionally performed with
sensory analysis, the scientific discipline that statistically and
experimentally analyze reactions to stimuli perceived through the
human senses (sight, smell, taste, touch and hearing). Sensory
evaluation is however time-consuming and afected by
subjectiveness and low-reproducibility due to the human component.</p>
      <p>
        Considering these limitations, objective analysis as chemical
techniques, electronic noises and data-driven approaces are
commonly exploited for cofee quality control. Diferent chemical
techniques adopt Gas Chromatography (GC) and Mass
Spectroscopy (MS) analysis. Several works study the efect of external
variables (e.g. water pressure, water temperature) or of cofee
characteristics on the final espresso quality. Some works are
focused on the influence of water, as its composition, pressure [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
temperature [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and of water pressure and temperature
combined [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Others studies instead consider the impact of cofee
features themselves, as the roasting conditions [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] or the type
of cofee and roast combined [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        However, GC and MS analysis often require a significant
amount of time and human intervention. Many studies exploit
Electronic Nose (EN) systems to overcome the complexity and
cost of GS/MS techniques. An electronic nose is a device intended
to mimic human olfaction. It consists of an array of chemical
sensors for chemical detection and a pattern recognition system
capable of identifying the specific components of an odor [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. EN
are frequently exploited for determining and discriminating
coffee characteristics. Several works aim at determining the roasting
degree [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], using PCA and Neural Networks (NN) coupled with
GRNN, while others focus on distinguishing cofee blends,
exploting both NN [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and Support Vector Machines techniques [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
EN systems are also used in conjunction with GS analysis, as
in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], to characterize roasting degree and cofee beans from
diferent countries. The analysis in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] studies espresso chemical
attributes when the extraction time and grinding level are varied.
The work emphasizes the importance of the first 8 seconds of the
espresso brew, because in this range the major amount of organic
acids, solids and cafeine are extracted. This result confirms the
relevance of analyzing the entire trend of cofee extractions to
characterize their quality.
      </p>
      <p>
        Finally, data-driven approaches can be applied for large-scale
and real-time espresso quality assessment, exploiting Internet of
Things (IoT) sensors in place of the more sensitive and unstable
EN devices. Recently, a data-driven approach that exploits
association rule mining has been proposed to analyze the correlation
of cofee-making machine parameters and espresso quality [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The work relies solely on metrics describing each espresso
brewing as a whole (e.g., average flow, total amount of water). In the
proposed work, instead, we focus on the brewing time series to
fully characterize the cofee extractions.
      </p>
      <p>
        Time series analysis is a popular and well-known approach in
many application fields [
        <xref ref-type="bibr" rid="ref10 ref13">10, 13</xref>
        ], from physiological data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to
energy and weather data [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. However, in our work, we exploit
a basic intuition on the time series trend and resort to feature
engineering to avoid a direct analysis of the time series itself.
Feature engineering from time series has been extensively
addressed for diferent applications, as in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for industrial one in
the context of IoT and Industry 4.0, or for pattern matching of
technical patterns in financial applications [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>With respect to the state of the art, the current work
contributes by cleverly transferring known and simple time-series
feature engineering techniques into the espresso quality
evaluation domain, leading to significant improvement in classification
performance with respect to the state of the art. To the best of the
authors’ knowledge, this is the first attempt to perform a
datadriven analysis of real-world espresso-brewing time series, as
until now the focus has been limited to whole-extraction metrics.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DATASET DESCRIPTION</title>
      <p>The dataset under analysis consists of real-world espresso
brewing data. Since the dataset is provided by a leading cofee
company, we cannot disclose exact details of the real-world settings
(e.g., the cofee-machine maker and model, the precise location
and name of the involved business activities). Each espresso
extraction has been performed on professional cofee-making
machines and the values of the quality-evaluation variables have
been collected every 300 ms. In particular, our time series consist
of the values of the amount of water at each time interval, as
provided by flow-meter pulse counter, then deriving the instant
lfow rate (i.e., the ratio of the amount of water and the time).</p>
      <p>Each extraction has been performed with specific values of
the quality-influencing variables, hence allowing us to know the
ground-truth labels of high-quality espresso cofees, i.e., those
having all optimal settings for (i) cofee ground size, (ii) ground
amount, and (iii) water pressure. An exhaustive set of cofees has
been produced to observe the efect of non-optimal values on the
espresso quality. For each quality-influencing variable, diferent
values are considered: ground size can be coarse, optimal, or fine;
ground amount can be high, optimal, or low; brewing water
pressure can be high, optimal, or low. All possible combinations of
the three external-variable values (e.g., optimal, high, low) have
been included in the dataset, hence generating 33 = 27 possible
input configurations. For each configuration among the 27
combinations of external variables (for instance: coarse ground size,
optimal ground amount, and high water pressure), 20 espresso
extractions have been performed. Experiments have been repeated
on a professional cofee-making machine, generating a datasets
consisting of 540 espresso extractions.</p>
      <p>
        The domain-expert quality thresholds used in our experiments
are as follows: espresso volume from 20–30 ml, extraction time
from 20–30 s. The values have been selected according to public
literature, e.g., those published by the Specialty Cofee
Association of Europe [
        <xref ref-type="bibr" rid="ref12 ref5">5, 12</xref>
        ]. The flow rate thresholds derive from the
above-mentioned ones, as the flow rate is the ratio of the volume
by the time, hence obtaining the range 0.67–1.50 ml/s.
      </p>
      <p>Given such thresholds, espresso extractions can be labelled
with their quality assessment. Quality labels are optimal, too
low or too high for each of the quality variables: volume, time,
and flow. Table 1 recaps the domain-based threshold values and
corresponding labels.</p>
      <p>
        The problem tackled by this work stems from the fact that
analyzing the standard quality-evaluation variables without the
additional time-series novel features, many false positives are
provided: some espresso extractions are characterized by
highquality values in terms of water amount, flow rate and extraction
time, however, their ground size, ground amount or water
pressure were not optimal (compensation efect [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>TIME-SERIES FEATURE ENGINEERING</title>
      <p>Feature engineering refers to the process of extracting features
from raw data. It is typically executed to improve the performance
of predictive or classification models. In the current work, we
exploit feature engineering to leverage the cofee-brewing time
series with the aim of improving the espresso quality assessment.</p>
      <p>For each cofee extraction, the time series of the flow-meter
pulses is stored, with sampling time equal to 300 ms. Flow-meter
pulses are firstly converted to quantity of brewed water q, as
follows:
(1)
q =
nump ∗ pulseq</p>
      <p>numc
where nump is the number of pulses of the flow-meter, pulseq
represents the quantity of brewed water per pulse of the
flowmeter and numc represents the number of brewed cofees. In the
experimental data under analysis, pulseq =0.5 ml, as given by the
cofee-machine datasheet, and numc =2, since two espresso
coffees are brewed for each extraction. The time series captures the
water quantity over time, hence the instant flow rate is known.</p>
      <p>Figure 1 shows an example of a real time series from the dataset.
We notice a clear two-segment trend that is observable for any
arbitrary extraction: a first steeper phase is followed by a second
part having a lower flow rate. This phenomenon is known by
domain experts. In the first, transient, phase of cofee brewing,
water is forced in the cofee panel inside the filter holder, and
cofee grounds do not slow the water flow yet. On the contrary,
in the second phase, water penetrate and dampen cofee grounds
yielding the actual cofee extraction.</p>
      <p>We propose to extract the following new features to capture
the two-fold behavior of the extraction. We firstly determine the
point where a significant flow variation is observed. We refer
to this point as trend point. The trend point is used to
approximate the water quantity time series as a polygonal chain. The
approximate polygonal chain is constituted by two line segments
that represent the two phases of the water flow and its vertex
of intersection is the trend point. The trend point is estimated
by considering the maximum variation of the slope average of
the points in two consecutive not-overlapping sliding windows
of size W . The slope si (or gradient) of two consecutive points
pi = (ti , qi ) and pj = (tj , qj ) is computed as follows.
si =
qj − qi
tj − ti
In Equation 2, t is the time reference and q is the water quantity,
and they represent the axes of Figure 1. The slope s describes the
steepness of the water flow.</p>
      <p>The procedure for the trend point estimation is reported in
Algorithm 1.</p>
      <p>The maximum variation of the slope and the corresponding
points are initialized in Lines 1 and 2. In Lines 4 and 5, two
consecutive not-overlapping sliding windows of size W are defined.</p>
      <p>Let wk be a time window of size W . The slope average wkmean
of all consecutive points of the time window is computed as
follows</p>
      <p>1
wkmean = W − 1 j=1 tj − tj−1</p>
      <p>WÕ−1 qj − qj−1
(2)
(3)
where pj = (tj , qj ) and pj−1 = (tj−1, qj−1) are consecutive points
of the time window.</p>
      <p>The slope average is estimated for the two sliding windows, as
reported in Lines 6 and 7. The two terms capture the average flow
rate in the corresponding time window. The diference of the
two slope averages is computed in Line 8. The maximum slope
variation and the corresponding point are updated in Line 9.</p>
      <p>The point of maximum variation corresponds to the
intersection point of the two considered sliding windows. The process
is repeated until all N points of the time series are considered.
Finally, the trend point is returned (Line 12).</p>
      <p>The trend point ptp = (ttp , qtp ) represents the intersect vertex
of an approximate polygonal chain of the water quantity time
series. It is exploited to compute two features that capture the two
phases of the espresso extraction. Let be p0 = (t0, q0) and pN =
(tN , wqN ) the first and last points of the time series, respectively.
We define s1 and s2 as follows.
l)125
m
(
r
e
ta100
w
f
o
ty 75
i
t
n
a
uQ50
25
0
0
5
10
Time (s)
15
20
25</p>
      <p>In Figure 2, the approximate polygonal chain of a cofee
extraction time series is reported. The dashed line indicates the
average water flow. The slope s1 represents the average flow of
the first phase of the espresso brewing while slope s2 the average
lfow of the second phase. These two features are exploited in the
analysis to better characterize the cofee extraction, providing
additional information with respect to the overall average flow.
The extracted features will also be exploited to compute new
ranges for the optimal quality parameters, hence improving the
recognition of high-quality cofees.
5</p>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTAL RESULTS</title>
      <p>This section provides a description of the data cleaning
procedures applied to the dataset (Section 5.1), a discussion of the data
characterization of the extracted features (Section 5.2), and their
contribution to the espresso quality assessment improvement
(Section 5.3).
5.1</p>
    </sec>
    <sec id="sec-6">
      <title>Data cleaning</title>
      <p>
        The dataset has been pre-processed by applying the data
cleaning steps described in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The original dataset consists of 1080
cofees, corresponding to 540 extractions. Among them, 30
extractions were missing the time series data due to low-level hardware
issues. Domain-driven thresholds, aimed at removing values
being unacceptable for the phenomena under exam, lead to other 38
extractions to be discarded. As described in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], domain-driven
threshold values of valid espresso extractions have been set to
10–40 ml and 10–40 s, according to leading industrial domain
experts. Finally, the statistical-based outlier removal approach
of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] removed 15 additional samples from the dataset. After the
cleaning procedure, 457 extraction time series remain out of the
540 original records.
5.2
      </p>
    </sec>
    <sec id="sec-7">
      <title>Data characterization</title>
      <p>We firstly analyze the relationship between the extracted features
and the quality-evaluation variables (i.e., total extraction time,
average flow rate, total water amount). The trend point and the
consequent slope values have been computed with a window size
W set to 10.</p>
      <p>The correlation analysis shows that slope s2 is highly
correlated with the average flow rate (over the whole extraction),
with a Pearson correlation coeficient equal to 0.95, and the total
brewing time, with a correlation coeficient of -0.94. As expected,
lower flow rates lead to longer extraction times, since the total
amount of cofee is an almost constant goal of the cofee machine.</p>
      <p>We then investigate the relationship between the two
average flows (i.e. s1 and s2) and the three external quality-influencing
variables: water pressure, cofee ground amount and cofee ground
size, also known as grinding setting).</p>
      <p>Figure 3 shows the pressure behavior with respect to s1 and s2.
The pressure values (low, optimal, and high) are represented by
the label in the scatter plot. We can observe that cofee extractions
in the (s1, s2) space are clearly divided in three macro-areas,
determined by s1 value. The central partition is characterized by
an optimal pressure, while the first and last areas by low and high
values of pressure respectively. Hence, the value of the external
variable highly influence the first phase of cofee extractions,
when water is forced into the cofee panel. To a low pressure
corresponds a low water flow in the initial phase and vice versa
for the high pressure. The flow in the second phase is instead
almost independent from the pressure value.</p>
      <p>Regarding the total amount of water, we report in Figure 4
the cofee extractions as a function of s1 and s2. Diferently from
the pressure-labeled scatter plot, it is not observable a sharp
distinction. We can however identify a relationship with s2. Higher
amounts of cofee ground lead to lower values of the flow s2. In
this case, the average flow in the second phase of the extraction
is hindered by the higher amount of cofee ground. Hence, the
water flow is reduced. Likewise, the lower quantity of cofee
ground facilitates the flow of water, with a consequent increase
in flow s2. The cofee ground amount, instead, do not influence s1,
since it captures the average flow of the water when it is forced
in the cofee panel and before the cofee ground tampering.</p>
      <p>Finally, we observe a similar behavior when considering the
cofee ground size (i.e., grinding settings), hence we do not report
the plot. A coarser grinding generally corresponds to a higher
lfow. The finer cofee grinding instead hinders the water flow.
This results in a lower flow s2 in the second phase of the cofee
extraction.
5.3</p>
    </sec>
    <sec id="sec-8">
      <title>Quality Evaluation</title>
      <p>
        In this section, we evaluate the extracted feature ability to
characterize espresso quality and to improve the detection of
highquality espresso cofees. All the three external variables are under
the barista control. However, brew pressure is set at first in the
espresso machine calibration phase and it is periodically checked
and configured, typically with the support of technicians. On
the other hand, the grinding settings and the amount of cofee
ground are determined by the barista at each espresso brewing.
Hence, it is particularly relevant to control that these two
external variables are set properly by the barista. In existing works,
domain-experts and data-driven thresholds on quality indexes,
such as espresso volume, extraction time and brewing flow rate,
have been applied to evaluate cofee quality. The analysis in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
4.0
4.5
6.0
      </p>
      <p>6.5
5.0</p>
      <p>Slope 1
5.5
5.0</p>
      <p>
        Slope 1
5.5
explored the phenomena of compensating sub-optimal values
of diferent external variables. A compensation efect is
observable when configurations of values of external variables allow
to achieve apparently high-quality cofees, in terms of quality
indexes, despite one or more values are, in fact, not optimal.
Interpretable exploration techniques highlighted that high amounts
of cofee ground, that generally hinder the water flow and lead
to long percolation times, could be compensated by a coarser
grinding that, on the other hand, facilitates the flow [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Similarly, the low amounts of cofee ground could be compensated
by a finer grinding. Despite the optimal quality-index values,
the low amount of cofee has generally a negative impact on
cofee intensity and body, and therefore on the final customer
experience, hence possibly afecting also the brand image of the
cofee supplier. To this aim, we exploit the time-series features
to better characterize the quality of espressos so that false
highquality cofees can be detected and, if not totally avoided, at least
significantly reduced.
      </p>
      <p>
        As a reference, we consider domain-driven thresholds on
coffee quality indexes. In Figure 5 the espresso extractions with
optimal values of quality indexes are reported in the s1 and s2
space. They can be grouped as follows. (i) True high-quality
extractions present optimal values for both the quality-evaluation
indexes and, in particular, for all external variables. (ii) False
high-quality extractions present optimal quality-index values
with respect to domain-expert thresholds, but at least an external
variable has a sub-optimal value [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Such espresso extractions
(ii) are the result of compensation efects.
      </p>
      <p>We refer to true high-quality extractions as optimal, and we
characterize them as a function of the proposed time-series
features s1 and s2. Let O be the set of optimal extractions {o1, o2, ..., oN },
where each point oi ∈ O is defined in terms of s1 and s2, i.e.,
oi = (oi_s1, oi_s2 ). We define novel quality thresholds for optimal
extractions To_min and To_max in the (s1, s2) space as follows:
To_min = (min(oi_s1 ), min(oi_s2 ))
(6)</p>
      <p>To_max = (max (oi_s1 ), max (oi_s2 )) (7)
Among the whole set of espresso extractions E = {e1, e2, ..., eM },
a generic sample ej = (ej_s1, ej_s2 ) ∈ E is labeled as optimal
e ∈ O, with O ⊆ E, if its values of flow rate ( ej_s1 , ej_s2 ) are
within the thresholds To_min and To_max .</p>
      <p>In Figure 5 two rectangular areas are shown. The green area
contains the optimal extractions. Its boundaries are defined by the
thresholds To_min and To_max . The orange dashed area contains
the false high-quality extractions, which current
state-of-theart solutions would (incorrectly) classify as high-quality cofees.
Exploiting the proposed thresholds in the new feature space, we
can detect many false positives (orange squared points in the
plot). Specifically, instead of assigning an optimal label to the
overall 67 extractions (green and orange ones), we can correctly
detect the 20 true optimal extractions (green ones), and we can
discard 31 out of 47 false positives (orange ones). State of the
art thresholds would lead to the same true positive detection (20
out of 67), while the proposed approach leads to a drastically
better accuracy (76% instead of 30%) and precision of high-quality
extractions (56% instead of 30%).</p>
      <p>To drill down the analysis, we further distinguished two types
of false positives, stemming from diferent compensation efects:
(i) low amount of cofee ground with fine grinding and (ii) high
amount of cofee ground with coarse grinding. The former is
less common, since very few baristas intentionally use higher
amounts of cofee ground, being a cost for them. On the contrary,
the latter is much more frequent, because it brings savings on
cofee ground costs. For this reason, extractions afected by the
latter are of greater interest.</p>
      <p>In Figure 6 three areas are shown. The green one still contains
the true optimal extractions, the blue one contains the extractions
belonging to the first type of compensation and the orange one
now contains only the extractions belonging to the second type
of compensation. Again, exploiting thresholds in the new feature
space, the target extractions can be correctly classified and the
compensation efect can be detected. Results show that all 23
extractions from type-(ii) compensation can be correctly detected,
besides 8 extractions out of 24 from type-(i) compensation, which
means improving from 30% accuracy of data-driven state of the
False high-quality extractions</p>
      <p>True optimal extractions
4.0
4.5
5.0</p>
      <p>Slope 1
5.5
6.0
6.5
2.0
1.5
4.0
4.5
6.0</p>
      <p>6.5
5.0</p>
      <p>Slope 1
5.5
art to 100% accuracy considering only true optimal and
type(ii) compensation extractions. To this aim, in our dataset, the
new feature thresholds have been set as 5.19 &lt; s1 &lt; 5.48 and
2.64 &lt; s2 &lt; 3.73.
6</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSIONS</title>
      <p>This work presented a data-driven analysis of a real-world
timeseries dataset of espresso brewing by professional cofee-making
machines. The proposed feature space, despite being simple and
easy to compute, brought a large improvement in the
classification accuracy of high-quality espresso with respect to current
state-of-the-art data-driven approaches: results yielded to
threefold improvements in accuracy, from 30% to 100%, with specific
focus on currently misclassified extractions due to common
compensation efects. The proposed methodology can be applied
in similar contexts to improve current data-driven analyses of
espresso quality.</p>
      <p>Future works aim to widen the scope of the analysis
including additional quality variables, definitely diferent models of
professional cofee-making machines, diverse cofee blends, and
environmental variables. Furthermore, we plan to apply
clustering techniques for determining the quality-index thresholds.</p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work is partially funded by the SmartData@PoliTO center.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Andueza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Maeztu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. P. de Peña</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Bello</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cid</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Influence of Water Pressure on the Final Quality of Arabica Espresso Cofee</article-title>
          .
          <article-title>Application of Multivariate Analysis</article-title>
          .
          <source>J. Agric. Food Chem</source>
          <volume>50</volume>
          ,
          <issue>25</issue>
          (
          <year>2002</year>
          ),
          <fpage>7426</fpage>
          -
          <lpage>7431</lpage>
          . https://doi.org/10.1021/jf0206623 PMID:
          <fpage>12452670</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Andueza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Maeztu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pascual</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ibáñez</surname>
          </string-name>
          , M Paz de Peña, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Cid</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Influence of extraction temperature on the final quality of espresso cofee</article-title>
          .
          <source>J. Sci. Food Agric</source>
          .
          <volume>83</volume>
          ,
          <issue>3</issue>
          (
          <year>2003</year>
          ),
          <fpage>240</fpage>
          -
          <lpage>248</lpage>
          . https://doi.org/10.1002/jsfa.1304
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Andueza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Vila</surname>
          </string-name>
          , M. Paz de Peña, and
          <string-name>
            <given-names>C.</given-names>
            <surname>Cid</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Influence of coffee/water ratio on the final quality of espresso cofee</article-title>
          .
          <source>J. Sci. Food Agric</source>
          .
          <volume>87</volume>
          ,
          <issue>4</issue>
          (
          <year>2007</year>
          ),
          <fpage>586</fpage>
          -
          <lpage>592</lpage>
          . https://doi.org/10.1002/jsfa.2720
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Apiletti</surname>
          </string-name>
          , E. Baralis, G. Bruno, and
          <string-name>
            <given-names>T.</given-names>
            <surname>Cerquitelli</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Real-time analysis of physiological data to support medical applications</article-title>
          .
          <source>IEEE Trans. Inf. Technol. Biomed</source>
          .
          <volume>13</volume>
          ,
          <issue>3</issue>
          (
          <year>2009</year>
          ),
          <fpage>313</fpage>
          -
          <lpage>321</lpage>
          . https://doi.org/10.1109/TITB.
          <year>2008</year>
          .2010702
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Apiletti</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Pastor</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Correlating Espresso Quality with CofeeMachine Parameters by Means of Association Rule Mining</article-title>
          .
          <source>Electronics 9</source>
          ,
          <issue>1</issue>
          (
          <year>2020</year>
          ),
          <fpage>100</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Caprioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cortese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cristalli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Maggi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Odello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ricciutelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sagratini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sirocchi</surname>
          </string-name>
          , G. Tomassoni, and
          <string-name>
            <given-names>S.</given-names>
            <surname>Vittori</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Optimization of espresso machine parameters through the analysis of cofee odorants by HSSPME-GC/MS</article-title>
          . Food Chemistry
          <volume>135</volume>
          ,
          <issue>3</issue>
          (
          <year>2012</year>
          ),
          <fpage>1127</fpage>
          -
          <lpage>1133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Christ</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. W</given-names>
            <surname>Kempa-Liehr</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Feindt</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Distributed and parallel time series feature extraction for industrial big data applications</article-title>
          .
          <source>arXiv preprint arXiv:1610.07717</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            and
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Luk</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ng</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Flexible time series pattern matching based on perceptually important points</article-title>
          . (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Di Corso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cerquitelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Apiletti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>METATECH: METeorological data analysis for thermal energy characterization by means of self-learning transparent models</article-title>
          .
          <source>Energies</source>
          <volume>11</volume>
          ,
          <issue>6</issue>
          (
          <year>2018</year>
          ). https://doi.org/10.3390/en11061336
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Esling</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Agon</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Time-series data mining</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 45</source>
          ,
          <issue>1</issue>
          (
          <year>2012</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Gardner</surname>
          </string-name>
          and 1956 Bartlett,
          <string-name>
            <surname>Philip N.</surname>
          </string-name>
          <year>1999</year>
          .
          <article-title>Electronic noses : principles and applications</article-title>
          . Oxford ; New York : Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Istituto</given-names>
            <surname>Nazionale Espresso Italiano</surname>
          </string-name>
          . [n.d.]. Espresso Italiano Certificato. http: //www.espressoitaliano.org/files/File/istituzionale_inei_hq_en.pdf/. [Online; accessed January-2020].
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E</given-names>
            <surname>Keogh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hart</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Pazzani</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Segmenting time series: A survey and novel approach</article-title>
          .
          <source>In Data mining in time series databases. World Scientific</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Michishita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Akiyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hirano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ikeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sagara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Araki</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Gas chromatography/olfactometry and electronic nose analyses of retronasal aroma of espresso and correlation with sensory evaluation by an artificial neural network</article-title>
          .
          <source>J. Food Sci. 75</source>
          ,
          <issue>9</issue>
          (
          <year>2010</year>
          ),
          <fpage>S477</fpage>
          -
          <lpage>S489</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pardo</surname>
          </string-name>
          , G. Niederjaufner,
          <string-name>
            <given-names>G.</given-names>
            <surname>Benussi</surname>
          </string-name>
          , E. Comini, G. Faglia, G. Sberveglieri,
          <string-name>
            <given-names>M.</given-names>
            <surname>Holmberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Lundstrom</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Data preprocessing enhances the classification of diferent brands of Espresso cofee with an electronic nose</article-title>
          .
          <source>Sensors and Actuators B: Chemical</source>
          <volume>69</volume>
          ,
          <issue>3</issue>
          (
          <year>2000</year>
          ),
          <fpage>397</fpage>
          -
          <lpage>403</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pardo</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Sberveglieri</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Classification of electronic nose data with support vector machines</article-title>
          .
          <source>Sensors and Actuators B: Chemical</source>
          <volume>107</volume>
          ,
          <issue>2</issue>
          (
          <year>2005</year>
          ),
          <fpage>730</fpage>
          -
          <lpage>737</lpage>
          . https://doi.org/10.1016/j.snb.
          <year>2004</year>
          .
          <volume>12</volume>
          .005
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Romani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cevoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fabbri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Alessandrini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. Dalla</given-names>
            <surname>Rosa</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Evaluation of cofee roasting degree by using electronic nose and artificial neural network for of-line quality control</article-title>
          .
          <source>J. Food Sci. 77</source>
          ,
          <issue>9</issue>
          (
          <year>2012</year>
          ),
          <fpage>C960</fpage>
          -
          <lpage>C965</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <article-title>Rossi writes</article-title>
          . [n.d.].
          <article-title>Cofee in Italy or 101 Facts about Italian Cofee Culture</article-title>
          . http://rossiwrites.com/italy/italy
          <article-title>-for-foodies/ cofee-in-italy-italian-cofee-culture</article-title>
          . [Online; accessed January-2020].
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schenker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Heinemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Huber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pompizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Perren</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R</given-names>
            <surname>Escher</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Impact of Roasting Conditions on the Formation of Aroma Compounds in Cofee Beans</article-title>
          .
          <source>J. Food Sci. 67</source>
          ,
          <issue>1</issue>
          (
          <year>2002</year>
          ),
          <fpage>60</fpage>
          -
          <lpage>66</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Severini</surname>
          </string-name>
          , I. Ricci,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Derossi</surname>
          </string-name>
          , and T. De Pilli.
          <year>2015</year>
          .
          <article-title>Changes in the Aromatic Profile of Espresso Cofee as a Function of the Grinding Grade and Extraction Time: A Study by the Electronic Nose System</article-title>
          .
          <source>J. Agric. Food Chem</source>
          <volume>63</volume>
          ,
          <issue>8</issue>
          (
          <year>2015</year>
          ),
          <fpage>2321</fpage>
          -
          <lpage>2327</lpage>
          . https://doi.org/10.1021/jf505691u PMID:
          <fpage>25665600</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>