<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Model for Detecting Changes in the Ease of Breathing of COPD Patients</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas T. Kok</string-name>
          <email>thomas.kok@ugent.be</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Willemijn Groenendaal</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dolores Blanco-Almazán</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lien Lijnen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christophe Smeets</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Ruttens</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Morales</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Dhaene</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Femke Ongenae</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Van Hoecke</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dirk Deschrijver</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Biomedical Research Networking Center in Bioengineering</institution>
          ,
          <addr-line>Biomaterials and Nanomedicine, CIBER-BBN</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Medicine and Life Sciences, Hasselt University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Respiratory Medicine</institution>
          ,
          <addr-line>Ziekenhuis Oost-Limburg</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Future Health department</institution>
          ,
          <addr-line>Ziekenhuis Oost-Limburg</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>IDLab, Ghent University - imec</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Imec The Netherlands/Holst Centre</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper introduces a new machine learning based comparator model to assess changes in the ease of breathing of COPD patients during loaded breathing. The comparator model is based on a random forest classifier that detects whether breathing becomes either more dificult, easier or remains stable. The designed model can accurately detect respiratory changes by comparing temporal segments of physiological signals measured during loaded breathing, with an  1 score of almost 80%, resp. 70% for the wearable solution. As the model is trained and tested with features derived from diferent signal modalities, such as respiratory flow, audio, bio-impedance and accelerometer data, we also did a systematic comparison of the signal modalities to assess their predictive power.</p>
      </abstract>
      <kwd-group>
        <kwd>machine learning</kwd>
        <kwd>biomedical signals</kwd>
        <kwd>respiratory status</kwd>
        <kwd>COPD</kwd>
        <kwd>ease of breathing prediction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Chronic obstructive pulmonary disease (COPD) is a
chronic inflammatory lung disease where the airflow
adverse symptoms, such as breathing dificulty and
shortness of breath, excess phlegm or sputum, and frequent
coughing or wheezing. In the United States, it is
estimated that 16 million Americans have this disease,
whereas millions more people sufer from COPD but
have not yet been diagnosed [1]. As a result, morbidity
and mortality in COPD patients are considerably high. In
order to diagnose COPD, a spirometer test is applied as a
gold standard test for pulmonary function. The patient
inhales and then exhales with a maximal efort through a
mouthpiece, while the airflow going into and out of the
lungs is measured and analyzed [2].</p>
      <sec id="sec-1-1">
        <title>A limitation of this test is that it has to be performed by trained medical personnel, and thus does not allow longitudinal monitoring of the respiratory status, which is desirable, because early detection and treatment of</title>
        <p>the predictive power of diferent physiological signals is</p>
        <p>Time
Time</p>
        <sec id="sec-1-1-1">
          <title>Feature extractor</title>
        </sec>
        <sec id="sec-1-1-2">
          <title>Feature extractor</title>
          <p>still largely unaddressed.</p>
          <p>This paper aims to address and explore this by
proposing a new machine learning based comparator
model that detects whether breathing becomes either
more dificult, easier or remains stable. The patient’s
ease of breathing is assessed by comparing pairs of
segments of non-invasive physiological signals that were
recorded during a loaded breathing test. By analysing
the performance of the model, valuable new insights
are obtained on the choice of signals that are most
relevant to monitor changes in respiration. As such, the
outcome of this study provides useful insights to decide
which signal modalities should be measured by the
wearable device, and forms a basis for the development
of next-generation wearable respiratory monitoring
technology. The availability of such a technology can
possibly result in a faster intervention to prevent disease
worsening, potentially leading to reduced health care
costs, hospital (re)admissions and improved quality of
life.</p>
          <p>In summary, the main contributions of this paper
are threefold:
• A systematic analysis and comparative study of
multiple physiological signals is performed, to
identify which ones are most influential to
correctly detect changes in the ease of breathing of
a COPD patient during inspiratory loading.
• Relevant features are automatically extracted
from the signals, and a state-of-the-art
classification algorithm for time series data is used to
build a comparator model that accurately detects
whether breathing becomes either more dificult,
easier or remains stable.
• The performance of the model is validated on data
from a clinical study and the results are evaluated.
Comparator
model
(load A – load B) ≥ t
“easier breathing ”
| load A – load B | &lt; t
“stable breathing ”
(load B – load A) ≥ t
“more difficult breathing ”</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Materials and Methods</title>
      <p>2.1. Data collection and setup</p>
      <sec id="sec-2-1">
        <title>The study uses data gathered from 50 patients with COPD</title>
        <p>and enrolled for an inspiratory loaded breathing test in
a clinical setting at Ziekenhuis Oost-Limburg (Belgium).</p>
        <p>The data from six patients were excluded in this study
due to one or more of the collected signal modalities
being either unavailable or saturated. The resulting data
set comprises 44 patients.</p>
        <p>The study was approved by the local institutional
medical ethics committee from Ziekenhuis Oost-Limburg with
reference 18/0047U. The study followed the World
Medical Association’s Declaration of Helsinki on Ethical
Principles for Medical Research Involving Humans Subjects.</p>
        <p>All patients provided written informed consent.</p>
        <p>An incremental inspiratory threshold loading
protocol was performed, where the patients were imposed to
increasing inspiratory loads proportional to their
maximal inspiratory pressure (MIP). The loads are quantified
as 0%, 12%, 24%, 36%, 48% and 60% of the MIP that was
measured before the start of the test [8]. As such, a
higher load corresponds to an overall higher inspiratory
efort that is required to breathe. Each load test was
applied for a period of thirty breaths (i.e. a variable time
length), with a two-minute resting period between the
tests for each diferent load. The loads were sequentially
applied in increasing order. The measurement of the
MIP and the imposition of the loads was both done
using an inspiratory muscle trainer (POWERbreathe KH2,
POWERbreathe International Ltd, Southam, UK).</p>
        <p>During the test, two systems were used to record
several physiological signals simultaneously. The first
system was a standard wired acquisition system (MP150,
Biopac Systems, Inc., Goleta, CA, USA), used to record
the accelerometer data and audio signals; the airflow
was measured using Biopac, together with a pneumotach
transducer (TSD107B, Biopac Systems, Inc.) connected
to a diferential amplifier (DA100C, Biopac Systems, Inc.).</p>
        <p>Given the fact that the lung sound signals have a
bandwidth around 4000 Hz [9], the Biopac was configured to the bio-impedance signals are upsampled using cubic
record at a sampling frequency of 10,000 Hz, and this was interpolation from 16 Hz to 100 Hz.
applied to all the channels. The second system was a low- After filtering, a bidirectional moving average and
power wearable device (imec the Netherlands, Eindhoven, moving variance filter is applied to remove artefacts in
the Netherlands) with an injecting current of 100 uAp-p the data where the patient did not breathe into the
pneuat 80 kHz that was used to record the bio-impedance sig- motach transducer correctly. When either of these values
nal with a sampling frequency of 16 Hz [10]. The location in either direction is close to zero, this is considered an
and configuration of the signals were as follows: artefact and removed.</p>
        <p>After removal of artefacts, the signal is divided into
• Respiratory flow (gold-standard, used as a base- multiple shorter signals. In order to define a
commonline) length input signal for the machine learning model, each
• Accelerometer: parasternal and diaphragm part of the signal that is not corrupted, is subdivided into
(lower intercostal spaces), three axes several smaller segments having a predefined window
• Microphone (audio): left lung, right lung, tracheal length  . For each combination of patient and inspiratory
• Bio-impedance (bioZ): 4 configurations load, one window is sampled from the longest stable
breathing period without interruptions or artefacts. This</p>
        <p>For the bio-impedance signals, the first of the four ensures that each segment has a consistent length.
channels described in [10] is used in this study, because It was found that a window length of  = 30 sec
it was shown to have a more robust performance for es- is the most adequate choice, because it ensures that at
timating respiratory volume changes when higher loads least 95% of the input space (i.e. all possible patient-load
are imposed. The full details about this data collection combinations) is still included in the dataset after
preprotocol for the respiratory flow, bio-impedance, and processing, while guaranteeing the inclusion of multiple
accelerometer signals, as well as the set-up of the ex- breaths. This window length strikes the balance between
periments are explained in [7, 11]. Regarding the audio including as much data as possible in the entire window
signals, these were recorded using three microphones and including as many patient-load combinations as
pos(TSD108, Biopac Systems, Inc) with a frequency response sible, avoiding biased results.
of 35-3500 Hz. Two microphones (for both lungs) were
positioned on the back, two to three centimeters below 2.3. Machine learning-based comparator
the shoulder blades, at each side of the spinal cord. The model
other microphone (for the tracheal sound) was positioned
on the right side of the patient’s neck. After amplifying
the sound 200 times, it was filtered with an analog
lowpass filter of 5 kHz and a high-pass filter of 0.05 Hz. In the
forthcoming sections of the paper, the respiratory flow is
referred to as an obtrusive signal modality, whereas all
others are considered as unobtrusive signal modalities.</p>
        <p>Fig. 1 shows an overview of the designed comparator
model. Pairs of segments from the same signal modalities
of the same patient are used as an input for the model,
and a three-way classification is calculated, depending
on whether the load of the second signal is lower than,
higher than, or comparable to the load of the first signal.</p>
        <p>It is hypothesized that this method can then also be used
2.2. Preprocessing and segmentation to detect increased dificulty in breathing, which may
be indicative for worsening of the respiratory condition.</p>
        <p>The signals contain noise due to subject movement, elec- The threshold  used to separate the categories is set to
trical inference, measurement noise and other distur- 12%, to match the granularity of variations in the loads.
bances. In order to extract all relevant information, all For each signal type, a comparator model is trained
signals except audio are first filtered with a low-pass using pairs of signals recorded at diferent load
combizero-phase Butterworth filter with the following orders nations, with the labels corresponding to an increase,
and cut of frequencies per signal modality: a fith-order decrease, or stability in inspiratory load. Since the model
40 Hz low-pass filter for the respiratory flow signal, a is trained on data from all the patients in the training
fourth-order 0.7 Hz low-pass filter for the bio-impedance set at once, a general model is obtained, rather than a
signal, and a eighth-order 40 Hz low-pass filter for the patient-specific one.
accelerometer data. The filter values are based on the As six diferent loads were considered in the breathing
characteristics of the noise present in the signals, remov- test, 21 pairs of load combinations can be generated. The
ing higher frequency noise while retaining the relevant swapped pairs are also included to train the model,
resultinformation below the cut of frequencies. ing in 42 load pairs. As such, it would have been possible</p>
        <p>All signals are resampled to a common sampling rate. to have 1848 pairs. However, one patient was able to
The respiratory flow, accelerometer, and audio signals perform only 5 loads instead of 6. For this reason, 1836
are downsampled from 10,000 Hz to 100 Hz, whereas pairs were available for the development of the model.
2.3.1. Feature extraction
2.3.2. Model training and cross-validation
The performance of the models is evaluated using
10fold cross-validation, ensuring no overlap of samples
from patients between diferent folds. For each training
and test step, features are extracted from the 9 training
folds with the exact same parameters each time. The
model performance is assessed for every pair of loads
from patients in the test fold. Note that the evaluation
is based on a weighted  1 score to take into account
the small class imbalance, as instances from the stable
condition are less common than from the worsening and
improving conditions.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <sec id="sec-3-1">
        <title>An overview of the model performance for various signal</title>
        <p>modalities is shown in Fig. 2.</p>
        <p>A one-way ANOVA analysis is performed on all results
of the unobtrusive signal modalities, with  = 0.05 [15].</p>
        <p>For this, all  1 scores are considered from each signal
modality as a separate group, with the null hypothesis
that the means for each group are sampled from the same
distribution. The resulting p-value of 0.029 &lt;  and
rejection of this null hypothesis confirms that there is a</p>
        <p>1</p>
        <p>the classification performance of the model. For
examEasier 0.82 0.90 0.82 0.90 0.87 0.82 ple, the model has a 10% chance of missing an increase
Stable 0.69 0.88 0.71 0.88 0.83 0.70 in dificulty and a 19% chance of being wrong when an
More dif- 0.82 0.90 0.81 0.90 0.87 0.82 increase is detected. As such, it is important to interpret
ficult the classification results carefully and weigh the
potential consequences of false positives and false negatives,
Table 1 as the acceptable margin of error may vary depending on
Quality metrics of best performing model (sensitivity, speci- its application. For example, this error margin may be
acficity, positive predictive value, negative predictive value, ac- ceptable for regular monitoring of stable COPD patients,
curacy and  1 score) while it may not be acceptable for high-risk patients.</p>
        <p>When considering other signals that can easily be
acquired with wearable devices, a lower  1-score is
obstatistically significant diference in the mean  1 scores served when compared to the use of the respiratory flow
of the signal modalities. Amongst those signal modal- signal. Nevertheless, the actual performance of these
ities, bioZ, accelerometer (parasternal) and audio (tra- models is less important from a clinical perspective,
becheal) seem to provide the best predictive power to detect cause the comparator model in this study is only used
changes in the ease of breathing. A paired t-test with to benchmark and rank the diferent signal modalities
 = 0.05 confirms that the  1 score for the tracheal audio according to their predictive power. Having these
insignal is significantly higher than the left lung (  = 0.002 ) sights can be valuable, because an identification of the
and right lung ( = 0.028 ) signals. Furthermore, the best-performing signal modalities can help to make an
 1 score for parasternal accelerometer signal is signifi- informed choice of sensors during the design of a
wearcantly higher than the diaphragm accelerometer signal able.
( = 0.025 ). There is no significant diference between Such a wearable can collect longitudinal data from
pathe  1 score of the bioZ and accelerometer (parasternal) tients during normal daily activities, on which the
comsignals ( = 0.821 ), the bioZ and audio (tracheal) sig- parator model can be retrained. Having more lengthy
nals ( = 0.401 ), and the accelerometer (parasternal) and signals makes it possible to consider multiple window
segaudio (tracheal) signals ( = 0.64 ). ments, which can further boost the model performance.</p>
        <p>Fig. 3 shows the confusion matrix of the overall best Furthermore, the availability of more extensive data sets
performing model, i.e. the random forest model based on creates new possibilities to apply advanced deep
learntsfresh features, that considers the (obtrusive) respiratory ing techniques that have shown to be efective on similar
lfow signal. It is seen that the majority of the instances problem settings with respiratory data [16], while also
enare classified correctly, and misclassifications are more hancing the generalizability of the current feature-based
common between neighbouring classes. Table 1 pro- model.
vides an overview of the performance metrics calculated Future work will focus on the identification of an
opto assess the quality of the model, including sensitivity, timal combination of unobtrusive signal modalities to
specificity, positive predictive value (PPV), negative pre- avoid redundancy in the selection of signal modalities
dictive value (NPV), accuracy and the  1 metric for each within a certain category. Longitudinal, clinical and
exof the three classes. From both Fig. 3 and Table 1, it is ternal validation of the approach will also be performed.
seen that changes in breathing dificulty can be identified Additionally, to increase trust in the model, interpretable
more accurately than stability. This is not unexpected, as machine learning methods will be explored. Generating
the stable class is more similar to the other two classes, explanations alongside classifications also allows for
phythan they are to each other. isicans to incorporate the reasoning of the model within
their own decisions.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <sec id="sec-5-1">
        <title>The results confirm that diferences in the respiratory</title>
        <p>pattern during the inspiratory load protocol applied to This paper presents a novel machine-learning based
comCOPD patients can be assessed by our machine learning parator model that detects changes in the ease of
breathbased comparator model. The best performance is ob- ing of COPD patients during inspiratory loaded breathing.
tained when considering the respiratory flow signal, with Numerical results provide a comparison of diferent input
an  1 score of 0.78. This demonstrates the ability of the signals and models. When applied to the respiratory flow,
model to discriminate between an increase, remaining a weighted  1 score of 0.78 is obtained. When considering
stable or decrease in the ease of breathing for patients. other signal modalities that are not as obtrusive, and can</p>
        <p>However, there is still an error margin associated with be measured with wearable devices, the ones that ofer</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors have no conflicts of interest to declare. This
work was partially funded by the Flemish Government
(AI Research Program).
the best predictive performance are bioZ, accelerometer
(parasternal) and audio (tracheal), with a weighted  1
score of 0.69, 0.69 and 0.68 respectively.</p>
    </sec>
    <sec id="sec-7">
      <title>Overview of extracted features</title>
      <p>cross power spectral density
autoregressive process
coefLempel-Ziv complexity
estiabsolute energy
quantiles (10ℎ )
kurtosis
Combinations of simple
features
 2 &gt; 
standard error ⋅ 1

average over diferences
longest subsequence above
and below mean
sum of reoccuring values
(10ℎ )
energy ratio by chunks
mean of  largest values
first and last location of
maximum
Complex features
Benford correlation
c3 statistic
autocorrelation statistics
binned entropy
fourier entropy
FFT statistics
linear trend statistics
symmetry
index of of mass quantiles
percentage of unique values
number of zero-crossings
median
variance ( 2)
(absolute) maximum
root mean square
skewness
value and range count
mean
values
below
percentage of reoccuring
 ⋅  &gt; 
minimum
first and last location of
time reversal asymmetry
statistic
tance
number of (unique) peaks
permutation entropy
CWT coeficients
ficient
test
Langevin coeficients
augmented</p>
      <p>Dickey-Fuller</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>complexity-invariant dis-</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>