<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting respiratory failure in patients with COVID-19 pneumonia: a case study from Northern Italy</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Federica Mandreoli</string-name>
          <email>ica.mandreoli@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Guaraldi</string-name>
          <email>vanni.guaraldi@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Missier</string-name>
          <email>paolo.missier@newcastle.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). This volume is published and copyrighted by its editors. Advances in Artificial Intelligence for Healthcare</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Covid-19 crisis caught health care services around the world by surprise, putting unprecedented pressure on Intensive Care Units (ICU). To help clinical staff to manage the limited ICU capacity, we have developed a Machine Learning model to estimate the probability that a patient admitted to hospital with COVID-19 symptoms would develop severe respiratory failure and require Intensive Care within 48 hours of admission. The model was trained on an initial cohort of 198 patients admitted to the Infectious Disease ward of Modena University Hospital, in Italy, at the peak of the epidemic, and subsequently refined as more patients were admitted. Using the LightGBM Decision Tree ensemble approach, we were able to achieve good accuracy (AUC = 0.84) despite a high rate of missing values. Furthermore, we have been able to provide clinicians with explanations in the form of personalised ranked lists of features for each prediction, using only 20 out of more than 90 variables, using Shapley values to describe the importance of each feature.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In this paper we report on a machine learning exercise using an
evolving, unstable, and limited training set, aimed at supporting
hospital clinical staff during the COVID-19 crisis in Italy. The pandemic
evolved very rapidly over the course of a few weeks, with Italy
recording the first severe clusters of virus spread in Europe between
February and March, 2020. This put frontline health services into
emergency mode, forcing them to adapt very rapidly to an overload
of patients with severe complications, primarily viral pneumonia. In
addition to the clinical challenges, hospital medical staff had to deal
with a shortage of critical care resources, mainly Intensive Care Unit
(ICU) beds.</p>
      <p>At the University Hospital in Modena, Italy (UHM), this
translated into the urgent need to rapidly design end deploy new
protocols for recording medical records, which had to integrate routine
patient assessment information, eg blood tests, with observations about
their complications (respiratory issues), a record of patients transfers
across departments, namely the Infectious Diseases clinic, the ICU,
and a number of clinics to deal with specific complications. Doctors
also started recording details of tentative therapies, whose
effectiveness was largely unknown at the time.</p>
      <p>A new patients’ database was created, however this was subject
to a continuously evolving schema, with the result that the hundreds
of precious daily data points exhibited sparsity problems, i.e., when
a new variable was introduced and not retrospectively populated for
the existing patients, as well as heterogeneity and inconsistencies.</p>
      <p>Against this backdrop, there was an immediate need for
analytics that would help clinicians to answer some of the more pressing
questions. These included, besides the more clinical questions on the
efficacy of therapies, the challenge of predicting the needs for limited
ICU resources within a limited time horizon.</p>
      <p>
        We focused specifically on the problem of predicting whether
a patient would develop acute respiratory distress syndrome
(ARDS) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], leading to moderate to severe respiratory failure within
hours, and thus to the need for assisted breathing and to admin the
patient to ICU. This question translates well into a quantitative
outcome that can be used in machine learning, namely by measuring the
respiratory rate (a critical cutoff is &gt; 30 breaths per minute), blood
oxygen saturation &lt; 93%, and more importantly, the P aO2=F iO2
ratio of partial pressure of arterial oxygen to the fraction of inspired
oxygen. The reach of a clinically-defined cutoff (150 mmHg) in at
least one of two two consecutive days after 48 after admission, was
taken as the proxy measure of choice to put the patient in a critical
state where the need for assisted breathing was assessed. Clearly, the
ability to predict this outcome with some advance notice would give
medical staff information for manging ICU resources. From a
machine learning persective, this is a well-defined binary classification
problem, where the respiratory condition becomes a binary outcome
that can be evaluated on the ground training set represented by the
patients’ database.
      </p>
      <p>In the rest of the paper we outline the the challenges associated
with the machine learning task, we describe the approach and
evaluate the results, and outline the direction of our ongoing research.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Challenges and requirements</title>
      <p>The specific context around data collection and curation, as well as
the urgent need to manage ICU resource allocation, translated into a
number of requirements and technical challenges.</p>
      <p>Firstly, the dataset has been evolving rapidly not only in number of
records but also, critically, in the schema, with new attributes added
most daily as the requirements of downstream analysis became
increasingly clear. As explained in more detail in Sec. 2, the dataset
consist of Electronic Health Records (EHR) including routinely
collected clninical tests, but also ad hoc observations, associated with
the specific symptoms. Existing EHR collection systems were
therefore inadequate. A consequence of this predicament is a continuously
changing set of features, some of which require experts’
explanations, which complicates the learning process.</p>
      <p>Secondly, different labs may operate different practices and adopt
slightly different standards, including different laboratory
biomarkers. Care must therefore be taken when aggregating their values, as
in general their source would have to be taken into account.
Furthermore, different types of data are collected with varying frequency
over time. Some systematically, some only on demand, and some,
such as information about co-morbidities are collected only once,
but not for every patient.</p>
      <p>Related to this is the problem of data sparsity and of missing data.
One example is the Interleukin 6 variable, which was found to be
relevant for analysis only after weeks of lockdown, thus is absent for a
substantial proportion of the patients. While missing data can
sometimes be inferred, or imputed, from available value distributions, this
was not an option when dealing with critical patients vital
parameters, which by their own nature are subject to abrupt changes and
thus should not be exrapolated from known distributions. In fact, one
may argue that the value of the data in this context is the change in
data values, signalling an impending crisis.</p>
      <p>Thus, a key challenge for the model is to be robust to missing data,
a property that is not enjoyed by the majority of the available
off-theshelf libraries.</p>
      <p>We should also mention that two milestone data extraction
processes took place, with intermediate data versions in between, as
explained in the next Section. The characteristics of the dataset changed
with respect, for instance, to class imbalance and data sparsity,
requiring different modelling strategies.</p>
      <p>A separate challenge concerns the trustworthiness and
transparency of the model itself, both of which are required if the model is
to be embraced in clinical practice. As a general principle in Machine
Learning, models should be parsimonious: we seek to reduce model
complexity (the number of features required to learn the model, as
well as the non-linearity of the function learnt by the model) without
sacrificing prediction accuracy. This principle is particularly relevant
in this case, as only a limited set of variables can be presented to
medical professionals to describe the nature of the model in simple
terms, despite its non-linear, “black box” nature. Thus, accurate
feature selection and ranking is a critical requirement. Furthermore, we
seek to achieve personalised explanations, by associating a
potentially different list of variables, along with their relative weight, for
each individual prediction.</p>
      <p>Finally, when assessing model accuracy, we need to minimizes the
risk of under-estimating the severity of a patient’s condition, that is,
by reducing false negatives possibly at the expense of an increase in
the number of false positives.</p>
      <p>Our modelling approach, which takes account of all of these
requirements, is presented in Sec. 3.
1.2</p>
    </sec>
    <sec id="sec-3">
      <title>Case study</title>
      <p>In addition to testing theoretical model performance as reported in
Sec. 3, we have empirically validated our model on an early patient, a
55 year old man who was initially admitted to the clinic with typical
COVID pneumonia symptoms. He was treated and discharged the
next day after his clinical assessment established a stable condition,
with our outcome variable P aO2=F iO2 = 420 mmHg well above
the 250mmHg cutoff. However, he was readmitted to hospital four
days later with severe symptoms, and at that point his condition has
worsened to P aO2=F iO2 = 230 mmHg.
)
%
(60
y
t
lii
b
a
b40
o
r
P
20
0
66.4 %
47.4 %
120
10.2 %
17.5 %
225
200
100
75</p>
      <p>2
175 iO</p>
      <p>F
/
150 2O</p>
      <p>a
125 P</p>
      <p>In the following 24 hours, the patient experienced a clinically
unpredictable dramatic worsening (P aO2=F iO2 = 88 mmHg,
respiratory rate higher than 35 breaths per minute). He was then transferred
to ICU where non-invasive mechanical ventilation (NIV) was started.
After 8 days of assisted spontaneous breathing, he was weaned from
NIV and discharged the following day without oxygen supply.</p>
      <p>We retrospectively used this patient’s baseline assessment and then
his repeated medical evaluations to predict the probability of adverse
outcome at different points in time. The model predicted a 36%
probability based on the first admission, followed by much higher
confidence after the second admission and prior to ICU treatment, as
shown in Fig. 1. If the model have been available at that time, it
would have correctly alerted clinical staff against complacence with
the first dismissal, which was objectively justified at the time but
could not account for the population data that was instead available
to the model.</p>
      <p>We conclude that, in this particular instance, an early deployment
of our model would have provided effective support to assist clinical
judgment.
1.3</p>
    </sec>
    <sec id="sec-4">
      <title>Contributions</title>
      <p>We present our experience in developing a machine learning model
that satisfies a number of special requirements and addresses the
challenges outlined in the previous Section. The model proves that
you can bootstrap decision support to clinical staff in a time of
respiratory crisis, by making the best of an approximately curated,
constantly evolving dataset by rapidly customising out-of-the-box ML
algorithms.</p>
      <p>Specifically:
we have adapted LightGBM and designed a bespoke loss function
which includes a penalty , that can be tuned to adjust the model’s
FNR (on a test dataset).
we have successfully experimented with SHAP in combination
with LighGBM to provide post hoc explanations of the model’s
prediction, both globally and locally.</p>
      <p>The model is currently available through a hospital private web
service that lets clinicians probe the model with new cases and get a
prediction using a simple Web interface completely integrated with
their usual medical records management system.</p>
      <p>
        This is ongoing work, as the models presented in this work are
being periodically re-trained as more patients are added to the dataset.
The global spread of the pandemic by COVID-19 has raised the
level of attention of researchers around the world with regard to
the monitoring, treatment and study of diseases related to it. One
of the most important and characterizing aspects of this pathology
is the rapid evolution of the patients’ health into a crisis, requiring
ICU treatment. This makes ICU allocation strategies a priority and a
challenge[
        <xref ref-type="bibr" rid="ref14 ref8 ref9">9, 14, 8</xref>
        ].
      </p>
      <p>
        Our work aims to provide support to medical decisions both in the
triage phase and in the continuative patient monitoring phase.
Similar studies have been presented very recently [
        <xref ref-type="bibr" rid="ref10 ref15 ref4 ref7">7, 15, 10, 4</xref>
        ]. These
study share the common goal to segregate the most critical patients,
i.e. those requiring ICU or even approaching death, from those who
are improving their health status. The approaches generally used
take into consideration very similar variables, including biomarkers,
symptoms and co-morbidities, while the outcome chosen may differ,
and typically includes a critical respiratory event or death.
      </p>
      <p>The value of our work which makes it stand out in this space is
primarily its focus on ensuring trust by clinical staff. As we have
explained, we achieve this through a combination of a data-driven
strategy for feature set reduction and prioritisation, combined with a
loss function that ensures conservative predictions that penalise false
negatives. Testimony to this focus is the current experimental
deployment of the model as part of the hospital’s Information Systems, and
its upcoming availability throughout the province.
2</p>
    </sec>
    <sec id="sec-5">
      <title>Dataset characterisation</title>
      <p>The datasets used in this work were extracted from the Hospital
Information System of Policlinico di Modena, where a tailored data
collection protocol was implemented in order to gather new data
which were deemed relevant to assess the health status of COVID-19
patients. While at the beginning of the pandemic the schema for the
new data was vey close to that of the existing EHR management
system, new elements were increasingly added, as it became clear that
new biomarkers and symptoms were going to be relevant.</p>
      <p>We performed two milestone data extractions, after 27 and after
37 days from the start of the covid-specific data collection, which
included 91 of the available 99 variables. Specifically, we collected
a set of static variables, specifically sex and age, and the 14 most
relevant co-morbidities such as diabetes, cardiovascular diseases,
neoplasms, and hypertension. We also collected a number of
timevarying variables, which are measured periodically, as follows:
39 blood and urine tests, including standard blood test and
COVID-related ones such as Interleukin-6, Lymphocytes,
Troponin and C-reactive protein (CRP);
7 blood gas analysis (BGA) measures: pH, P aCO2, SO2,
Lactates, P aO2, HCO3 and F iO2;
29 different disease specific symptoms, e.g. dyspnoea, cough,
fever, conjunctivitis and shivers, and signs, e.g. heart rate, body
temperature and respiratory rate.</p>
      <p>Time-varying information were collected at different times. For
instance, most blood tests are collected daily but some specific tests
are collected “on-demand” based on clinical needs. An example is
Interleukin 6 which is collected only twice a week and only for the
patients that were treated with the immune active drug Tocilizumab.</p>
      <p>It is worth noting that we were forced to exclude variables related
to both drug therapies because they were collected only for the few
patients that received specific treatments and then were almost
always missing.</p>
      <p>A separate training set was derived from each of the two data
extractions. As mentioned in the introduction, the selected outcome was
a binary variable stating whether the patient would develop ARDS in
the next two days, measured as a P aO2=F iO2 ratio lower than 150
mmHg. The EHR of each patient p was therefore sliced in daily
snapshots and each sample, exemplified in Fig. 2, was represented by the
pair (xip; yi+1;i+2) where xip is the feature vector of the day i that
p
contains
the values of the static variables for the patient p;
the value recorded at day i for each daily-collected variables;
the last recorded value for each ”on demand”-collected variable.
As far as the outcome yip+1;i+2 is concerned, we considered the two
p
alternative options: the value of yi+1;i+2 is set to TRUE either when
P aO2=F iO2 &lt; 150 both at day i + 1 and at day i + 2 (AND
condition), or when P aO2=F iO2 &lt; 150 at least one the two days
i + 1 or i + 2 (OR condition). The two extractions contained the
records for 224 and 287 patients and 2454 and 2888 observations,
respectively.</p>
      <p>Time [days] i
● Blood Exams
● Urine Exams
● BGA
○ . .
○○ …PaO2/FiO2
● Symptoms
● Signs
● Therapies
● Sex
● Age
● Comorbidities
Daily Updated Data
Constant Data
Outcome</p>
      <p>i + 1
● Blood Exams
● Urine Exams
● BGA
○ . .
○○ …PaO2/FiO2
● Symptoms
● Signs
● Therapies</p>
      <p>i + 2
● Blood Exams
● Urine Exams
● BGA
○ . .
○○ …PaO2/FiO2
● Symptoms
● Signs
● Therapies</p>
      <p>...
● Blood Exams
● Urine Exams
● BGA
○ . .
○○ …PaO2/FiO2
● Symptoms
● Signs
● Therapies
Prediction
[PaO2/FiO2 ]i+1 &lt; 150</p>
      <p>OR
[PaO2/FiO2 ]i+2 &lt; 150</p>
      <p>Data sparsity and balancing issues for each of the two datasets are
summarised in Table 1, showing the mean and the variance of the
percentage of completeness for all the 91 variable and the
population distribution with respect to the two alternative outcomes, AND
conditions and OR conditions.</p>
      <p>The table reveals several problems having to do with missing data.
With only sex and age complete, the averages are predictably low.
Importantly, the number of collected values for some important
variables were very low. For instance 168 values were available for
lymphocytes (7.5%) and 515 for Interleukin-6 (20%) in the first
extraction. This increased slightly, to 557 (7.8%) and 642 (22.2%),
respectively, in the second extraction. Furthermore, the completeness of the
”on-demand” variables generally decreased in the second batch
because these variables are missing for most of the snapshots added in
the second batch. As a consequence, the mean percentage of
completeness decreased from 62% to 57%.</p>
      <sec id="sec-5-1">
        <title>First Data Extraction</title>
      </sec>
      <sec id="sec-5-2">
        <title>Second Data Extraction OR Condition AND Condition</title>
      </sec>
      <sec id="sec-5-3">
        <title>OR Condition</title>
        <p>AND Condition</p>
        <p>Sparsity represents a problem because most off-the-shelf learning
algorithms do not tolerate missing data. At the same time, removing
variables or data points was not an option, as sparsity affects most
variables and the training set is of very limited size. Imputation is not
an option, either, as the interesting signal about some of the important
variables is actually their abrupt changes within a patient’s timeline.</p>
        <p>On the other hand, the data in Table 1 suggests that using the “OR”
outcome, leads to a larger training set, because it only requires P aO2
and F iO2 to be available on any single day. This is in contrast to the
“AND” condition, where two consecutive data points are required.
For instance, considering the most recent data extraction, the
outcome built on the OR condition is available for 198 patients (68.9%)
and 1068 snapshots (36.9%) that actually represents the potential
number of samples. The distribution between the two classes, FALSE
and TRUE, is quite balanced, i.e. 417 (43%) and 557 (57%) samples,
respectively. In contrast, the distribution for the outcome built on the
AND condition is available only for 165 patients (57.4%) and 796
snapshots (27.5%) distributed between the two classes in 419 (59%)
and 287 (41%) samples respectively.</p>
        <p>Based on these considerations, we therefore decided to use the
OR condition, which is at the same time clinically valid, and suitable
from a ML modelling point of view, providing a larger and more
balanced training set, with a better representation of the positive class.
3</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Modelling approach and results</title>
      <p>Our main requirements for modelling include: (i) robustness to
dataset evolution, sparsity, and possible class unbalance; (ii) feature
selection (parsimony); and (iii) control over False Negatives. Here
we present our approach to modelling that meets these requirements.
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Robustness</title>
      <p>
        Robustness is achieved mainly through a choice of a learning
algorithm that can tolerate missing data, and the appropriate tuning of
the hyper-parameters. For this binary classification problem the main
candidates are XGBoost6 and LightGBM7. Both implement a form
of decision tree ensembles that can tolerate missing data in a tunable
way, where imputation is not an option as explained above. Decision
tree algorithms are based on the idea that at each node, the feature
is chosen to maximise some measure of information gain. XGBoost
extend the idea to missing values, by assigning a default direction
to each node in the tree, using a technique known as Sparsity-aware
split finding [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This determines which way the decision proceeds
when a feature value is missing and the corresponding node condition
cannot be evaluated. While both algorithms can also be retrofitted
      </p>
      <sec id="sec-7-1">
        <title>6 https://github.com/dmlc/xgboost 7 https://github.com/microsoft/LightGBM</title>
        <p>with explanations, as described below. LightGBM was selected
owing to its greater flexibility.
3.2</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Model and Feature selection</title>
      <p>Following standard experimental practice, the choice of learning
algorithm and features used to learn the model go hand in hand. Our
process involved (i) selecting a category of features with sufficient
predictive power, and (ii) reducing the feature space within that class,
in order to facilitate model interpretability. For this learning problem
we compared three models, which were built using the three sets of
features described earlier, namely (i) the entire set of 91 variables
in the dataset, (ii) only the 39 biomarker variables; and (iii) the 31
“symptoms and signs” variables. The results are reported in Table 2.</p>
      <p>The first model yields the best performance, however it was
unlikely it would have been useful in practice as it required extensive
and expensive data collection for each patient. It was also
featurerich and possibly redundant. The simpler subset of biomarkers, used
for the second model, is attractive as the data acquisition workflow
is entirely standard, and it provides an objective assessment of the
patient’s health status. Its performance is comparable with that of the
first, with a slight advantage in the number of FN.</p>
      <p>Finally, the 31-variables model was appealing in terms of data
collection, as the “signs and symptoms” variables included only
questions to patients and simple measurements such as heart rate and
temperature, which are easy to obtain. However it exhibited
inferior performance (AUC=0.69) along with a higher number of both
FN and FP relative to the previous two models. This result suggest
that such subjective data is quite less informative than objective and
instrumental data collection.</p>
      <p>All models were generated using standard ML practice, namely a
75=25 training / test split, random selection from the majority class
for balancing, 10-fold cross-validation, and hyper-parameter tuning
session using Grid Search.</p>
      <p>For feature set reduction we adopted a data-driven approach
centred on Shapley values8. These are generated by a values-based ML
model interpretation framework that provides both a global-level
assessment of the relative importance of each feature used by the
model, as well as a local view of how each feature is weighted when
making individual predictions. Furthermore, Shapley values offer an
interpretation of feature importance as a function of the value of the
feature. This leads to an intuitive interpretation, for instance “high
values of Dyspnea contribute strongly to an adverse outcome, while
low values make the feature relatively less important”. This sort of
explanation provides clinicians with a way to validate the model
against their own expertise, and thus may contribute to building trust</p>
      <sec id="sec-8-1">
        <title>8 https://github.com/slundberg/shap</title>
        <p>Dataset
91 Mixed variables
39 Biomarkers only
31 Signs and Symptoms</p>
        <p>Acc.
0.77
0.75
0.63
in the model’s predictions. Also, the framework uses a “post hoc”
method for generating such explanations, which can be used in
conjunction with non-linear models such as the decision trees generated
by LightGBM (for a linear model, feature ranking is simply given by
the model’s parameters).</p>
        <p>Using this approach and global-level feature rankings, and starting
from the original 91 variables, we repeatedly pruned features from
the set and retrained the model with the remaining features, using a
combination of performance and FNR to guide the process. The final
set consists of 20 core features, shown in Fig. 3. Note that these are
drawn from each of the variable categories: biomarkers, BGA, signs,
symptoms and co-morbidities, confirming that multiple views on a
patient’s status are required to generate reliable predictions.</p>
        <p>The first row of Table 3 shows the performance of the resulting
20-variables model 3.3.
The model emerging from the two steps described above can be
further tuned to achieve a trade-off between False Negatives and
accuracy. This was achieved by adding a new hyper-parameter to the
loss function that was minimised during the learning process. The
customized loss-function is therefore defined as follows:
When = 1, we get the standard Log-loss function. The effect
of tuning when training the 20-variables version of the model is
reported in Table 3. The table shows a clear trade-off between FN
and FP, as expected. The obvious measure that balances the two is
F P =F N , which is closest to 1 for = 2. As this is also the setting
where accuracy begins to drop, it is the one we used in the rest of
the experiments. Moreover, note that the FNR of this model is even
lower than in the earlier 39-variables model. The ROC curve of this
final model is depicted in figure 4.</p>
        <p>Receiver Operating Characteristic</p>
        <p>Train AUC = 0.98
Test AUC = 0.83
0.2
0.4
0.6
0.8
1.0</p>
        <p>False Positive Rate
The ML model presented in this paper is currently deployed behind
a web service and integrated into the Hospital’s Information
Systems where it can be used for medical decision support systems in
two phases of the patient assessment. Firstly, as a first triage
evaluation, e.g. in the Emergency Room, and secondly after admission and
through patient’s monitoring during their stay in hospital.</p>
        <p>
          An ongoing challenge with this project is the rapidly moving
dataset, both in schema as well as in content, as pointed out through
the paper. This translates into rapid changes in data distribution,
sparsity, and balance with respect to the outcome, and thus into the need
to periodically re-configure and re-train the model, including
potentially changing the implementation when it is no longer suitable.
This situation is also ideal for experimenting with “AutoML”
approaches, which are becoming increasingly popular both in the
scientific community and as part of commercial offering. These solutions
aims to build automatically tailored ML pipelines and include
operators for data transformation such as data imputation, feature
processing, classification and calibration algorithms. They have proven to be
of remarkable reliability and performance. One notable example is
AutoPrognosis [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], specifically tailored to clinical ML pipelines and
uses ML itself to proper configure and chose the best configuration
for the ML models it’s creating.
        </p>
        <p>Adopting AutoML solutions is planned for future work, in the
hope to improve the efficiency of the whole ML pipeline design, and
also to test such approaches in new emergency contexts, in which
there is no time for the manual development of the pipeline itself
(COVID-19 pandemic would have been an example).</p>
        <p>ML systems which are also capable of explaining the behavior of
the model can go far beyond mere prediction. Model interpretation
approaches in medical applications can be extremely useful, for
example, to discover new predictors for the chosen outcome. Also, with
the appropriate choice of population it may also be possible to
identify different predictors for different sub-populations, leading to new
insights into data-driven personalisation of predictive models.</p>
        <p>
          These approaches are becoming more and more effective and
popular in medical applications [
          <xref ref-type="bibr" rid="ref1 ref3 ref5">3, 1, 5</xref>
          ], also in the recent COVID-19
pandemic context [
          <xref ref-type="bibr" rid="ref10 ref15 ref7">7, 10, 15</xref>
          ]. As mentioned, our approach to
explanatory models involves Shapley values [
          <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
          ], which provide
a measure of the impact of each variable on the construction of the
predictions result for every instance. More specifically, these are
positive or negative numerical values and, in a binary classification task,
the sign indicates whether the variable contributes to the positive or
to the negative class. This provides an immediate and perception of
which of the variables have a positive vs negative impact on the
patient’s outcome.
        </p>
        <p>For instance, in our models it appears that lab variables are
stronger predictors of a patient’s clinical condition. however, Shapley
values show that these are best combined with non laboratory
variables when it comes to explaining the patient’s outcome and
obtaining better personalized insights on the individual. This can be seen
in the global view of these variables, shown in Fig. 3, and in Fig. 5,
where we show the local interpretation of the case study patient data
on the most critical day we have from his medical records (95%
probability of having respiratory failure within the next 2 days predicted
exactly the day before being transferred to ICU under mechanical
ventilation). Two of the most important variables here, respiratory
rate and dyspnoea, are not laboratory variables. This indicates how
some of these elements can be relevant to the analysis of health status
and in the formulation of a therapeutic plan.</p>
        <p>Finally, an interesting further study is to extend the model to
forecasting the entire patient’s journey through stages of disease and
treatment while in hospital, from admission to discharge.
Understanding the evolution of the patient’s state can be of great
importance both in the personalisation of the care plan, in the prevention
of adverse outcomes, as well as in the planning of hospital resource
organization.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ahmed</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Alaa</surname>
          </string-name>
          , Thomas Bolton, Emanuele Di Angelantonio,
          <string-name>
            <surname>James H. F. Rudd</surname>
          </string-name>
          , and Mihaela van der Schaar, '
          <article-title>Cardiovascular disease risk prediction using automated machine learning:</article-title>
          <source>A prospective study of 423</source>
          ,604 uk biobank participants',
          <source>PLOS ONE</source>
          ,
          <volume>14</volume>
          (
          <issue>5</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          , (05
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Ahmed</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Alaa</surname>
          </string-name>
          and Mihaela van der Schaar, 'Autoprognosis:
          <article-title>Automated clinical prognostic modeling via bayesian optimization with structured kernel learning'</article-title>
          , CoRR, abs/
          <year>1802</year>
          .07207, (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ahmed</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Alaa</surname>
          </string-name>
          and Mihaela van der Schaar, '
          <article-title>Prognostication and risk factors for cystic fibrosis via automated machine learning'</article-title>
          ,
          <source>Scientific Reports</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <fpage>11242</fpage>
          , (
          <year>Jul 2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Egon</given-names>
            <surname>Burian</surname>
          </string-name>
          , Friederike Jungmann, Georgios A.
          <string-name>
            <surname>Kaissis</surname>
          </string-name>
          , Fabian K. Loho¨fer, Christoph D. Spinner, Tobias Lahmer, Matthias Treiber, Michael Dommasch, Gerhard Schneider, Fabian Geisler, Wolfgang Huber, Ulrike Protzer,
          <string-name>
            <surname>Roland M. Schmid</surname>
          </string-name>
          , Markus Schwaiger,
          <string-name>
            <surname>Marcus R. Makowski</surname>
            ,
            <given-names>and Rickmer F.</given-names>
          </string-name>
          <string-name>
            <surname>Braren</surname>
          </string-name>
          , '
          <article-title>Intensive care risk estimation in covid-19 pneumonia based on clinical and imaging parameters: Experiences from the munich cohort'</article-title>
          ,
          <source>Journal of Clinical Medicine</source>
          ,
          <volume>9</volume>
          (
          <issue>5</issue>
          ), (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Rich</given-names>
            <surname>Caruana</surname>
          </string-name>
          , Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad, '
          <article-title>Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission'</article-title>
          ,
          <source>in Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '15</source>
          , p.
          <fpage>1721</fpage>
          -
          <lpage>1730</lpage>
          , New York, NY, USA, (
          <year>2015</year>
          ).
          <article-title>Association for Computing Machinery</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tianqi</given-names>
            <surname>Chen</surname>
          </string-name>
          and Carlos Guestrin, '
          <article-title>Xgboost: A scalable tree boosting system'</article-title>
          ,
          <source>in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining</source>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          , (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Yuanfang</given-names>
            <surname>Chen</surname>
          </string-name>
          , Liu Ouyang, Sheng Bao,
          <string-name>
            <given-names>Qian</given-names>
            <surname>Li</surname>
          </string-name>
          , Lei Han, Hengdong Zhang, Baoli Zhu, Ming Xu, Jie Liu, Yaorong Ge, and Shi Chen, '
          <article-title>An interpretable machine learning framework for accurate severe vs nonsevere covid-19 clinical type classification'</article-title>
          , medRxiv, (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ezekiel</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Emanuel</surname>
            , Govind Persad, Ross Upshur, Beatriz Thome, Michael Parker, Aaron Glickman, Cathy Zhang, Connor Boyle,
            <given-names>Maxwell</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>and James P.</given-names>
          </string-name>
          <string-name>
            <surname>Phillips</surname>
          </string-name>
          , '
          <article-title>Allocating medical resources in the time of covid-19'</article-title>
          ,
          <source>New England Journal of Medicine</source>
          ,
          <volume>382</volume>
          (
          <issue>22</issue>
          ),
          <year>e79</year>
          , (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ezekiel</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Emanuel</surname>
            , Govind Persad, Ross Upshur, Beatriz Thome, Michael Parker, Aaron Glickman, Cathy Zhang, Connor Boyle,
            <given-names>Maxwell</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>and James P.</given-names>
          </string-name>
          <string-name>
            <surname>Phillips</surname>
          </string-name>
          , '
          <article-title>Fair allocation of scarce medical resources in the time of covid-19'</article-title>
          ,
          <source>New England Journal of Medicine</source>
          ,
          <volume>382</volume>
          (
          <issue>21</issue>
          ),
          <fpage>2049</fpage>
          -
          <lpage>2055</lpage>
          , (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Frank</given-names>
            <surname>Stefan</surname>
          </string-name>
          <string-name>
            <given-names>Heldt</given-names>
            , Marcela P Vizcaychipi, Sophie Peacock, Mattia Cinelli,
            <surname>Lachlan</surname>
          </string-name>
          <string-name>
            <surname>McLachlan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Andreotti</surname>
          </string-name>
          , Stojan Jovanovic, Robert Durichen, Nadezda Lipunova,
          <article-title>Robert A Fletcher,</article-title>
          and
          <string-name>
            <surname>Anne</surname>
          </string-name>
          et al Hancock,
          <article-title>'Early risk assessment for covid-19 patients from emergency department data using machine learning'</article-title>
          , medRxiv, (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Scott</surname>
            <given-names>M Lundberg</given-names>
          </string-name>
          , Gabriel Erion, Hugh Chen, Alex DeGrave,
          <string-name>
            <surname>Jordan M Prutkin</surname>
          </string-name>
          , Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and
          <string-name>
            <surname>Su-In</surname>
            <given-names>Lee</given-names>
          </string-name>
          , '
          <article-title>Explainable ai for trees: From local explanations to global understanding'</article-title>
          , arXiv preprint arXiv:
          <year>1905</year>
          .
          <volume>04610</volume>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Scott</surname>
            <given-names>M Lundberg</given-names>
          </string-name>
          and
          <string-name>
            <surname>Su-In</surname>
            <given-names>Lee</given-names>
          </string-name>
          , '
          <article-title>A unified approach to interpreting model predictions'</article-title>
          ,
          <source>in Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , eds., I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Garnett</surname>
          </string-name>
          ,
          <volume>4765</volume>
          -
          <fpage>4774</fpage>
          , Curran Associates, Inc., (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Lingzhong</surname>
            <given-names>Meng</given-names>
          </string-name>
          , Haibo Qiu, Li Wan,
          <source>Yuhang Ai</source>
          , Zhanggang Xue, Qulian Guo, Ranjit Deshpande, Lina Zhang, Jie Meng, Chuanyao Tong, Hong Liu, and Lize Xiong, '
          <article-title>Intubation and Ventilation amid the COVID-19 Outbreak: Wuhan's Experience</article-title>
          .', Anesthesiology,
          <volume>132</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1317</fpage>
          -
          <lpage>1332</lpage>
          , (jun
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Robert</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Truog</surname>
          </string-name>
          , Christine Mitchell, and
          <string-name>
            <surname>George</surname>
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Daley</surname>
          </string-name>
          , '
          <article-title>The toughest triage - allocating ventilators in a pandemic'</article-title>
          ,
          <source>New England Journal of Medicine</source>
          ,
          <volume>382</volume>
          (
          <issue>21</issue>
          ),
          <fpage>1973</fpage>
          -
          <lpage>1975</lpage>
          , (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Akhil</surname>
            <given-names>Vaid</given-names>
          </string-name>
          , Sulaiman Somani, Adam J Russak,
          <string-name>
            <surname>Jessica K De Freitas</surname>
          </string-name>
          , and Fayzan F et al. Chaudhry, '
          <article-title>Machine learning to predict mortality and critical events in covid-19 positive new york city patients'</article-title>
          , medRxiv, (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>