<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Temporal Feature Selection for Characterizing Antimicrobial Multidrug Resistance in the Intensive Care Unit</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>O´scar Escudero-Arnanz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Inmaculada Mora-Jime´nez</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergio Mart´ınez-Ag u¨ero</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joaqu´ın A´ lvarez-Rodr´ıguez</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cristina Soguero-Ruiz</string-name>
          <email>cristina.soguero@urjc.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University</institution>
          ,
          <addr-line>Madrid 28943</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University</institution>
          ,
          <addr-line>Madrid 28943, Spain</addr-line>
          ,
          <institution>Copyright © 2020 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0). This volume is published and copyrighted by its editors. Advances in Artificial Intelligence for Healthcare</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University</institution>
          ,
          <addr-line>Madrid 28943, Spain, inmac-</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Signal Theory and Communications, Telematics and Computing Systems, Rey Juan Carlos University</institution>
          ,
          <addr-line>Madrid 28943, Spain, ser-</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Intensive Care Department, University Hospital of Fuenlabrada</institution>
          ,
          <addr-line>Madrid 28942</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The emergence and increase of antimicrobial multidrug resistance (AMR) is a demographic and economic problem for current health systems. AMR is particularly problematic in clinical units such as the intensive care unit (ICU), where the risk of infection is high, principally due to the extensive use of antimicrobials and invasive devices. In this work, we propose the use of different temporal feature selection and classification approaches to ascertain the most informative features and extract knowledge for characterizing AMR in the ICU. For this purpose, a set of demographic and temporal features such as antibiotics taken daily by the patient and the use of mechanical ventilation are considered. According to the results obtained in this work, it could be concluded that temporal features such as mecanic ventilation provide powerful insights to predict AMR in ICU.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The discovery of antibiotics and their subsequent use in the clinical
practice represented a great scientific advance, improving the
treatment of infectious diseases and thus saving millions of lives [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
However, the excessive and incorrect use of antibiotics is
contributing a downturn in their effectiveness against bacterial infections,
caused by mutations and the acquisition of genetic information from
other germs [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This fact makes infection control difficult and
increases the morbidity and mortality of previously treatable infectious
diseases such as malaria or acute respiratory diseases [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        The impact of antimicrobial multidrug resistance (AMR) can
cause an economic burden in hospitals and in the healthcare
systems, whose real outcomes still remain unknown. Following the
report by the World Health Organisation (WHO), it is estimated an
increase in deaths by 2050 caused by antimicrobial resistance, mainly
affecting countries such as Africa and Asia [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This is a growing
problem which needs to be alleviated to avoid the consequences
that this could cause. In addition to the demographic effects, the
increase in antimicrobial multidrug resistance, this is the resistance
of a single bacterium to more than one antibiotic, has a major
economic impact, resulting in loss to the world economy of
approximately 7% of the Gross Domestic Product by 2050 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. From an
economic viewpoint, patients infected with antimicrobial resistant
bacteria present a higher cost for the healthcare system in
comparison to patients who are susceptible to microbial infection [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This
is caused by the increasing difficulty in treating resistant organisms,
making it necessary the breakthrough of new strategies to combat
antibiotic resistance. Previous studies have proposed initial analysis
based on machine learning models to determine the result
(susceptible/resistance) of the antibiogram (a test to measure the in vitro
activity of an antibiotic against a given bacterium, which is previously
isolated in the culture [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]) or to predict the probability of acquiring
a hospital-acquired infection (nosocomial infection), specifically in
the ICU [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Focusing on a hospital environment, antimicrobial resistance can
be acquired by any hospitalised patient, increasing the probability of
acquisition for patients admitted to the Intensive Care Unit (ICU).
The main reasons are the use of invasive devices, the intensity of
treatment and its duration, the high risk of transmission and
exposure to antibiotics. The ICU can be considered as the epicenter of
development of antimicrobial resistance due to the high rate of
nosocomial infections (20-30% of all ICU admissions) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However,
the period just before the patient is admitted to the ICU is
beginning to take great importance, caused by the increase in the
number of patients arriving in the ICU infected by multi-resistant
microorganisms [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. A culture is usually performed to assess bacteria
susceptibility/resistance to series of antibiotics. Firstly, an organic
sample from the patient (blood or urine samples, among others) is
obtained which allows the study of the microorganisms present in their
system. Then, the antibiogram is carried out. The result of the
antibiogram represents the pair antibiotic/sensibility. Therefore, based
on this results, we consider that patients did not acquired the
multiresistant bacteria in the ICU if the culture’s result is positive within
the first 48 hours of the patient’s admission, otherwise, the AMR
occur during the ICU stay.
      </p>
      <p>
        The excessive use of antimicrobials during the stay of patients in
the ICU (some studies corroborate that more than 60% of patients
take antibiotics during their ICU stay [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) along with other factors
discussed above, facilitate the emergence of AMR, making this
problem the target to be treated. We will study the daily use of antibiotics
and mechanical ventilation (MV) in the ICU at University Hospital
of Fuenlabrada, Madrid, Spain. The final aim consists in determining
the risk factors that best characterize the evolution of critical patients
as well as the relevance to identify patients with AMR. To this end,
we apply hypothesis tests, linear and non-linear learning algorithms.
      </p>
      <p>The rest of the paper is organized as follows. Section 2 introduces
the methods used for the temporal patient characterization. In
Section 3, a brief description of the data set is presented, while in
Section 4 the experimental work and prediction results are shown.
Finally, discussion and conclusions are presented in Section 5.
2</p>
    </sec>
    <sec id="sec-2">
      <title>METHODS</title>
    </sec>
    <sec id="sec-3">
      <title>Notation</title>
      <p>In this paper, each sample is a patient represented by a set of D
features, being each feature composed by a time series of T consecutive
time slots. Therefore, the data associated to the i-th patient can be
arranged in a feature matrix Xi = [xi1; xi2; : : : ; xiT ] 2 RD T . Where
the column vector xit contains the D features of the i-th patient in
the time slot t. Thus, xit can be represented as the column vector
xit = [xit;1; xit;2; ; xit;D]T , where [:]T denotes the transpose
opt
erator and xi;d shows the value of the d-th feature associated to the
i-th patient in the t-th time slot. Since we are tackling with a
binary classification task, we have considered the label ‘1’ to identify
patients with AMR, and the label ‘0’ to identify patients with
nonAMR. Therefore, the label (desired output) for the i-th patient is
defined by yi, whereas the output provided by the model is represented
as y^i.
2.1</p>
    </sec>
    <sec id="sec-4">
      <title>Feature Selection</title>
      <p>
        There are different methods for feature selection in the literature.
The goal is to eliminate features that may be noisy, irrelevant or
redundant when building a data-driven model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Also, selecting the
most important features can increase the knowledge and the model
interpretability. In this work, we want to select features based on
hypothesis tests. For each feature, our null hypothesis is that there is
no difference between the two populations (AMR patients and
nonAMR patients). If there is no evidence to rule out the null hypothesis,
then the tested feature is not selected. Since we are dealing with
binary and numerical features, we evaluate a test of proportions for the
first kind of features, and a two-sample Kolmogorov-Smirnov test for
the latter.
      </p>
      <p>
        Two-proportion z-test. This hypothesis test evaluates whether the
presence on a single feature differs in two populations [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The null
hypothesis states that there is no evidence of difference in the
proportion between both populations, whereas the opposite applies for
the alternative hypothesis.
      </p>
      <p>
        Two-sample Kolmogorov-Smirnov test. It is a hypothesis test based
on the empirical distribution function and used to estimate whether
values of the same feature in two populations are from the same
continuous distribution [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. An advantage of this test over parametric
test is the independence of the statistic from the expected frequency
distribution, depending only on the sample size.
2.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Imbalanced sampling</title>
      <p>
        In healthcare-related data sets, it is very common to deal with
imbalanced data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], i. e., one class predominates over the other. This
imbalance is a challenge for designing data-driven models, since
conventional approaches will mostly learn from the majority class and
lead to biased models, reducing the performance for the minority
class. Data-driven approaches tend to learn better the mapping of
patients belonging to the majority class (far more numerous) than
that of the minority class. To tackle this challenge, several strategies
could be followed [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this work, we followed a random
undersampling strategy [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] with no replacement for the majority class.
The final sample size is such that the class frequency is similar. Thus,
the number of patients of the majority class is matched before
training the model according to the number of patients of the minority
class. The undersampling process and subsequent model training is
repeated several times not to be conditioned to a particular
subsampling, providing statistics on the performance. We benchmark the
results obtained with random undersmpling with a synthetic minority
oversampling technique (SMOTE), which consists of oversampling
the examples in the minority class [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
2.3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Classification Approaches</title>
      <p>
        Classification approaches encompasses statistical techniques to build
models based on the underlying relationships among data. The set
of N available samples is split into two independent subsets, named
training set and test set. The former is used to create the classifier
following a learning process, whereas the latter is used to evaluate
the performance of the built model. Normally, the 70% of samples
are randomly assigned to the training set and the rest to the test [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
2.3.1
      </p>
      <sec id="sec-6-1">
        <title>Logistic Regression</title>
        <p>
          The model provided by Logistic Regression (LR) is a linear
combination of the different features. Despite its name, it is a classification
approach since the result of the linear combination is the input to
a logistic function. To carry out the linear combination of the
features, a set of coefficients wiid=1 should be found by optimizing a
binary cross-entropy cost function. In this work, we considered a
regularized term in the cost function, in particular the Ridge
regularization [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for preventing the model from overfitting. To find an
appropriate value for the hyperparameter weighting the penalization
term in the cost function, named penalty coefficient C &gt; 0, we
followed a 5 fold cross-validation approach on the training set.
2.3.2
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>Decision Trees</title>
        <p>
          Decision trees (DT) are non-parametric classifiers which can be
graphically represented in a tree shape as a hierarchical structure
starting from a root node [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. For building the tree, a recursive
splitting process is carried out dividing the decision space into subspaces
based on a criteria related to entropy or Gini index. In this work,
we have chosen the Gini criterion to make the splitting process [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
When a node is created, a region in the feature space is splitted in two
parts. A label is assigned to each partition according to the majority
class among the training samples in that particular partition. One
advantage of DT is the model interpretability, that partly relies on the
fact that the most discriminative features are closest to the root node,
what implicitly could be considered as a feature selection process.
        </p>
        <p>
          In this work we considered DT built following the classification
and regression tree algorithm named CART [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], since it has been
extensively used in the literature when dealing with heterogeneous
features (numerical and categorical).
3
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>DATASET DESCRIPTION AND TEMPORAL</title>
    </sec>
    <sec id="sec-8">
      <title>FEATURES</title>
      <p>In this work, an anonymized dataset provided by the University
Hospital of Fuenlabrada (UHF) in Madrid (Spain) has been analysed.
This dataset contains demographic and clinical features of 2889
patients admitted in the ICU of the UHF during a period of 13 years,
from 2004 to 2016. The goal is leverage these data to
characterize AMR in the ICU. From a clinical viewpoint, clinicians at UHF
considered that patients with a positive culture (presence of
multiresistant germs) in the first 48 hours, had acquired the AMR before
their ICU admission. On the contrary, we considered that patients
with a positive culture after the early 48 hours of their admission,
had acquired the AMR during their ICU stay. Therefore, 507 of the
total number of patients acquired antimicrobial resistance, of which
171 (33.73%) acquired AMR before their ICU admission and 336
(66.27%) during their ICU stay. The average age of AMR patients
is 62.39 years, and 59.29 for non-AMR patients. In both cases, the
standard deviation is high (13.00 and 16.02, respectively). Regarding
gender, the percentage of men is higher for both AMR and non-AMR
patients (63.71% and 61.13%, respectively).</p>
      <p>The dataset has been preprocessed to characterize the evolution of
the patient’s health status by a set of features suitable to feed the
predictive model inputs. Thus, the d-th temporal feature corresponding
to the i-th patient is represented by a a row vector associated to a T
days time window, and it is given by: xi;d = [xi1;d; xi2;d; ; xiT;d],
with d = 1; ; D. In this work, we have considered T = 7 time
slots, i.e, the temporal characterization of a patient has been done
in a 7-days time window, with t0 the first 24 hours from the ICU
admission for the non-AMR patients. Regarding AMR patients, the
time slot t0 represents the time slot furthest from the first positive
culture, and therefore, closest to the ICU admission. Since the length
of the ICU stay can be shorter than 7 days for some patients, we
created a new binary feature, called mask, which takes a value of
‘1’ if the patient was in the ICU at this time slot, or ‘0’ otherwise.
The upper panel in Fig. 1 illustrates ficticious values for the mask
and the D features associated to one AMR patient. In this example,
since the culture flagged as positive the fifth day since the patient’s
ICU admission, all features assigned to t0 and t1 have null values.
The bottom panel in Fig. 1 represents the hypothetical values for the
mask and features associated to a potential non-AMR patient with a
stay of at least 7 days, being t0 the time slot nearest to the patient’s
ICU admission.</p>
      <p>The features represented as xi;d in Fig. 1 are associated to the
family of antibiotics taken by the patient (23 features), as well
as to the mechanical ventilation (MV), to the result of the
albumin blood test and to the number of times this blood test was
required. The families of the antibiotics the patient can take are
the following: Aminoglycosides (AMG), Antifungals (ATF),
Carbapenemes (CAR), 1st generation Cephalosporins (CF1), 2nd
generation Cephalosporins (CF2), 3rd generation Cephalosporins (CF3),
4th generation Cephalosporins (CF4), unclassified antibiotics
(Others), Glycyclines (GCC),Glycopeptides (GLI), Lincosamides (LIN),
Lipopeptides (LIP), Macrolides (MAC), Monobactamas (MON),
Nitroimidazolics (NTI), Miscellaneous (OTR), Oxazolidinones (OXA),
Broad-Spectrum Penicillins (PAP), Penicillins (PEN), Polypeptides
(POL), Quinolones (QUI), Sulfamides (SUL) and Tetracyclines
(TTC). Regarding the feature associated to MV, for each time slot
we have considered the number of hours the patient was assisted with
mechanical ventilation. The use of these features is supported by the
fact that the incorrect and excessive use of antibiotics or external
devices are one of the main causes for the AMR onset. In addition, two
demographic features (not time-dependent), the age and the gender
of the patient, have been used as input of the models.</p>
      <p>We present in Fig. 2 the percentage of AMR and non-AMR
patients who take each family of antibiotics. Note that this percentage
is similar for some families of antibiotics such as Broad-Spectrum
Penicillins, Quinolones and Lipopeptides. However, the percentage
of Antifungals, Glycopeptides and Carbapenemes is higher for AMR
patients, while non-AMR patients present a higher percentage of
Penicillins, among others.</p>
    </sec>
    <sec id="sec-9">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>
        The goal of this work was twofold. On the one hand, a feature
selection strategy was applied to find the most relevant features to
discriminate between AMR and non-AMR patients. On the other hand,
the chosen features were considered to evaluate the potential of
different prediction models when classifying AMR and non-AMR
patients. Towards that end, we start this section by discussing the
experimental set-up, then we present the feature selection process and
the prediction results.
The methodology to select relevant features and train different
classifiers is as follows. First, patients in the dataset were randomly
separated, assigning the 70% to the train set and the 30% to the test set [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
In order to reduce the potential bias in the results produced by good
or bad partitions, we repeat this process 1000 times. The metrics used
for measuring the performance of the classifiers are the mean and the
standard deviation of the Accuracy, Specificity, Sensitivity, F1-score
and the Area Under the Curve (AUC).
      </p>
      <p>For tuning hyperparameters, a 5-fold cross-validation strategy was
considered in the training set. For the LR models the hyperparameter
used was the penalty coefficient C 2 f0.001, 0.005, 0.01, 0.05, 0.1,
0.5, 0.75, 1.0g. The hyperparameters associated to the decision tree
were the depth of the tree (ranging from 4 to 22) and the minimum
of samples per leaf (between 6 and 15).
We performed a hypothesis test for each time slot and for all features
described in Section 3, except for demographic features due to the
non-dependence in time of this kind of features. For both imbalanced
and balanced data, we considered the p-value provided by the
twoproportion z-test for antibiotics and by the two-sample
KolmogorovSmirnov test for MV and albumin (see Table 1). In the case of
imbalanced data, we determined as significant features those with a p-value
&lt; 0.1. When considering balanced datasets, we perform N = 1000
subsamplings of the majority class and obtain the median of the
pvalues, selecting those features such that the median of the p-values
is lower than 0.1.</p>
      <p>To perform the experiments, we have used those features selected
by the above tests when using balanced subsets, together with the
demographic features of the patient. We have selected those
features that are statistically significant (p-value ¡ 0.1) during the first
48 hours (t0 and t1), from 48 hours (t2,t3,t4,t5, and t6) onwards or
throughout the time window (from t0 to t6). According to these
conditions, we have obtained the following features: all time slots of
ATF, PEN, OXA, and Albumin (Value), from time slot t2 to t6 for
Others and MV (hours), and the first two time slots for AMG, CF3,
GLI, NTI, QUI and Albumin (Count). Some of these features are
clinically relevant. For example, QUI and AMG are antimicrobial
families employed to treat the pseudomona aeruginosa infections,
OXA family are the main antimicrobial given to tackle the
staphylococcus aureus (both pseudomonas aeruginosa and staphylococcus
aureus are the most common MDR bacteria). The mechanical
ventilation and the level of albumin in the blood are related to the
patient’s state of health. The p-values associated to these features and
time slots are in bold in Table 1.
4.3</p>
    </sec>
    <sec id="sec-10">
      <title>Prediction Results</title>
      <p>In this subsection, the results of predicting whether a patient will
be considered AMR or non-AMR are presented in Table 2. For the
prediction, we considered both a linear (LR) and non-linear (DT)
models, designed using the features selected in Subsection 4.2.</p>
      <p>Several conclusions can be obtained from Table 2, where the mean
and standard deviation of several performance measurements on the
test subsets of 1000 subsamplings are provided using random
undersampling and SMOTE to balance the data. In general, the LR model
(a linear model) achieves better results, especially in terms of
Sensitivity (69.97 3.68). On the contrary, better results in term of
Specificity (86.13 1.87) are obtained when considering DT (non-linear
model). The results obtained through the use of SMOTE for LR
improve, except Sensitivity. On the other hand, in DT, better results
are obtained for Specificity and Accuracy, while the other metrics
worsen.</p>
      <p>Figure 3 shows the importance of the features provided by 1000
different models when considering LR and DT. To estimate the
feature importance in LR, we have considered the absolute values of
the weights associated to the features, while we have used the Gini
index in DT. The results are presented in box-plots, with features
sorted increasingly according to median of the p-values provided
by 1000 models. Features with the highest importance are
approximately the same in both classifiers, highlighting MV in some time
slots, the blood albumin value and the age of the patient.</p>
      <p>0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0.0
(a)
0.4
0.3
0.2
0.1
0.0
(b)
Nowadays, AMR has become a real and growing problem due to
the inappropriate use of antimicrobials. Bacteria that were previously
standard deviation of several performance measurements (Specificity, Sensitivity, Accuracy, F1-score and AUC) on 1000 test sets when
designing a lineal model (LR) and a non linear model (DT).</p>
      <p>Training Strat.
easily treatable have now become an issue difficult to deal with,
especially in the ICUs. In these units, AMR has created a great impact
on morbidity, hospital costs, and sometimes patient survival.</p>
      <p>It is necessary to be aware of the growing problem caused by the
expansion of AMR, for which new research, efforts, and approaches
are needed to prevent further spread of AMR. The use of automatic
learning methods is a very useful tool to solve problems related to the
clinical environment following a data-driven strategy. These methods
allows us to reduce the time of detection of infectious diseases,
resulting in a reduction in the number of deaths as well as in health
economic costs.</p>
      <p>In this work, we proposed the use of feature selection and machine
learning approaches to extract knowledge and predict the appearance
of AMR of patients admitted in the ICU. Features such as the
performance provided by LR (71.71% AUC) suggests that the analysis
presented in this paper could be a first step to identify the bacteria
appearance and isolate the patients at risk of AMR.</p>
      <p>As future work, we propose the analysis of more features related to
the patients such as blood samples or vital signs, as wells as the use
of more advanced machine learning methods, as for example, long
short-term memory networks which are capable of learning
longterm dependencies.</p>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work has been partly supported by the Institute of Health
Carlos III, Spain (grant DTS 17/00158), by the Spanish Ministry of
Economy, Industry and Competitiveness under the Research Project
Klinilycs (TEC2016-75361-R), by the Science and Innovation
Ministry Grants AAVis-BMR (PID2019-107768RA-I00) and BigTheory
(PID2019-106623RB-C41), by Project Ref. F656 financed by Rey
Juan Carlos University, by the Young Researchers R&amp;D Project Ref.
2020-661, financed by Rey Juan Carlos University and Community
of Madrid (Spain), and by the Youth Employment Initiative (YEI)
R&amp;D Project Ref. TIC-11649 financed by the Community of Madrid
(Spain).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Bassetti</surname>
          </string-name>
          , Garyphallia Poulakou, Etienne Ruppe, Emilio Bouza,
          <string-name>
            <surname>Sebastian J Van Hal</surname>
          </string-name>
          , and Adrian Brink, '
          <article-title>Antimicrobial resistance in the next 30 years, humankind, bugs and drugs: a visionary approach'</article-title>
          ,
          <source>Intensive care medicine</source>
          ,
          <volume>43</volume>
          (
          <issue>10</issue>
          ),
          <fpage>1464</fpage>
          -
          <lpage>1475</lpage>
          , (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W</given-names>
            <surname>Baumgartner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P</given-names>
            <surname>Weiß</surname>
          </string-name>
          , and H Schindler,
          <article-title>'A nonparametric test for the general two-sample problem'</article-title>
          ,
          <source>Biometrics</source>
          ,
          <fpage>1129</fpage>
          -
          <lpage>1135</lpage>
          , (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L</given-names>
            <surname>Breiman</surname>
          </string-name>
          ,
          <article-title>JH Friedman, RA Olshen, and CJ Stone, Classification and Regression Trees</article-title>
          ,
          <source>Chapman and Hall</source>
          ,
          <year>1984</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Nele</given-names>
            <surname>Brusselaers</surname>
          </string-name>
          , Dirk Vogelaers, and Stijn Blot, '
          <article-title>The rising problem of antimicrobial resistance in the intensive care unit'</article-title>
          ,
          <source>Annals of intensive care</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <fpage>47</fpage>
          , (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Rich</given-names>
            <surname>Caruana</surname>
          </string-name>
          , Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad, '
          <article-title>Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission'</article-title>
          ,
          <source>in Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining</source>
          , pp.
          <fpage>1721</fpage>
          -
          <lpage>1730</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nitesh</surname>
            <given-names>V Chawla</given-names>
          </string-name>
          , Kevin W Bowyer, Lawrence O Hall, and
          <string-name>
            <given-names>W Philip</given-names>
            <surname>Kegelmeyer</surname>
          </string-name>
          , '
          <article-title>Smote: synthetic minority over-sampling technique'</article-title>
          ,
          <source>Journal of artificial intelligence research</source>
          ,
          <volume>16</volume>
          ,
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          , (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Sara</surname>
            <given-names>E Cosgrove</given-names>
          </string-name>
          , '
          <article-title>The relationship between antimicrobial resistance and patient outcomes: mortality, length of hospital stay, and health care costs'</article-title>
          ,
          <source>Clinical Infectious Diseases</source>
          ,
          <volume>42</volume>
          (
          <issue>Supplement 2</issue>
          ),
          <fpage>S82</fpage>
          -
          <lpage>S89</lpage>
          , (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] Jose´
          <string-name>
            <given-names>F D</given-names>
            <surname>´</surname>
          </string-name>
          ıez-Pastor,
          <article-title>Juan J Rodr´ıguez, Ce´sar Garc´ıa-</article-title>
          <string-name>
            <surname>Osorio</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ludmila</surname>
            <given-names>I Kuncheva,</given-names>
          </string-name>
          '
          <article-title>Random balance: ensembles of variable priors classifiers for imbalanced data'</article-title>
          ,
          <source>Knowledge-Based Systems</source>
          ,
          <volume>85</volume>
          ,
          <fpage>96</fpage>
          -
          <lpage>111</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Jianqing</given-names>
            <surname>Fan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Runze</given-names>
            <surname>Li</surname>
          </string-name>
          , '
          <article-title>Statistical challenges with high dimensionality: Feature selection in knowledge discovery'</article-title>
          ,
          <source>arXiv preprint math/0602133</source>
          , (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Infectious Diseases Society of America (IDSA), 'Combating antimicrobial resistance: policy recommendations to save lives'</article-title>
          ,
          <source>Clinical Infectious Diseases</source>
          ,
          <volume>52</volume>
          (
          <issue>suppl 5</issue>
          ),
          <fpage>S397</fpage>
          -
          <lpage>S428</lpage>
          , (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G</given-names>
            <surname>James</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Witten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Hastie</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          ,
          <article-title>An Introduction to Statistical Learning with Applications in</article-title>
          R, Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Mart</surname>
          </string-name>
          <article-title>´ınez-Agu¨ero, Inmaculada Mora-Jime´nez, Jon Le´ridaGarc´ıa, Joaqu´ın A´lvarez-Rodr´ıguez, and Cristina Soguero-Ruiz, 'Machine learning techniques to identify antimicrobial resistance in the intensive care unit'</article-title>
          ,
          <source>Entropy</source>
          ,
          <volume>21</volume>
          (
          <issue>6</issue>
          ),
          <fpage>603</fpage>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Marc</given-names>
            <surname>Mendelson</surname>
          </string-name>
          and Malebona Precious Matsoso, '
          <article-title>The world health organization global action plan for antimicrobial resistance'</article-title>
          ,
          <source>SAMJ: South African Medical Journal</source>
          ,
          <volume>105</volume>
          (
          <issue>5</issue>
          ),
          <fpage>325</fpage>
          -
          <lpage>325</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>J. Ross</surname>
            <given-names>Quinlan</given-names>
          </string-name>
          , '
          <article-title>Induction of decision trees'</article-title>
          ,
          <source>Machine learning</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <fpage>81</fpage>
          -
          <lpage>106</lpage>
          , (
          <year>1986</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Paz</given-names>
            <surname>Revuelta-Zamorano</surname>
          </string-name>
          ,
          <article-title>Alberto Sa´nchez, Jose´ Luis Rojo- A´lvarez, Joaqu´ın A´lvarez-Rodr´ıguez, Javier Ramos-Lo´pez, and Cristina Soguero-Ruiz, 'Prediction of healthcare associated infections in an intensive care unit using machine learning and big data tools'</article-title>
          ,
          <source>in XIV Mediterranean Conference on Medical and Biological Engineering and Computing</source>
          <year>2016</year>
          , pp.
          <fpage>840</fpage>
          -
          <lpage>845</lpage>
          . Springer, (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Pilar</surname>
            <given-names>Talo´</given-names>
          </string-name>
          <article-title>n-</article-title>
          <string-name>
            <surname>Ballestero</surname>
          </string-name>
          ,
          <article-title>Lydia Gonza´lez-</article-title>
          <string-name>
            <surname>Serrano</surname>
          </string-name>
          ,
          <article-title>Cristina SogueroRuiz, Sergio Mun˜oz-</article-title>
          <string-name>
            <surname>Romero</surname>
          </string-name>
          , and Jose´ Luis Rojo- A´lvarez, '
          <article-title>Using big data from customer relationship management information systems to determine the client profile in the hotel sector'</article-title>
          ,
          <source>Tourism Management</source>
          ,
          <volume>68</volume>
          ,
          <fpage>187</fpage>
          -
          <lpage>197</lpage>
          , (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Jiliang</surname>
            <given-names>Tang</given-names>
          </string-name>
          , Salem Alelyani, and Huan Liu, '
          <article-title>Feature selection for classification: A review', Data classification: Algorithms and</article-title>
          applications,
          <volume>37</volume>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Unemo</surname>
          </string-name>
          and William M Shafer, '
          <article-title>Antimicrobial resistance in neisseria gonorrhoeae in the 21st century: past, evolution</article-title>
          , and future',
          <source>Clinical microbiology reviews</source>
          ,
          <volume>27</volume>
          (
          <issue>3</issue>
          ),
          <fpage>587</fpage>
          -
          <lpage>613</lpage>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Dirk</surname>
            <given-names>Vogelaers</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Bels</surname>
          </string-name>
          , et al.,
          <article-title>'Patterns of antimicrobial therapy in severe nosocomial infections: empiric choices, proportion of appropriate therapy, and adaptation rates-a multicentre, observational survey in critically ill patients'</article-title>
          ,
          <source>International journal of antimicrobial agents</source>
          ,
          <volume>35</volume>
          (
          <issue>4</issue>
          ),
          <fpage>375</fpage>
          -
          <lpage>381</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Show-Jane Yen</surname>
          </string-name>
          and
          <string-name>
            <surname>Yue-Shi</surname>
            <given-names>Lee</given-names>
          </string-name>
          , '
          <article-title>Under-sampling approaches for improving prediction of the minority class in an imbalanced dataset'</article-title>
          ,
          <source>in Intelligent Control and Automation</source>
          ,
          <volume>731</volume>
          -
          <fpage>740</fpage>
          , Springer, (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>