<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Classification Models in Intensive Care Outcome Prediction-can we improve on current models?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicholas A. Barnes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lynnette A. Hunt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael M. Mayo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Corresponding Author: Nicholas A. Barnes.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Waikato</institution>
          ,
          <addr-line>Hamilton</addr-line>
          ,
          <country country="NZ">New Zealand</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Statistics, University of Waikato</institution>
          ,
          <addr-line>Hamilton</addr-line>
          ,
          <country country="NZ">New Zealand</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Intensive Care Unit, Waikato Hospital</institution>
          ,
          <addr-line>Hamilton</addr-line>
          ,
          <country country="NZ">New Zealand</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>5</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>Classification models (“machine learners” or “learners”) were developed using machine learning techniques to predict mortality at discharge from an intensive care unit (ICU) and evaluated based on a large training data set from a single ICU. The best models were tested on data on subsequent patient admissions. Excellent model performance (AUCROC (area under the receiver operating curve) =0.896 on a test set), possibly superior to a widely used existing model based on conventional logistic regression models was obtained, with fewer perpatient data than that model.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Intensive care clinicians use explicit judgement and heuristics to formulate
prognoses as soon as reasonable after patient referral and admission to an intensive care
unit [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Models to predict outcome in such patients have been in use for over 30 years [
        <xref ref-type="bibr" rid="ref3">2</xref>
        ]
but are considered to have insufficient discriminatory power for individual decision
making in a situation where patient variables that are difficult or impossible to
measure may be relevant. Indeed even variables that have little or nothing to do with the
patient directly (such as bed availability or staffing levels [
        <xref ref-type="bibr" rid="ref4">3</xref>
        ]) may be important in
determining outcome.
      </p>
      <p>There are further challenges for model development. Any model used should be
able to deal with the problem of class imbalance, which refers in this case to the fact
that mortality should be much less common than survival. Many patient data are
probably only loosely or indeed not related to outcome and many are highly
correlated. For example, elevated measurements of serum urea, creatinine, urine output,
diagnosis of renal failure and use of dialysis will all be closely correlated.</p>
      <p>Nevertheless, models are used to risk adjust for comparison within an institution
over time or between institutions, and model performance is obviously important if
this is to be meaningful. It is also likely that a model with excellent performance
could augment clinical assessment of prognosis. Furthermore, a model that performs
well while requiring fewer data would be helpful as accurate data acquisition is an
expensive task.</p>
      <p>
        The APACHE III-J (Acute Physiology and Chronic Health Evaluation revision
IIIJ [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ]) model is used extensively within Australasia by the Centre for Outcomes
Research of the Australian and New Zealand Intensive Care Society (ANZICS) and a
good understanding of its local performance is available in the published literature
[
        <xref ref-type="bibr" rid="ref5">4</xref>
        ]. It should be noted that death at hospital discharge is the outcome variable usually
considered by these models. Unfortunately the coefficients for all variables for this
model are no longer in the public domain so direct comparison with new models is
difficult. The APACHE (Acute Physiology and Chronic Health Evaluation) models
are based largely on baseline demographic and illness data and physiological
measurements taken within the first day after ICU admission.
      </p>
      <p>This study aims to explore machine learning methods that may outperform the
logistic regression models that have previously been used.</p>
      <p>
        The reader may like to consult a useful introduction to the concepts and practice of
machine learning [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ] if terms or concepts are unfamiliar.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>The study is comprised of three parts:
1. An empirical exploration of raw and processed admission data with a variety of
attribute selection methods, filters, base classifiers and metalearning techniques
(which are overarching models that have other methods nested within them) that
were felt to be suitable to develop the best classification models. Metamodels and
base classifiers may be nested within other metamodels and learning schemes can
be varied in very many ways .These experiments are represented below in Figure 1
where we used up to two metaclassifiers with up to two base classifiers nested
within a metaclassifier.</p>
      <sec id="sec-2-1">
        <title>Choose</title>
      </sec>
      <sec id="sec-2-2">
        <title>Dataset</title>
      </sec>
      <sec id="sec-2-3">
        <title>Metamodel 1</title>
      </sec>
      <sec id="sec-2-4">
        <title>Metamodel 2</title>
      </sec>
      <sec id="sec-2-5">
        <title>Base Classifier (s)</title>
      </sec>
      <sec id="sec-2-6">
        <title>Evaluate</title>
      </sec>
      <sec id="sec-2-7">
        <title>Classifier</title>
      </sec>
      <sec id="sec-2-8">
        <title>Results</title>
        <p>2. Further testing with the best performing data set (full unimputed training set) and
learners with manual hyperparameter setting. A hyperparameter is a particular
model configuration that is selected by the user, either manually or following an
automatic tuning process. This is represented in a schematic below:
3. Testing of the best models from phase 2 above on a new set of test data to better
understand generalizability of the models. This is depicted in Figure 3 below.
Matching
Test Set</p>
        <p>Four Best Models
based on 4
Evaluation</p>
        <p>Measures</p>
        <p>The training data for adult patients (8122 patients over 16 years of age) were
obtained from the database of a multidisciplinary ICU in a tertiary referral centre from a
period between July 2004 and July 2012.Data extracted were comprised of a
demographic variable (age), diagnostic category (with diagnostic coefficient from the
APACHE III-J scoring system, including ANZICS modifications), and an extensive
list of numeric variables relating to patient physiology and composite scores based on
these, along with the classification variable: either survival, or alternatively, death at
ICU discharge (as opposed to death at hospital discharge as in the APACHE models).
Much of the data collected is used in APACHE III-J model mentioned above, and
represents a subset of the data used in that model. Training data, prior to the
imputation process, but following discretization of selected variables are represented in
Table 1. Test data for the identical variable set were obtained from the same database for
the period July 2012 to March 2013.</p>
        <p>Of particular interest is that the data is clearly class imbalanced with mortality
during ICU stay of approximately 12%. This has important implications for modelling
the data.</p>
        <p>There were many strongly correlated attributes within the data sets. Many of the
model variables are collected as highest and lowest measures within twenty four
hours of admission to the ICU. Correlated variables may bring special problems with
conventional modelling including logistic regression. The extent of correlation is
demonstrated in Figure 4.</p>
        <p>Patterns of missing data are indicated in Table 1 and represented graphically in
Figure 5.</p>
        <p>Fig. 5. Patterns of missing data in the raw training set. Missing data is represented by red
colouration.</p>
        <p>
          Missing numeric data in the training set was imputed using multiple imputation
with the R program [
          <xref ref-type="bibr" rid="ref7">6</xref>
          ] and the R package Amelia [
          <xref ref-type="bibr" rid="ref8">7</xref>
          ], which utilises bootstrapping of
non-missing data followed by imputation by expectation maximisation. We initially
used the average of five multiple imputation runs.
        </p>
        <p>Using the last imputed set was also trialled, as it may be expected to be the most
accurate based on the iterative nature of the Amelia algorithm. No categorical data
were missing. Date of admission was discretized to the year of admission, age was
converted to months of age, and the diagnostic categories were converted to five to
eight (depending on study phase) ordinal risk categories by using coefficients from
the existing APACHE III-J risk model.</p>
        <p>A summary of data is presented below in Table 1.</p>
        <p>Phase 1 consisted of an exploration of machine learning techniques thought
suitable to this classification problem, and in particular those thought to be appropriate to
a class imbalanced data set. Attribute selection, examining the effect of using imputed
and unimputed data sets and application of a variety of base learners and
metaclassifiers without major hyperparameter variation occurred in this phase. The importance of
attributes was examined in multiple ways including using random forest methodology
for variable selection, using improvement in Gini index using particular attributes.
This information is displayed in figure 6.
riables used in the study are ranked by their contribution to Gini index.</p>
        <p>
          A comprehensive evaluation of all techniques is nearly impossible given the
enormous variety of techniques and the ability to combine up to several of these at a
time in any particular model. Techniques were chosen based on the likely success of
their application. WEKA [
          <xref ref-type="bibr" rid="ref9">8</xref>
          ] was used to apply learners and all models were
evaluated with tenfold cross validation. WEKA default settings were commonly used in
phase 1 and the details of these defaults are widely available [
          <xref ref-type="bibr" rid="ref10">9</xref>
          ]. Unless otherwise
stated all settings in all study phases were the default settings of WEKA for each
classifier or filter. Two results were used to judge overall model performance during
phase 1. These were:
1. Area under the receiver operating curve (AUC ROC)
2. Area under the precision recall curve (AUC PRC)
The results are presented in Table 3 in the results section.
        </p>
        <p>
          Phase 2 of our study involved training and evaluation on the same data sets with
learners that had performed well in phase 1. Hyperparameters were mostly selected
manually, as automatic hyperparameter selection in any software is limited and
hampered by a lack of explicitness. Class imbalance issues were addressed with
appropriate WEKA filters (spread subsample and SMOTE, a filter which generates a synthetic
data set to balance the classes [
          <xref ref-type="bibr" rid="ref11">10</xref>
          ]), or the use of cost sensitive learners [
          <xref ref-type="bibr" rid="ref12">11</xref>
          ]. Unless
otherwise stated in Table 3, WEKA default settings were used for each filter or
classifier. Evaluation of these models proceeded with tenfold cross-validation and the
results were examined in light of four measures:
1. Area under the receiver operating curve with 95% confidence intervals by the
method of Hanley and McNeill [
          <xref ref-type="bibr" rid="ref13">12</xref>
          ]
2. Area under the precision recall curve
3. Matthews correlation coefficient and,
4. F-measure
Additionally, scaling the quantitative variables by standardizing or normalizing the
data was explored as this is known to sometimes improve model performance [
          <xref ref-type="bibr" rid="ref14">13</xref>
          ].
The results of phase 2 are presented in Table 2 in the results section.
        </p>
        <p>Phase 3 involved evaluating the accuracy of the best classification models from phase
2 on a new test set of 813 patient admissions. Missing data in the test set were not
imputed. Results are shown in Table 3.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>Base
classifier
2
NA
NA</p>
      <p>ROC</p>
      <p>PRC
Preprocess
best of all models on any of the four classification methods is shaded in red to
emphasise that no one performance measure dominates a classifier’s overall utility.
NA
J 48
NA
NA
NA
NA
NA
J48
NA
NA
NA</p>
      <p>NA
0.901
(0.881,0.921)
NA
Spread
subsample
uniform
Spread
subsample
Spread
subsample
Spread
subsample
Spread
subsample
uniform
Spread
subsample
uniform
0.888
(0.864,0.912)</p>
      <p>Normalizing or standardizing the data did not improve model performance and
indeed tended to moderately worsen it.</p>
      <p>Table 4 presents the results of applying four of the best models from phase 2 on a
test data set of 813 patient admissions which should be from the same population
distribution (if date of admission is not a relevant attribute). Evaluation is based on
AUC ROC, AUC PRC, Matthews’s correlation coefficient and F-measure. These
evaluations were obtained by WEKA’s knowledge flow interface.
NA</p>
      <p>ROC
0.896
0.893
ROC-area under receiver operating characteristic curve
CI-confidence interval
PRC-area under precision-recall curve
MCC-Matthews correlation coefficient</p>
      <p>F-meas-F-measure
4</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>It is unrealistic to expect models to perfectly represent such a complex reality as
that of survival from critical illness. Perfect classification is impossible because of the
limitations of any combination of currently available measurements made on such
patients to accurately reflect survival potential. Patient factors such as attitudes
towards artificial support and presumably health practitioner and institution related
factors are important. Additionally non-patient related factors which may be purely
logistical will continue to thwart perfect prediction by any future model. For instance,
a patient may die soon after discharge from the ICU if a ward bed is available and
conversely will die within the ICU if a ward bed is not available and transfer cannot
proceed. Models currently employed generally consider death at hospital discharge,
but new factors that increase randomness can enter in the hospital stay following ICU
discharge, so problems are not necessarily decreased with this approach.</p>
      <p>The best models we have studied have excellent performance when evaluated
following tenfold cross validation in the single ICU setting with use of fewer data points
than the current gold standard model. Machine learning techniques usually make few
distributional assumptions about the data when compared with the traditional logistic
regression model. Missing data are often dealt with effectively with machine learning
techniques while complete cases are generally used in traditional general linear
modelling such as logistic regression. Clinical data will never be complete, as some data
will not be required for a given patient, while some patients may die prior to
collection of data which cannot subsequently be obtained. Imputation may be performed on
data prior to modelling but has limitations. It is interesting that models trained on
unimputed data tend to perform better than imputed data, both in phase 2 and with the
test set in phase 3.</p>
      <p>
        The best comparison we can make in the published literature is the work of Paul et
al [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ] which demonstrates that the AUC ROC of the APACHE-III-J model has varied
between 0.879 and 0.890 when applied to over half a million adult admissions to
Australasian ICUs between 2000 and 2009. Routine exclusions in this study included
readmissions, transfers to other ICUs, and missing outcome and other data, and
admission post coronary artery bypass grafting prior to introduction of the ANZICS
modification to APACHE-III-J for this category. None of these were exclusions in our
study. The Paul et al paper looks at outcome at hospital discharge, while ours
examines outcome at ICU discharge. For these reasons the results are not directly
comparable but our results for AUC ROC of up to 0.896 on a separate validation set
clearly demonstrate excellent model performance.
      </p>
      <p>
        The techniques associated with the best performance involve addressing class
imbalance (i.e. pre-processing data to create a dataset with similar numbers of those who
survive and those that die). This class imbalance is a well-known problem in
classification. Mortality data from any healthcare setting tend to be class imbalanced. Our
study shows that any approach to class imbalance in the data greatly enhance model
performance. Cost sensitive metalearners [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ], synthetic minority generation
techniques (SMOTE [
        <xref ref-type="bibr" rid="ref11">10</xref>
        ]) and creating a uniform class distribution by subsampling across
the data all improve model performance.
      </p>
      <p>
        A cost sensitive learner indicates a technique that reweights cases according to a
cost matrix that the user sets to reflect differing “cost” of misclassification of positive
and negative cases. This intuitively lends itself to the intensive care treatment process
where such a framework is likely implemented at least subconsciously by the
intensive care clinician. For instance the cost of clinically “misclassifying” a patient may
be substantial and clinicians would likely try hard to avoid this situation.
In our study, the ensemble learner random forests [
        <xref ref-type="bibr" rid="ref15">14</xref>
        ] with or without a technique to
address class imbalance tends to outperform many more complex metalearners, or
enhancements of single base classifiers such as bagging [
        <xref ref-type="bibr" rid="ref16">15</xref>
        ] and boosting [
        <xref ref-type="bibr" rid="ref17">16</xref>
        ].
Random forests involve generation of many different tree models, each of which splits the
cases based on different variables and a criterion to increase information gain. Voting
then occurs across the “forest” to decide on the best way to split the cases and this
produces the model. The term ensemble simply represents the fact that multiple
learners are involved, rather than a single tree. As many as 500 or 1000 trees are
commonly required before the error of the forest is at a minimum. The number of
variables to be considered by each tree may also be set to try and improve performance.
The other techniques that produced excellent results were rotation forests either alone,
with a cost sensitive classifier, or in combination with a technique known as
alternating decision tree. Alternating decision tree takes a “weak” classifier (such as a tree
classifier) and uses a technique similar to boosting to improve performance.
The reason extensive experimentation may be required to produce the best model is
attributed to Wolpert [
        <xref ref-type="bibr" rid="ref18">17</xref>
        ] and described as the “no free lunch theorem”, meaning that
there is no one single technique that will model the best in every given scenario. Of
course the same is true of any conventional statistical technique applied to
multidimensional problems. Data processing and model selection are crucial to performance
although if prediction alone is important, a pragmatic approach can be taken to the
usual statistical assumptions. Machine learning techniques are generally not a “black
box” approach however and deserve the same credibility as any older method, if
application is appropriate.
      </p>
      <p>Similarly, no single evaluation measure can summarize a classifier’s performance and
different model strengths and weaknesses may be more or less tolerable depending on
the circumstances of model use and hence a range of measures are usually presented
as we have done.</p>
      <p>There are several weaknesses to our study. It is clearly from a single centre and
may not generalize to other ICUs in other healthcare systems. Mortality remains a
crude measure of ICU performance but remains simple to measure and of great
relevance nevertheless. The existing gold standard models usually measure classification
of survival or death at hospital discharge, so are not necessarily directly comparable
to our models which measures survival or death at ICU discharge.</p>
      <p>
        We are unable to directly compare our models with what may be considered gold
standards as some of these (e.g. APACHE IV) are only commercially available, and
as mentioned before, even the details of APACHE-III-J are not in the public domain.
The best comparison involving Australasian data using APACHE-III-J comes from
the paper of Paul et al. [
        <xref ref-type="bibr" rid="ref5">4</xref>
        ] but as with all APACHE models, this predicts death at
hospital discharge. Additionally, re-admissions were excluded which may be a
significant factor beyond what are often relatively small numbers of re-admissions in any
given ICU, as re-admissions suffer a disproportionately high mortality.
      </p>
      <p>
        Exploration of the available hyperparameters of the many models examined has
been relatively limited. The ability to do this automatically, and explicitly or in a
reproducible way in WEKA and indeed any available software is limited although this
may be changing [
        <xref ref-type="bibr" rid="ref19">18</xref>
        ]. Yet minor changes to these hyperparameters may produce
meaningful enhancements in model performance. Tuning hyperparameters runs the
risk of overfitting a model, but we have tried to guard against this by testing the data
on a separate validation set.
      </p>
      <p>
        Likewise, the ability to combine models with the best characteristics [
        <xref ref-type="bibr" rid="ref20">19</xref>
        ], which is
becoming more common in prediction of continuous variables [
        <xref ref-type="bibr" rid="ref21">20</xref>
        ] is not yet easily
performed with the available software.
      </p>
      <p>We have not examined the calibration of our models. Good calibration is not
required for accurate classification. Accurate performance across all risk categories is
highly desirable in a model. Similarly, performance including calibration for different
diagnostic categories that may become more significant in an ICU’s case mix is not
accounted for.</p>
      <p>Modelling using imputed data in every phase of our study tends to show
inconsistent or suboptimal performance. It may be that imputation could be applied more
accurately by another approach that would improve model performance.</p>
      <p>
        The major current use of these scores is in quality improvement activities. Once a
score is developed which accurately quantitates risk, the expected number of deaths
may be compared to those observed [
        <xref ref-type="bibr" rid="ref22">21</xref>
        ]. The exact risk for a given integer valued
number of deaths may be derived from the Poisson binomial distribution and
compared to the number observed [
        <xref ref-type="bibr" rid="ref23">22</xref>
        ]. A variety of risk adjusted control charts can be
constructed with confidence intervals [
        <xref ref-type="bibr" rid="ref24">23</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We have presented alternative approaches to the classification problem involving
prediction of mortality at ICU discharge using machine learning techniques. Such
techniques may hold substantial advantage over traditional logistic regression
approaches and should be considered to replace these. Complete clinical data may be
unnecessary when using machine learning techniques, and in any case are frequently
not available. Out of the techniques studied, random forests seems to be the
modelling approach with the best performance and has an advantage that it is relatively easy
to conceptualise and implement with open source software. During model training a
method to address class imbalance should be used.
6</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]. Downar,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2013</year>
          , April 18).
          <article-title>Even without our biases, the outlook for prognos-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>http://ccforum.com/content/13/4/168</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]. Knaus WA, W. D. (
          <year>1981</year>
          ).
          <article-title>APACHE-acute physiology and chronic health evaluation: a physiologically based classification system</article-title>
          .
          <source>Crit Care Med</source>
          ,
          <fpage>591</fpage>
          -
          <lpage>597</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]. Tucker,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Patient volume, staffing, and workload in relation to riskadjusted outcomes in a random stratified sample of UK neonatal intensive care units: a prospective evaluation</article-title>
          .
          <source>Lancet</source>
          ,
          <volume>99</volume>
          -
          <fpage>107</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]. Paul, E.,
          <string-name>
            <surname>Bailey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Lint</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Pilcher</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Performance of APACHE III over time in Australia and New Zealand: a retrospective cohort study</article-title>
          .
          <source>Anaesthesia and Intensive Care</source>
          ,
          <fpage>980</fpage>
          -
          <lpage>994</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]. Domingos,
          <string-name>
            <surname>P.</surname>
          </string-name>
          (
          <year>2013</year>
          , May 6).
          <article-title>A few useful things to know about machine learning</article-title>
          . Available from Washington University: http://homes.cs.washington.edu/~pedrod/papers/cacm12.pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6
          <string-name>
            <given-names>]. R Core</given-names>
            <surname>Team</surname>
          </string-name>
          . (
          <year>2013</year>
          , April 25). Available from CRAN: http://www.Rproject.org/.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]. Honaker,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , &amp;
            <surname>Blackwell</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          (
          <year>2013</year>
          , April 25).
          <article-title>Amelia II: a program for missing data</article-title>
          .
          <source>Available from Journal of Statistical</source>
          Software: http://www.jstatsoft.org/v45/i07/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]. Hall,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Eibe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Holmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Pfahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            , &amp;
            <surname>Reutemann</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          (
          <year>2009</year>
          , 1).
          <source>The WEKA Data Mining Software: An Update. SIGKDD Explorations.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>[9]. Weka overview</article-title>
          . (
          <year>2013</year>
          , April 25). Available from Sourceforge: http://weka.sourceforge.net/doc/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]. Chawla,
          <string-name>
            <given-names>N. O.</given-names>
            ,
            <surname>Bowyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. W.</given-names>
            ,
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. O.</given-names>
            , &amp;
            <surname>Kegelmeyer</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. P.</surname>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>SMOTE: Synthetic Minority Over-sampling Technique</article-title>
          .
          <source>Journal of Ariticial Intelligence Research</source>
          ,
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]. Ling,
          <string-name>
            <given-names>C. X.</given-names>
            , &amp;
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <surname>V. S.</surname>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Cost-sensitive learning and the class imbalance problem</article-title>
          . In C. Sammat; G. Webb, editors.
          <source>Encyclopaedia of Machine Learning</source>
          . Springer.p.
          <fpage>231</fpage>
          -
          <lpage>235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]. Hanley,
          <string-name>
            <given-names>J.</given-names>
            , &amp;
            <surname>McNeil</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          (
          <year>1982</year>
          ).
          <article-title>The meaning and use of the area under a receiver operating characteristic (ROC) curve</article-title>
          . Radiology,
          <volume>29</volume>
          -
          <fpage>36</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [13]. Aksoy,
          <string-name>
            <given-names>S.</given-names>
            , &amp;
            <surname>Haralick</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. M.</surname>
          </string-name>
          (
          <year>2013</year>
          , May 20).
          <article-title>Feature Normalization and Likelihood-based Similarity Measures for Image Retrieval. Available from cs</article-title>
          .bilkent.edu: http://www.cs.bilkent.edu.tr/~saksoy/papers/prletters01_likelihood.pdf
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [14]. Breiman,
          <string-name>
            <surname>L.</surname>
          </string-name>
          (
          <year>2001</year>
          ).
          <source>Random Forests. Machine Learning</source>
          ,
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [15]. Breiman,
          <string-name>
            <surname>L.</surname>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>Bagging predictors</article-title>
          .
          <source>Machine Learning</source>
          ,
          <fpage>123</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [16]. Freund,
          <string-name>
            <given-names>Y.</given-names>
            , &amp;
            <surname>Schapire</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. E.</surname>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>Experiments with a new boosting algorithm</article-title>
          .
          <source>In Machine Learning:Proceedings of the Thirteenth International Conference on Machine Learning</source>
          , (pp.
          <fpage>148</fpage>
          -
          <lpage>156</lpage>
          ). San Francisco.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [17]. Wolpert,
          <string-name>
            <surname>D.</surname>
          </string-name>
          (
          <year>1996</year>
          ).
          <article-title>The lack of a priori distinctions between learning algorithms</article-title>
          .
          <source>Neural computation</source>
          ,
          <fpage>1341</fpage>
          -
          <lpage>1390</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [18]. Thornton,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Hutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Hoos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            , &amp;
            <surname>Leyton-Brown</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          (
          <year>2013</year>
          , April 21).
          <article-title>Auto-WEKA: Combined selection and hyperparameter optimisation of classification algorithms</article-title>
          . Available from arxiv.org: http://arxiv.org/pdf/1208.3719.pdf
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [19]. Carauna,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Nikilescu-Mizil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Crew</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , &amp;
            <surname>Ksikes</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2013</year>
          , May 20).[Internet]
          <article-title>Ensemble selection from libraries of models. Available from cs</article-title>
          .cornell.edu: http://www.cs.cornell.edu/~caruana/ctp/ct.papers/caruana.icm
          <year>l04</year>
          .
          <year>icdm06long</year>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [20]. Meyer,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          (
          <year>2013</year>
          , April 21).
          <article-title>New package for ensembling R models [Internet]</article-title>
          . Available from Modern Toolmaking: http://moderntoolmaking.blogspot.co.nz/
          <year>2013</year>
          /03/new-packagefor-ensembling
          <string-name>
            <surname>-</surname>
          </string-name>
          r-models.html
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [21]. Gallivan,
          <string-name>
            <surname>S</surname>
          </string-name>
          ; (
          <year>2003</year>
          )
          <article-title>How likely is it that a run of poor outcomes is unlikely?</article-title>
          <source>European Journal of Operational Research</source>
          ,
          <volume>150</volume>
          <fpage>46</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [22]. Hong,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>On computing the distribution function for the Poisson binomial distribution</article-title>
          .
          <source>Computational Statistics and Data Analysis</source>
          <volume>59</volume>
          <fpage>41</fpage>
          -
          <lpage>51</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [23].
          <string-name>
            <surname>Sherlaw-Johnson</surname>
            <given-names>C.</given-names>
          </string-name>
          <year>2005</year>
          <article-title>A method for detecting runs of good and bad clinical outcomes on Variable Life-Adjusted Display (VLAD) charts</article-title>
          .
          <source>Health Care Manag Sci. Feb</source>
          ;
          <volume>8</volume>
          (
          <issue>1</issue>
          ):
          <fpage>61</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>