<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>autoBagging: Learning to Rank Bagging Work ows with Metalearning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fabio Pinto</string-name>
          <email>fhpinto@inesctec.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V tor Cerqueira</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlos Soares</string-name>
          <email>csoares@fe.up.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jo~ao Mendes-Moreira</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>INESC TEC/Faculdade de Engenharia, Universidade do Porto Rua Dr. Roberto Frias</institution>
          ,
          <addr-line>s/n Porto, Portugal 4200-465</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Machine Learning (ML) has been successfully applied to a wide range of domains and applications. One of the techniques behind most of these successful applications is Ensemble Learning (EL), the eld of ML that gave birth to methods such as Random Forests or Boosting. The complexity of applying these techniques together with the market scarcity on ML experts, has created the need for systems that enable a fast and easy drop-in replacement for ML libraries. Automated machine learning (autoML) is the eld of ML that attempts to answers these needs. We propose autoBagging, an autoML system that automatically ranks 63 bagging work ows by exploiting past performance and metalearning. Results on 140 classi cation datasets from the OpenML platform show that autoBagging can yield better performance than the Average Rank method and achieve results that are not statistically different from an ideal model that systematically selects the best work ow for each dataset. For the purpose of reproducibility and generalizability, autoBagging is publicly available as an R package on CRAN.</p>
      </abstract>
      <kwd-group>
        <kwd>automated machine learning</kwd>
        <kwd>metalearning</kwd>
        <kwd>bagging</kwd>
        <kwd>classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Ensemble learning (EL) has proven itself as one of the most powerful techniques
in Machine Learning (ML), leading to state-of-the-art results across several
domains. Methods such as bagging, boosting or Random Forests are considered
some of the favourite algorithms among data science practitioners. However,
getting the most out of these techniques still requires signi cant expertise and
it is often a complex and time consuming task. Furthermore, since the number
of ML applications is growing exponentially, there is a need for tools that boost
the data scientist's productivity.</p>
      <p>
        The resulting research eld that aims to answers these needs is Automated
Machine Learning (autoML). In this paper, we address the problem of how to
automatically tune an EL algorithm, covering all components within it: generation
(how to generate the models and how many), pruning (which technique should
be used to prune the ensemble and how many models should be discarded) and
integration (which model(s) should be selected and combined for each
prediction). We focus speci cally in the bagging algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and four components of
the algorithm: 1) the number of models that should be generated 2) the pruning
method 3) how much models should be pruned and 4) which dynamic
integration method should be used. For the remaining of this paper, we call to a set of
these four elements a bagging work ow.
      </p>
      <p>Our proposal is autoBagging, a system that combines a learning to rank
approach together with metalearning to tackle the problem of automatically
generate bagging work ows. Ranking is a common task in information retrieval.
For instance, to answer the query of a user, a search engine ranks a plethora of
documents according to their relevance. In this case, the query is replaced by
new dataset and autoBagging acts as ranking engine. Figure 1 shows an overall
schema of the proposed system. We leverage the historical predictive performance
of each work ow in several datasets, where each dataset is characterised by
a set of metafeatures. This metadata is then used to generate a metamodel,
using a learning to rank approach. Given a new dataset, we are able to collect
metafeatures from it and feed them to the metamodel. Finally, the metamodel
outputs an ordered list of the work ows, taking into account the characteristics
of the new dataset.</p>
      <p>Algorithm
AlAgorithm
AlAgorithm</p>
      <p>a
dd dd
d</p>
      <p>Estimates of
Performance
Metafeatures
Extraction
New
Dataset</p>
      <p>Metadata
Learning to Rank</p>
      <p>Metamodel
Ranked
Algorithms</p>
      <p>
        We tested the approach in 140 classi cation datasets from the OpenML
platform for collaborative ML [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and 63 bagging work ows, that include two
pruning techniques and two dynamic selection techniques. Results show that
autoBagging has a better performance than two strong baselines, Bagging with 100
trees and the average rank. Furthermore, testing the top 5 work ows
recommended by autoBagging guarantees an outcome that is not statistically di erent
from the Oracle, an ideal method that for each dataset always selects the best
work ow. For the purpose of reproducibility and generalizability, autoBagging is
available as an R package.1
1 https://github.com/fhpinto/autoBagging
      </p>
      <p>
        autoBagging
We approach the problem of algorithm selection as a learning to rank problem [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Lets take D as the dataset set and A as the algorithm set. Y = f1; 2; :::; lg is
the label set, where each value represents a relevance score, which represents the
relative performance of a given algorithm. Therefore, l l 1 ::: 1, where
represents an order relationship.
      </p>
      <p>Furthermore, Dm = fd1; d2; :::; dmg is the set of datasets for training and di
is the i-th dataset, Ai = fai;1; ai;2; :::; ai;ni g is the set of algorithms associated
with dataset di and yi = fyi;1; yi;2; :::; yi;ni g is the set of labels associated with
dataset di, where ni represents the sizes of Ai and yi; ai;j represents the j-th
algorithm in Ai; and yi;j 2 Y represents the j-th label in yi, representing the
relevance score of ai;j with respect to di. Finally, the meta-dataset is denoted as
m
S = f(di; Ai); yigi=1.</p>
      <p>We use metalearning to generate the metafeature vectors xi;j = (di; ai;j ) for
each dataset-algorithm pair, where i = 1; 2; :::; m; j = 1; 2; :::; ni and represents
the metafeatures extraction functions. These metafeatures can describe di, ai;j or
even the relationship between both. Therefore, taking xi = fxi;1; xi;2; :::; xi;ni g
we can represent the meta-dataset as S0 = f(xi; yi)gim=1.</p>
      <p>Our goal is to train a meta ranking model f (d; a) = f (x) that is able to
assign a relevance score to a given new dataset-algorithm pair d and a, given x.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Metafeatures</title>
      <p>
        We approach the problem of generating metafeatures to characterize d and a
with the aid of a framework for systematic metafeatures generation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
Essentially, this framework regards a metafeature as a combination of three
components: meta-function, a set of input objects and a post-processing function.
The framework establishes how to systematically generate metafeatures from all
possible combinations of object and post-processing alternatives that are
compatible with a given meta-function. Thus, the development of metafeatures for a
MtL approach simply consists of selecting a set of meta-functions (e.g. entropy,
mutual information and correlation) and the framework systematically generates
the set of metafeatures that represent all the information that can be obtained
with those meta-functions from the data.
      </p>
      <p>
        For this task in particular, we selected a set of meta-functions that are able
to characterize the datasets as completely as possible (measuring information
regarding the target variable, the categorical and numerical features, etc) the
algorithms and the relationship between the datasets and the algorithms (who
can be seen as landmarkers [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). Therefore, the set of meta-functions used is:
skewness, Pearson's correlation, Maximal Information Coe cient (MIC [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]),
entropy, mutual Information, eta squared (from ANOVA test) and rank of each
algorithm [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Each meta-function is used to systematically measure information from all
possible combination of input objects available for this task. We de ned the
input objects available as: discrete descriptive data of the datasets, continuous
descriptive data of the datasets, discrete output data of the datasets, ve sets of
predictions (discrete predicted data) for each dataset (naive bayes, decision tree
with depth 1, 2 and 3, and majority class).</p>
      <p>For instance, if we take the example of using Entropy as meta-function, it
is possible to measure information in discrete descriptive data, discrete output
data and discrete predicted data (if the base-level problem is a classi cation
task). After computing the entropy of all these objects, it might be necessary to
aggregate the information in order to keep the tabular form of the data. Take for
the example the aggregation required for the entropy values computed for each
discrete attribute. Therefore, we choose a palette of aggregation functions to
capture several dimensions of these values and minimize the loss of information
by aggregation. In that sense, the post-processing functions chosen were: average,
maximum, minimum, standard deviation, variance and histogram binning.</p>
      <p>Given these meta-functions, the available input objects and post-processing
functions, we are able to generate a set of 131 metafeatures. To this set we add
eight metafeatures: the number of examples of the dataset, the number of
attributes and the number of classes of the target variable; and ve landmarkers
(the ones already described above) estimated using accuracy as error measure.
Furthermore, we add four metafeatures to describe the components of each
workow: the number of trees, the pruning method, the pruning cut point and the
dynamic selection method. In total, autoBagging uses a set of 143 metafeatures.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Metatarget</title>
      <p>In order to be able to learn a ranking meta-model f (d; a), we need to compute
a metatarget that represents a score z to each dataset-algorithm pair (d; a), so
that: F : (D; A) ! Z, where F is the ranking meta-models set and Z is the
metatarget set.</p>
      <p>
        To compute z, we use a cross validation error estimation methodology (4-fold
cross validation in the experiments reported in this paper, Section 3), in which
we estimate the performance of each bagging work ow for each dataset using
Cohen's kappa score [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. On top of the estimated kappa score, for each dataset,
we rank the bagging work ows. This ranking is the nal form of the metatarget
and it is then used for learning the meta-model.
3
      </p>
      <sec id="sec-3-1">
        <title>Experiments</title>
        <p>
          Our experimental setup comprises 140 classi cation datasets extracted from
the OpenML platform for collaborative machine learning [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. We limited the
datasets extracted to a maximum of 5000 instances, a minimum of 300 instances
and a maximum of 1000 attributes, in order to speed up the experiments and
exclude datasets that could be too small for some of bagging work ows that we
wanted to test.
        </p>
        <p>
          Regarding bagging work ows, we limited the hyperparameters of the bagging
work ows to four: number of models generated, pruning method, pruning cut
point and dynamic selection method. Speci cally, each hyperparameter could
take the following values:
{ Number of models: 50, 100 or 200. Decision trees was chosen as learning
algorithm.
{ Pruning method: Margin Distance Minimization(MDSQ) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], Boosting-Based
        </p>
        <p>
          Pruning (BB) [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] or none.
{ Pruning cut point: 25%, 50% or 75%.
{ Dynamic integration method: Overall Local Accuracy (OLA), a dynamic
selection method [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]; K-nearest-oracles-eliminate (KNORA-E) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], a dynamic
combination method; and none.
        </p>
        <p>
          The combination of all these hyperparameters described above generated 63
valid work ows. We tested these bagging work ows in the datasets extracted
from OpenML with 4-fold cross validation, using Cohen's kappa as evaluation
metric. We used the XGBoost learning to rank implementation for gradient
boosting of decision trees [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to learn the metamodel as described in Section 2.
        </p>
        <p>As baselines, at the base-level, we use 1) bagging with 100 decision trees 2)
the average rank method, which basically is a model that always predicts the
bagging work ow with the best average rank in the meta training set and the
3) oracle, an ideal model that always selects the best bagging work ow for each
dataset. As for the meta-level, we use as baseline the average rank method.</p>
        <p>
          As evaluation methodology, we use an approach similar to the leave-one-out
methodology. However, each test fold consists of all the algorithm-dataset pairs
associated with the test dataset. The remaining examples are used for training
purposes. The evaluation metric at the meta-level is the Mean Average Precision
at 10 (MAP@10) and at the base-level, as mentioned before, we use Cohen's
kappa. The methodology recommended by Demsar [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] was used for statistical
validation of the results.
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Figure 2 shows a loss curve, relating the average loss in terms of performance
with the number of work ows tested following the ranking suggested by each
method. The loss is calculated as the di erence between the performance of
the best algorithm ranked by the method in comparison with the ground truth
ranking. The loss for all datasets is then averaged for aggregation purposes. We
can see, as expected, that the average loss decreases for both methods as the
number of work ows tested increases.</p>
      <p>In terms of comparison between autoBagging and the Average Rank method,
it is possible to visualize that autoBagging shows a superior performance for all
the values of the x axis. Interestingly, this result is particularly noticeable in
the rst tests. For instance, if we test only the top 1 work ow recommended by
autoBagging, on average, the kappa loss is half of the one we should expect from
the suggestion made by the average rank method.</p>
      <p>0.08
0.06
%
s
s
o
L0.04
a
p
p
a
K
0.02
0.00
Method</p>
      <p>Autobagging
Average Rank
0
20</p>
      <p>Number of Tests 40
60</p>
      <p>We evaluated these results to assess their statistical signi cance using Demsar's
methodology. Figures 3 and 4 show the Critical Di erence (CD) diagrams for
both the meta and the base-level.</p>
      <p>AutoBagging</p>
      <p>CD
1
2
Average
Rank</p>
      <p>Oracle
AutoBagging@5
AutoBagging@3
1
CD
2
3
4
5
6
Bagging
AverageRank
AutoBagging@1</p>
      <p>At the meta-level, using MAP@10 as evaluation metric, autoBagging presents
a clearly superior performance in comparison with the Average Rank. The
difference is statistically signi cant, as one can see in the CD diagram. This result
is in accordance with performance that visualized in Figure 2 for both methods.</p>
      <p>At the base-level, we compared autoBagging with three baselines, as
mentioned before: bagging with 100 decision trees, the Average Rank method and
the oracle. We test three versions of autoBagging, taking the top 1, 3 and 5
bagging work ows ranked by the meta-model. For instance, in autoBagging@3,
we test the top 3 bagging work ows ranked by the meta-model and we choose
the best.</p>
      <p>Starting by the tail of the CD diagram, both the Average Rank method and
autoBagging@1 show a superior performance than Bagging with 100 decision
trees. Furthermore, autoBagging@1 also shows a superior performance than the
Average Rank method. This result con rms the indications that we visualized
in Figure 2.</p>
      <p>The CD diagram shows also autoBagging@3 and autoBagging@5 have a
similar performance. However, and we must highlight these results, autoBagging@5
shows a performance that is not statistically di erent from the oracle. This is
extremely promising since it shows that the performance of autoBagging excels
if the user is able to test the top 5 bagging work ows ranked by the system.</p>
      <sec id="sec-4-1">
        <title>Conclusion</title>
        <p>This paper presents autoBagging, an autoML system that makes use of a learning
to rank approach and metalearning to automatically suggest a bagging ensemble
speci cally designed for a given dataset. We tested the approach on 140
classication datasets and the results show that autoBagging is clearly better than
the baselines to which was compared. In fact, if the top ve work ows suggested
by autoBagging are tested, results show that the system achieves a performance
that is not statistically di erent from the oracle, a method that systematically
selects the best work ow for each dataset. For the purpose of reproducibility and
generalizability, autoBagging is publicly available as an R package.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Acknowledgements</title>
        <p>This work was partly funded by the ECSEL Joint Undertaking, the framework
programme for research and innovation horizon 2020 (20142020) under grant
agreement number 662189-MANTIS-2014-1.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Brazdil</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carrier</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilalta</surname>
          </string-name>
          , R.: Metalearning:
          <article-title>Applications to data mining</article-title>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Bagging predictors</article-title>
          .
          <source>Machine learning 24(2)</source>
          ,
          <volume>123</volume>
          {
          <fpage>140</fpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          . pp.
          <volume>785</volume>
          {
          <fpage>794</fpage>
          . KDD '16,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A coe cient of agreement for nominal scales</article-title>
          .
          <source>Educational and psychological measurement 20(1)</source>
          ,
          <volume>37</volume>
          {
          <fpage>46</fpage>
          (
          <year>1960</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Demsar</surname>
          </string-name>
          , J.:
          <article-title>Statistical comparisons of classi ers over multiple data sets</article-title>
          .
          <source>JMLR 7(Jan)</source>
          ,
          <volume>1</volume>
          {
          <fpage>30</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Ko</surname>
            ,
            <given-names>A.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabourin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Britto</surname>
            <given-names>Jr</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.S.:</surname>
          </string-name>
          <article-title>From dynamic classi er selection to dynamic ensemble selection</article-title>
          .
          <source>Pattern Recognition</source>
          <volume>41</volume>
          (
          <issue>5</issue>
          ),
          <volume>1718</volume>
          {
          <fpage>1731</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , T.Y.:
          <article-title>Learning to rank for information retrieval</article-title>
          .
          <source>Foundations and Trends® in Information Retrieval</source>
          <volume>3</volume>
          (
          <issue>3</issue>
          ),
          <volume>225</volume>
          {
          <fpage>331</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mart</surname>
          </string-name>
          nez-Mun~oz, G.,
          <string-name>
            <surname>Hernandez-Lobato</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suarez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An analysis of ensemble pruning techniques based on ordered aggregation</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>31</volume>
          (
          <issue>2</issue>
          ),
          <volume>245</volume>
          {
          <fpage>259</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bensusan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giraud-Carrier</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Tell me who can learn you and i can tell you who you are: Landmarking various learning algorithms</article-title>
          . In: ICML. pp.
          <volume>743</volume>
          {
          <issue>750</issue>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Pinto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendes-Moreira</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Towards automatic generation of metafeatures</article-title>
          .
          <source>In: PAKDD</source>
          . pp.
          <volume>215</volume>
          {
          <fpage>226</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Reshef</surname>
            ,
            <given-names>D.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reshef</surname>
            ,
            <given-names>Y.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Finucane</surname>
            ,
            <given-names>H.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grossman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McVean</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turnbaugh</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lander</surname>
            ,
            <given-names>E.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitzenmacher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabeti</surname>
            ,
            <given-names>P.C.</given-names>
          </string-name>
          :
          <article-title>Detecting novel associations in large data sets</article-title>
          .
          <source>Science</source>
          <volume>334</volume>
          (
          <issue>6062</issue>
          ),
          <volume>1518</volume>
          {
          <fpage>1524</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Vanschoren</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Van Rijn</surname>
            ,
            <given-names>J.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bischl</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torgo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Openml: networked science in machine learning</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          <volume>15</volume>
          (
          <issue>2</issue>
          ),
          <volume>49</volume>
          {
          <fpage>60</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Woods</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kegelmeyer</surname>
            <given-names>Jr</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>W.P.</given-names>
            ,
            <surname>Bowyer</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Combination of multiple classi ers using local accuracy estimates</article-title>
          .
          <source>Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>19</volume>
          (
          <issue>4</issue>
          ),
          <volume>405</volume>
          {
          <fpage>410</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>