<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multi-K Machine Learning Ensembles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matthew Whitehead</string-name>
          <email>matthew.whitehead@coloradocollege.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Larry S. Yaeger</string-name>
          <email>larryy@indiana.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Colorado College</institution>
          ,
          <addr-line>Mathematics and Computer Science, 14 E. Cache La Poudre St., Colorado Springs, CO 80903</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Indiana University, School of Informatics and Computing</institution>
          ,
          <addr-line>919 E. 10th St., Bloomington, IN 47408</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Ensemble machine learning models often surpass single models in classification accuracy at the expense of higher computational requirements during training and execution. In this paper we present a novel ensemble algorithm called Multi-K which uses unsupervised clustering as a form of dataset preprocessing to create component models that lead to effective and efficient ensembles. We also present a modification of Multi-K that we call Multi-KX that incorporates a metalearner to help with ensemble classifications. We compare our algorithms to several existing algorithms in terms of classification accuracy and computational speed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Groups of machine learning models, called ensembles, can
help increase classification accuracy over single models.
The use of multiple component models allows each to
specialize on a particular subset of the problem space,
essentially becoming an expert on part of the problem. The
component models are trained as separate, independent
classifiers using different subsets of the original training dataset or
using different learning algorithms or algorithm parameters.
The components can then be combined to form an
ensemble that has a higher overall classification accuracy than a
comparably trained single model. Ensembles often increase
classification accuracy, but do so at the cost of increasing
computational requirements during the learning and
classification stages. For many large-scale tasks, these costs can
be prohibitive. To build better ensembles we must increase
final classification accuracy or reduce the computational
requirements while maintaining the same accuracy level.</p>
      <p>In this paper, we discuss a novel ensemble algorithm
called Multi-K that achieves a high-level of classification
accuracy with a relatively small ensemble size and
corresponding computational requirements. The Multi-K
algorithm works by adding a training dataset preprocessing step
that lets training subset selection produce effective
ensembles. The preprocessing step involves repeatedly
clustering the training dataset using the K-Means algorithm at
different levels of granularity. The resulting clusters are then
used as training datasets for individual component
classifiers. The repeated clustering helps the component
classifiers obtain different levels of classification specialization,
ultimately leading to effective ensembles that rarely overfit.
We also discuss a variation on the Multi-K algorithm called
Multi-KX that includes a gating model in the final ensemble
to help merge the component classifications in an effective
way. This setup is similar to a mixture-of-experts system.
Finally, we show the classification accuracy and
computational efficiency of our algorithms on a variety of publicly
available datasets. We also compare our algorithms with
well-known existing ensemble algorithms to show that they
are competitive.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        One simple existing ensemble algorithm is called bootstrap
aggregating, or bagging
        <xref ref-type="bibr" rid="ref2">(Breiman 1996)</xref>
        . In bagging,
component models are given different training subsets by
randomly sampling the original, full training dataset. The
random selection is done with replacement, so some data points
can be repeated in a training subset. Random selection
creates a modest diversity among the component models.
      </p>
      <p>
        Bagging ensembles typically improve upon the
classification accuracies of single models and have been shown to
be quite accurate
        <xref ref-type="bibr" rid="ref2">(Breiman 1996)</xref>
        . Bagging ensembles
usually require a large number of component models to achieve
higher accuracies and these larger ensemble sizes lead to
high computational costs.
      </p>
      <p>
        The term boosting describes a whole family of
ensemble algorithms
        <xref ref-type="bibr" rid="ref12">(Schapire 2002)</xref>
        , perhaps the most famous
of which is called Adaboost
        <xref ref-type="bibr" rid="ref5">(Domingo &amp; Watanabe 2000)</xref>
        ,
        <xref ref-type="bibr" rid="ref4">(Demiriz &amp; Bennett 2001)</xref>
        . Boosting algorithms do away
with random training subset selection and instead have
component models focus on those training data points that
previously trained components had difficulty classifying. This
makes each successive component classifier able to improve
the final ensemble by helping to correct errors made by other
components.
      </p>
      <p>
        Boosting has been shown to create ensembles that have
very high classification accuracies for certain datasets
        <xref ref-type="bibr" rid="ref11">(Freund &amp; Schapire 1997)</xref>
        , but the algorithm can also lead to
model overfitting, especially for noisy datasets
        <xref ref-type="bibr" rid="ref9">(Jiang 2004)</xref>
        .
      </p>
      <p>
        Another form of random training subset selection is called
random subspace
        <xref ref-type="bibr" rid="ref7">(Ho 1998)</xref>
        . This method includes all
training data points in each training subset, but the included
data point features are selected randomly with replacement.
Adding in this kind of randomization allows components to
focus on certain features while ignoring others. Perhaps
predictably, we found that random subspace performed better
on datasets with a large number of features than on those
with few features.
      </p>
      <p>
        An ensemble algorithm called mixture-of-experts uses
a gating model to combine component classifications to
produce the ensemble’s final result (Jacobs et al.
        <xref ref-type="bibr" rid="ref8">1991),
(Nowlan &amp; Hinton 1991</xref>
        ). The gating model is an extra
machine learning model that is trained using the outputs of the
ensemble’s component models. The gating model can help
produce accurate classifications, but overfitting can also be
a problem, especially with smaller datasets.
      </p>
      <p>
        The work most similar to ours is the layered, cluster-based
approach of (Rahman &amp; Verma 2011). This work appears
to have taken place concurrently with our earlier work in
        <xref ref-type="bibr" rid="ref13">(Whitehead 2010)</xref>
        . Both approaches use repeated
clusterings to build component classifiers, but there are two key
differences between the methods. First, Rahman and Verma
use multiple clusterings with varying starting seed values at
each level of the ensemble to create a greater level of
training data overlap and classifier diversity. Our work focuses
more on reducing ensemble size to improve computational
efficiency, so we propose a single clustering per level.
Second, Rahman and Verma combine component classifications
using majority vote, but our Multi-KX algorithm extends
this idea by placing a gating model outside of the
ensemble’s components. This gating model is able to learn how to
weight the various components based on past performance,
much like a mixture-of-experts ensemble.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Multi-K Algorithm</title>
      <p>We propose a new ensemble algorithm that we call
MultiK, here formulated for binary classification tasks, but
straightforwardly extensible to multidimensional
classification tasks. Multi-K attempts to create small ensembles with
low computational requirements that have a high
classification accuracy.</p>
      <p>To get accurate ensembles with fewer components, we
employ a training dataset preprocessing step during
ensemble creation. For preprocessing, we repeatedly cluster the
training dataset using the K-Means algorithm with different
values of K, the number of clusters being formed. We have
found that this technique produces training subsets that are
helpful in building components that have a good mix of
generalization and specialization abilities.</p>
      <p>During the preprocessing step the value for the number
of clusters being formed, k, varies from Kstart to Kend.
Kstart and Kend were fixed to two and eight for all our
experiments as these values provided good results during
pretesting. Each new value of k then yields a new clustering
of the original training dataset. The reason that k is varied is
to produce components with different levels of classification
specialization ability.</p>
      <p>With small values of k, the training data points form larger
clusters. The subsequent components trained on those
subsets typically have the ability to make general classifications
well: they are less susceptible to overfitting, but are not
experts on any particular region of the problem space. Figure
1 shows the limiting case of k = 1, for which a single
classifier is trained on all of the training data. Figure 2 shows that
as the value of k increases, classifiers are trained on smaller
subsets of the original data.</p>
      <p>Larger values of k allow the formation of smaller, more
specialized clusters. The components trained on these
training subsets become highly specialized experts at classifying
data points that lie nearby. These components can overfit
when there are very few data points nearby, so it is
important to choose a value for Kend that is not too large. This
type of repeated clustering using varying values of k forms
a pseudo-hierarchy of the training dataset.</p>
      <p>Figure 3 shows that as k is further increased, the
clustered training datasets decrease in size allowing classifiers
to become even more highly specialized. In this particular
example, we see that classifier 6 will be trained on the same
subset of training data as classifier 3 was above. In this way,
tightly grouped training data points will be focused on since
they may be more difficult to discriminate between. The
clustering algorithm partitions the data in such a way as to
foster effective specialization of the higher-k classifiers, thus
maximizing those classifiers’ ability to discriminate.</p>
      <p>Following the clustering preprocessing step, each
component model is trained using its assigned training data subset.
When all the components are trained, then the ensemble is
ready to make new classifications. For each new data point
to classify, some components in the ensemble make
classification contributions, but others do not. For a given new data
point to classify, p, for each level of clustering, only those
components with training data subset centroids nearest to p
influence the final ensemble classification. Included
component classifications are inversely weighted according to their
centroid’s distance from p.</p>
      <p>The final classification is the weighted average of the n
ensemble components’ classifications in the current
ensemble formation:
n
X wi · Ci(p)
i=0
n
X wi
i=0
where Ci(p) is classifier i’s classification of target data
point p and each wi is an inverse distance of the form:
1
wi = dist(p,centroidi)</p>
      <p>Pseudocode listing 1 shows the algorithm for clustering
and training the component classifiers in Multi-K. Once the
clustering and component classifier training is complete,
then the ensemble is ready to classify new data points.
Pseudocode listing 2 shows the algorithm for choosing the
appropriate component classifiers given each new target data point
to classify.</p>
      <p>Given</p>
      <sec id="sec-3-1">
        <title>D : training dataset</title>
        <p>Kstart: The number of clusters in the first level of clustering.
Kend: The number of clusters in the last level of clustering.
Ensemble Training Pseudocode
for k from Kstart to Kend:
cluster D into k clusters, Dki, i ∈ [1, k]
for i from 1 to k:
train classifier Cki on data Dki</p>
      </sec>
      <sec id="sec-3-2">
        <title>Pseudocode 1: Multi-K ensemble training.</title>
        <p>Ensemble Formation Pseudocode</p>
        <sec id="sec-3-2-1">
          <title>Given, p, a data point to classify</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>For each clustering (k):</title>
          <p>Find the cluster, Dki, with centroid
&lt; Dki &gt; nearest to p
Add Cki, trained on Dki, to ensemble
Compute weight of Cki with distance from
&lt; Dki &gt; to p</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Pseudocode 2: Multi-K ensemble formation.</title>
        <p>Finally, once the appropriate component classifiers have
been selected, then the final ensemble classification can be
calculated. Pseudocode listing 3 shows the algorithm for
calculating the final classification based on a weighted sum
of the outputs of the selected components.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Multi-KX Algorithm</title>
      <p>The Multi-K algorithm used training dataset preprocessing
to form effective component models. Final classifications
were then performed by the entire ensemble by combining
component classifications together based on the distance
between xpredict and the centroids of each of the component
Ensemble Classification Pseudocode
sum = 0
sum weights = 0
for each classifier, C, in ensemble:
1
weightC = dist(C,p)
sum += weightC * C′s classification of p
sum weights += weightC
final_classification = sum / sum_weights</p>
      <sec id="sec-4-1">
        <title>Pseudocode 3: Multi-K ensemble classification.</title>
        <p>training datasets. This technique works well, but we also
thought that there may be non-linear interactions between
component classifications and a higher accuracy could be
gained by using a more complex method of combining
components.</p>
        <p>With this in mind, we propose a variation to Multi-K,
called Multi-KX. Multi-KX is identical to Multi-K except
in the way that component classifications are combined.
Instead of using a simple distance-scaled weight for each
component, Multi-KX uses a slightly more complex method that
attempts to combine component outputs in intelligent ways.
This intelligent combination method is achieved by the use
of a gating metanetwork. This type of metanetwork is used
in standard mixture-of-experts ensembles. This
metanetwork’s job is to learn to take component classifications and
produce the best possible final ensemble classification.
Figure 6 shows the basic setup of the ensemble.</p>
        <p>The metanetwork can then learn the best way to combine
the ensemble’s components. This can be done by weighting
certain components higher for certain types of new problems
and ignoring or reducing weights for other components that
are unhelpful for the current problem.</p>
        <p>Building a Metapattern For the metanetwork to
effectively combine component classifications, it must be trained
using a labeled set of training data points. This labeled
training set is similar to any other supervised learning problem:
it maps complex inputs to a limited set of outputs. In this
case, the metanetwork’s training input patterns are made up
of two different kinds of values. First, each classification
value is included from all the ensemble’s components. Then
the distance between xpredict and each cluster’s centroid is
also included. Figure 7 shows the general layout for a
metapattern.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Results</title>
      <p>To test the performance of our algorithms, we performed
several experiments. First, we tested the classification
accuracy of our algorithms against existing algorithms. Second,
we measured the diversity of our ensembles compared to
existing algorithms. Finally, we performed an accuracy vs.
computational time test to see how each algorithm performs
given a certain amount of computational time for ensemble
setup and learning.</p>
    </sec>
    <sec id="sec-6">
      <title>Accuracy Tests</title>
      <p>
        Ensembles need to be accurate in order to be useful. We
performed a number of tests to measure the classification
accuracy of our proposed algorithms and we compared these
results with other commonly-used ensemble techniques. We
tested classification accuracy on a series of datasets from the
UCI Machine Learning Repository
        <xref ref-type="bibr" rid="ref1">(Asuncion &amp; Newman
2007)</xref>
        along with a sentiment mining dataset from
        <xref ref-type="bibr" rid="ref13">(Whitehead &amp; Yaeger 2009)</xref>
        . We performed a K-fold cross
validation (with K=25) test using each algorithm on each dataset
and we repeated each test ten times to ensure that the results
were statistically stable. Each reported accuracy value is the
mean of the resulting 250 test runs.
      </p>
      <p>
        All accuracy tests were performed using support vector
machines
        <xref ref-type="bibr" rid="ref3">(Chang &amp; Lin 2001)</xref>
        with linear kernels as the
component classifiers, except we also compare our
accuracies with boosting ensembles of decision stumps since the
boosting algorithm is known to suffer less from overfitting
with these component classifiers. For these tests, ensembles
created with commonly used algorithms each had 50
component classifiers, as in
        <xref ref-type="bibr" rid="ref2">(Breiman 1996)</xref>
        .
      </p>
      <p>Table 1 shows the classification accuracies for each tested
algorithm and dataset. For each tested dataset, the most
accurate result is shown in bold face. These results show that
the proposed algorithms are competitive with existing
ensemble algorithms and are able to outperform all of those
algorithms for some datasets. The telescope dataset in
particular yielded encouraging results with more than a 3%
increase in classification accuracy (a 16% reduction in error)
obtained by the Multi-KX algorithm. Performance was also
good on the ionosphere dataset with an almost 2% higher
accuracy (a 17% reduction in error) than other ensemble
algorithms.</p>
      <p>The dataset that Multi-K performed most poorly on was
the restaurant sentiment mining dataset, where it was more
than 2% behind the subspace ensemble. Since that dataset
uses an N-gram data representation model, the data is
considerably more sparse than the other tested datasets. We
hypothesize that the sparsity made the clustering and the
resulting component classifiers less effective. None of the
ensemble algorithms were able to outperform a single SVM
on the income dataset. This again suggests that the nature of
the dataset will occasionally determine which algorithms do
well and which do poorly.</p>
    </sec>
    <sec id="sec-7">
      <title>Diversity Measurements</title>
      <p>
        To form an effective ensemble, a certain amount of
diversity among component classifiers is required. We measured
the diversity of the ensembles formed by Multi-K using
four different pairwise diversity metrics from
        <xref ref-type="bibr" rid="ref10">(Kuncheva &amp;
Whitaker 2003)</xref>
        :
• Q statistic - the odds ratio of correct classifications
between the two classifiers scaled to the range -1 to 1.
• ρ - the correlation coefficient between two binary
classifiers.
• Disagreement measure - proportion of the cases where the
two classifiers disagree.
• Double-fault measure - proportion of the cases
misclassified by both classifiers.
      </p>
      <p>Figure 8 shows that the diversity of Multi-K ensembles
generally falls in between algorithms that rely on random
subsampling (bagging and subspace) and the one tested
algorithm that particularly emphasizes diversity by focusing
on previously misclassified training points (boosting). For
example, values for the Q statistic and ρ were higher for the
random methods and lower for boosting. The disagreement
measure again shows Multi-K in the middle of the range.</p>
      <p>Double fault values were nearly identical across all
algorithms, suggesting that double fault rate is a poor metric for
measuring the kind of diversity that is important to create
ensembles with a good mix of generalization and
specialization.</p>
    </sec>
    <sec id="sec-8">
      <title>Combining Accuracy and Computational</title>
    </sec>
    <sec id="sec-9">
      <title>Efficiency</title>
      <p>Since our main goal was to provide an algorithm that yielded
high classification accuracies with the minimal amount of
computational overhead, we performed a final combined
accuracy and complexity test to show the relationship between
the two for various ensemble algorithms. To do this, we ran
existing ensemble algorithms with a wide variety of
parameters that affect accuracy and training time. We then plotted a
number of these training time/classification accuracy points,
hoping to provide a simple, but informative way to compare
the two across ensemble algorithms. We also ran each of the
Multi-* algorithms and plotted each result as a single point
on the graph because they have no free parameters to change.
Each plot line is a best-logarithmic-fit for each existing
algorithm to help see general trends. Figure 9 shows the results
and has been normalized against the classification accuracy
and computational time of a single SVM. Averaging across
all tested datasets, Multi-K provided higher accuracy than
other algorithms using anything less than about three times
the compute time, and Multi-KX provided the highest
accuracy of all, out to at least twice its computational costs.</p>
    </sec>
    <sec id="sec-10">
      <title>Conclusions and Future Directions</title>
      <p>We found that ensembles created using the Multi-*
algorithms showed a good amount of diversity and had strong
classification performance. We attribute this performance
to a good mix of components with varying levels of
generalization and specialization ability. Some components are
effective over a large number of data points and thus exhibit
the ability to generalize well. Other components are highly
specialized at making classifications in a relatively small
region of the problem space. The mix of both these kinds of
components seems to work well when forming ensembles.</p>
      <p>In the future, we hope to extend our method beyond
binary to multi-class classification. In addition, we
speculate that including additional characterizations of datasets
and models as inputs to the gating network may further
improve the accuracy of Multi-KX. We also are continuing
work investigating alternative ways of preprocessing
training datasets.</p>
      <p>LIBmachines.
tive mixtures of competing experts. Advances in Neural
Information Processing Systems.
purpose cross-domain sentiment mining model. In
Proceedings of the 2009 CSIE World Congress on Computer
Science and Information Engineering. IEEE Computer
Society.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Asuncion</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>UCI machine learning repository</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>1996</year>
          .
          <article-title>Bagging predictors</article-title>
          .
          <source>Machine Learning</source>
          <volume>24</volume>
          (
          <issue>2</issue>
          ):
          <fpage>123</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>C. C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Demiriz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bennett</surname>
            ,
            <given-names>K. P.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Linear programming boosting via column generation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Domingo</surname>
          </string-name>
          , and
          <string-name>
            <surname>Watanabe</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Madaboost: A modification of adaboost</article-title>
          .
          <source>In COLT: Proceedings of the Workshop on Computational Learning Theory</source>
          , Morgan Kaufmann Publishers.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>JCSS: Journal of Computer and System Sciences</source>
          <volume>55</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Ho</surname>
            ,
            <given-names>T. K.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>The random subspace method for constructing decision forests</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on Volume
          <volume>20</volume>
          (
          <issue>Issue 8</issue>
          ):
          <fpage>832</fpage>
          -
          <lpage>844</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          1991.
          <article-title>Adaptive mixtures of local experts</article-title>
          .
          <source>Neural Computation</source>
          <volume>3</volume>
          :
          <fpage>79</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Boosting with noisy data: Some views from statistical theory</article-title>
          .
          <source>Neural Computation</source>
          <volume>16</volume>
          (
          <issue>4</issue>
          ):
          <fpage>789</fpage>
          -
          <lpage>810</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Kuncheva</surname>
            ,
            <given-names>L. I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Whitaker</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <year>2003</year>
          .
          <article-title>Measures of diversity in classifier ensembles</article-title>
          .
          <source>Machine Learning</source>
          <volume>51</volume>
          :
          <fpage>181</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Freund</surname>
          </string-name>
          , and
          <string-name>
            <surname>Schapire</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>A decision-theoretic generalNowlan, S</article-title>
          . J., and
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G. E.</given-names>
          </string-name>
          <year>1991</year>
          .
          <article-title>Evaluation of adapRahman, A</article-title>
          ., and
          <string-name>
            <surname>Verma</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Novel layered clustering-based approach for generating ensemble of classifiers</article-title>
          .
          <source>In IEEE Transactions on Neural Networks</source>
          ,
          <fpage>781</fpage>
          -
          <lpage>792</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Schapire</surname>
            ,
            <given-names>R. E.</given-names>
          </string-name>
          <year>2002</year>
          .
          <article-title>The boosting approach to machine learning: An overview</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Whitehead</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Yaeger</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2009</year>
          . Building a general Whitehead,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Creating fast and efficient machine learning ensembles through training dataset preprocessing</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>