<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Diversity as The Basis for Effective Clustering-Based Classification</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dept. of Computer Science</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khmelnytskyi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ukraine alexander.barmak@gmail.com</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer Science and Information Technologies Khmelnytskyi National University</institution>
          ,
          <addr-line>Khmelnytskyi</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Theoretical Cybernetics Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>- Diversity is the basis of approaches using multiple classification systems. Solution sets are formed in ensembles of models. Many models give poorly distinguishable results. The use of strongly correlated model results in ensembles significantly reduces their effectiveness. Therefore, here the main influence is exerted by the variety of decisions made by the models. The value of each model decreases with an increase in the group of models. Accordingly, it is necessary to reduce the contribution of an individual model to the solution of the model. An approach based on clustering is proposed, according to which the influence of an individual model is inversely proportional to the volumes of aggregated groups. With this approach, the influence of an individual solution of the model, which differs from others, is significantly increased. Aggregation of groups is made in direct proportion to the correlation of decisions. Moreover, the aggregation of groups of models is performed according to the hierarchical structure of the ensemble. The solutions of strongly correlated groups of models are replaced by a single cluster solution. This solution at the next level can be grouped with other closest groups of models. Due to this architecture, the level of influence of a single solution of the model is increased. The main advantage of the proposed approach is the determination of the structure of the ensemble depending on the correlation of model decisions. Clusterization of decisions for features of similarity enhances the role of diversity and allows leveling out the error of an individual decision at a local level and to provide acceptable global indicators of cluster efficiency. Advantage of the proposed approach is the possibility of building an ensemble based on the properties of the correlation parameters of the models.</p>
      </abstract>
      <kwd-group>
        <kwd>diversity</kwd>
        <kwd>correlation classification</kwd>
        <kwd>hierarchical clustering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The basis for the use of group classification methods such as ensembles is the
diversity of solutions of the models. For successful use, models must provide diverse and at
the same time accurate solutions. Each model complements its solution with other
models. This underlies the application of the ensemble. However, obtaining a variety
of solutions is difficult, since the models are often trained on the same data, and they
are based on similar mathematical approaches. The consequence of this is similar
results, which have a strong correlation. Variety is also necessary for using the
accuracy of solutions since combining less accurate models often gives better results.
Supplementing informativeness with models is effective with a low correlation of
decisions. Correlation is one of the most important indicators of the need to use the model
in determining group decisions. And it can also serve as a criterion for determining
the need to use a model in an ensemble. Therefore, the influence of each model
should be determined depending on the correlation of decisions on the overall result
of the ensemble. Consideration of the peculiarities of the model should be displayed
in the architecture of the ensemble. At the same time, the model should improve the
outcome of the overall solution.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>Mutual addition of a solution in order to obtain an objective assessment in the form of
a general solution is a fundamental principle on which the methods of group
determination of solutions are based, as an example of ensembles of models [1]. A variety of
methods are used to combine group models with data manipulation, with enhanced
features. There are various modifications of bugging [2], boosting [3], stacking [4] as
the area of the most well-known ensembles.</p>
      <p>Despite the fact that the concept of diversity is intuitive, attempts to develop a
system for measuring diversity are quite extensive. An example is [5, 6, 7]. The
definition of a measure of diversity would allow the development of an approach based on
that measure. However, the variety of factors affecting the effectiveness of the
application of known solutions is very wide, requires extensive experience in use and does
not always give the necessary result.</p>
      <p>Another approach is to reduce the number of models that make up an ensemble.
Pruning can reduce the redundancy of models in a group since models can also worse
the outcome of ensemble predictions. At the same time, the computational complexity
of the combination of models still remains significantly at the modern level of
technology. Known a pruning method based on forward selection [8], ensemble pruning
based on objection maximization [9] and others. This allows optimizing the
calculations to achieve the expected result.</p>
      <p>The most interesting in our opinion are the clustering-based [10, 11] methods. The
rationale for this is the fact that the correlation of solutions contributes to the grouping
of objects into certain agglomerations. Moreover, the difference in solutions within
the agglomeration is relatively small. This may indicate that the correlation between
the solutions of these models is high. And accordingly, they can be reduced by the
required size. The correlation of models is an important factor, as it can be assumed
that correlation is the inverse of diversity. By decreasing correlation, it is possible to
increase the variety of solutions of ensemble models.</p>
      <p>The selection of ensemble models is also important. The difference in models can
often be determined based on the results of their predictions. The informative features
of the data on which the model is based also play an important role. To expand the
diversity of models, approaches using human intellectual abilities can be used. This
allows building machine models based on the mental models of humans [12, 13, 14].</p>
      <p>Methods of increasing the diversity of models and their effective use in ensembles
is an important area of research. In this paper, we consider the basics of building
ensembles based on the correlation of models and determine the effectiveness of the
result based on the estimation of the prediction error.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The individual influence of classifiers on the ensemble</title>
      <p>The solution of the ensemble is a set of individual predictions of classifiers. In
accordance with this, it is necessary to determine how the prediction of an individual
classifier can affect the general decision.
1. The classifier may have the wrong solution. However, with the vast majority of
correct decisions, the effect of an incorrect decision may generally be of little
importance.
2. An incorrect solution to each of the classifiers introduces noise into the overall
solution, thereby adversely affecting the quality of general solutions as a noise
source. A noise source can be understood as errors of one of the classifiers in a
significant percentage bias. Also, the noise source can be a set of errors of all
classifiers that appear relativity to each prediction element.
3. The correct decision of an individual classifier can be refuted by a combination of
incorrect decisions of a group of classifiers. And the presence in many solutions of
the right one will not affect the general wrong decision.
4. A strong correlation of incorrect decisions of classifiers may have a dominant
effect on the result of the work of the ensemble, despite the presence of correct
predictions of other classifiers.
5. In the general case, a strong correlation of solutions for all data elements leads to
worse results, since there is a greater probability that classifiers receive less
information from the data and have less overall value within the training sample in
comparison with the total data set. For arithmetic means, this is based on the
dispersion formula for the sums of correlated random variables. In this case, the
question arises of the influence of individual correct predictions of classifiers on the
general incorrect decision of the ensemble.</p>
      <p>Of primary interest is the definition of the influence of an individual classifier
solution on the overall incorrect decision. The obvious influence of the majority on the
decisions of the group, in this case, is erroneous. The interconnection of decisions
indicates the incorrectness of the general approach for making decisions regarding a
particular data element. This means that this element has a negligible content of
general informativeness, on which the decisions of the majority were based. At the same
time, if another decision was made on it by any classifier, there is another
informativeness that was not determined by the majority. This is significant in the event of a
majority group error.</p>
      <p>The positive correlation of group decisions has an undesirable effect. The presence
of a negative correlation may indicate the presence of an alternative opinion. This
aspect indicates the variability of predictions. A prerequisite is the presence of a
positive relationship between the prediction of the ensemble and the expected result.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The hierarchical structure of ensemble as the globalization of local solutions</title>
      <p>The presence of a strong positive correlation between the individual decisions of the
models within the ensemble facilitates their aggregation into groups. The decision of
each group model is strongly correlated with the decisions of other members of this
group. With a large measure of generalization, it is possible to formulate a general
decision of the group, which to one degree or another will represent the solution of
each model. Since differences in model decisions are insignificant within the group,
the generalizing ability of aggregation will be significant within the ensemble. This
allowing to divide the aggregate of models in the ensemble into aggregations on the
basis of strongly positively correlated groups and representing clusters. An individual
decision of a model within a cluster is of insignificant value, and it can be replaced by
a generalized solution - a cluster decision. A cluster decision represents a solution to
the models that form its solutions, and each individual model delegates its opinion to
the cluster. Further subsequent aggregation of cluster decisions forms the ensemble
solution. Such a process of delegating a decision to a higher-order level allows
creating a hierarchical structure for the formation of the ensemble decision. Using this
approach, the set of highly correlated solutions is replaced by a single cluster solution.
The influence of an individual element on the ensemble solution decreases under
conditions of strong correlation with other elements, and as a result of this, is determined
by its location in any cluster. The larger the group size, the less influence the model
has on the ensemble decision.</p>
      <p>The formation of groups allowing to gradually reduce the variance in the ensemble.
The variance of the group is replaced by a bias of a higher hierarchical order. A
hierarchical structure of the ensemble forms the conditions for the separation of decisions
by levels of locality, and the delegation process gradually globalizes local solutions.
An important consequence of this process is that global solutions may be wrong at the
local level, and local solutions may differ from the global one. The globalization of
local solutions through delegation through the hierarchical structure of the ensemble
improves the bias-variance tradeoff. The cluster decision is formed on the basis of an
unbiased estimate. In this case, the dispersion of the cluster at the highest level of the
hierarchy is not taken into account, and in fact, the distribution within the group is
converted into a solution of the cluster, which has a certain bias at the next level. This
allowed creating conditions for heterogeneous accounting for model predictions in the
ensemble solution. The participation of the model in the global solution is made
dependent on the correlation strength of the solution with relation to other models. This
translates into a general rule for the structure of the hierarchy: the more general
information contained in the model, the less its participation in global prediction. It is
also, the correctness of a local solution may be weakly correlated with a global
solution.</p>
      <p>Consider a simple example in a one-dimensional space, which is shown in Figure
1. This example demonstrates a general unbiased hierarchical estimate. Group I
consists of six strongly positively correlated models. Group II consists of one model.
However, the participation of this one model is high in relation to the general decision
of the groups. The figure shows the hierarchy of the second level and the unbiased
estimation of the highest level.
Aggregation of model predictions and location relative to the expected target
In the general case, group II can increase the bias and therefore worsen the result. As
an example, an asymmetric arrangement of groups I and II relative to the expected
result (target). This is a consequence of increasing the value of a weakly correlated
model. An important condition for this approach is the presence of a hierarchy of a
higher level. In our case, this is a level three hierarchy. At the third level of the
hierarchy (and this is the level of the ensemble), the model should have acceptable
indicators of unbiased estimation. Only in this case, an increase in local bias can indicate
that the model takes into account some information content that other models could
not determine. Moreover, information content is exclusively local in nature and
weakly correlates with the general. Using models with a biased global estimate introduces
uncertainty in which we cannot determine whether the prediction of the model at the
local level is an outlier or the model was able to determine the hidden local
information content of the data. Thus, under the condition of a global unbiased assessment
of the model, the hierarchical structure of the ensemble makes it possible to
strengthen the latent information content of the data. And in the general case, it tells us that it
is necessary to use some form of cascading classification as a consequence of the
appearance of uncertainty and the process of strengthening local information content
in relation to the global one.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Correlation of models and bias of ensemble solutions</title>
      <p>Let us examine how the presence of model correlation affects the error of ensemble
decisions. The ensemble makes decisions using the hierarchical structure of
delegation. The model is accepted into the ensemble if it is overall unbiased. We will
use the normal distribution of model predictions. For the simplest hierarchy structure
with symmetrical distribution, the number of groups is two. This is also consistent
with a minimum number of alternative solutions.</p>
      <p>We take the relation between the predictions of two cluster groups
CL1 = C1, C2 ,..., Cn1 and CL2 = C1, C2 ,..., Cn2 , n1, n2  N . Let the prediction
Err (C L1 ) and Err (C L2 ) errors of two clusters (groups) of models. We accept the
presence of cluster decision errors in general. Within each cluster, both variability of
predictions and the presence of a strong positive correlation are possible. The studied
variables Err (C L1 ) and Err (C L2 ) are both distributed normally and have large
unbiased dispersion  CL of cluster predictions. Group predictions are opposite
signP(CLi (x))  signP(CLk (x)) , x  X , i  k 1,2
Groups are correlated with a certain Pearson coefficient  CL . The combination of the
two groups should lead to the expected result with a minimum ensemble prediction
error min Err (E ) . Based on the opposite of the predictions of the groups of models,
the ensemble prediction error</p>
      <p>Err (E ) =Err (CLi ) + (1 − )Err (CLk ) , i  k 1,2,   0,1
Accordingly, the variance of the error</p>
      <sec id="sec-5-1">
        <title>Since covariance</title>
        <p>Var(Err(E)) = 2 C2Li + (1 − )2 C2Lk +
2 (1 − )Cov(Err(CLi ), Err(CLk ))</p>
        <p>, i  k 1,2</p>
        <p>Cov(Err (CL )) =  CL CLi CLk , i  k 1,2
Dispersion of ensemble error</p>
        <p>
          Var(Err(E )) =  2 C2Li + (1 − )2 C2Lk + 2 (1 − ) CL CLi CLk , i  k 1,2
The parameter   0,1 is a nondeterministic quantity and allows one to study the
variance of the ensemble error at the extremum
Var(Err(E ))

= 0
Solution (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ) with relation to the Pearson coefficient  CL has the form
 CL =
        </p>
        <p>
          (2 − 1) CLi CLk
The minimum coefficient value  CL , based on equation (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ), can be obtained under
the condition
Relating the parameter 
 C2Li −  C2Lk + C2Lk = 0 , i  k 1,2
 =
        </p>
        <p>
           C2Lk
 C2Li +  C2Lk
, i  k 1,2
Provided  C2Li =  C2Lk , i  k 1,2 we get the parameter value
We solve (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ) with relation to the parameter 
We set the condition  C2Li =  C2Lk , i  k 1,2
        </p>
        <p>CLk  CL CLi −</p>
        <p>CLk
 C2Li +  C2Lk − 2 CL CLi CLk</p>
        <p>
          , i  k 1,2
 = −
 2 ( CL −1)
2 2 (1−  CL )
, i  k 1,2
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
(
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
(
          <xref ref-type="bibr" rid="ref11">11</xref>
          )
(
          <xref ref-type="bibr" rid="ref12">12</xref>
          )
(
          <xref ref-type="bibr" rid="ref13">13</xref>
          )
(
          <xref ref-type="bibr" rid="ref14">14</xref>
          )
(
          <xref ref-type="bibr" rid="ref15">15</xref>
          )
This corresponds to the result (
          <xref ref-type="bibr" rid="ref10">10</xref>
          ). Taking into account (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) and (
          <xref ref-type="bibr" rid="ref11">11</xref>
          ), the main
condition for minimizing the ensemble prediction error is the equality of the variance of the
clusters of the ensemble of models and the opposite of their predictions.
  C2Li = C2Lk , i  k 1,2;

signP(CLi (x))  signP(CLk (x)), x  X , i  k 1,2.
        </p>
        <p>
          System (
          <xref ref-type="bibr" rid="ref13">13</xref>
          ) corresponds to the unbiased variance of the ensemble of models at the
level of global estimation of the data set. The dispersion symmetry ensures
complementarity of predictions of model clusters when used in ensembles with unbiased
estimates and of individual models in the general case.
        </p>
        <p>
          We use equation (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) under the condition of the maximum value of the coefficient
 CL = 1
Relating the parameter 
        </p>
        <p>
lim  −
 i →0</p>
        <p> CLk  = 1 , i  k  1,2
 CLi −  CLk 
Provided  CLi → 0 value  → 1 .</p>
        <p>Based on (16), bias from the median value of the parameter  is accompanied by
an increase in intergroup correlation. In the general case, the bias of the dispersed
estimate is manifested in the correlation dependence. Thus, it can be assumed that the
presence of a correlation between the clusters is the cause of the bias.</p>
        <p>The dispersion of solutions of a cluster (group) has a dependence on the internal
dispersion of solutions of the models forming this cluster = f ( ). Solutions within
the cluster are also correlated.</p>
        <p>Consider a situation in which  i → 0 value  → 1 . The general correlation of the
ensemble is dependent on the internal correlation of the groups. If the correlation of
one of the clusters tends to zero, then it can be assumed that the intragroup correlation
value of the other group is largely represented in the general correlation at  → 1 .
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Experimental studies</title>
      <p>To verify the effectiveness of using the hierarchical structure, it is necessary to
conduct experimental studies. It is necessary to form a data set for classification, and sets
of correlated predictions of solutions. The most optimal from the point of view of
understanding the effectiveness will be synthetic tests that provide the necessary
parameters for research management.</p>
      <p>Based on the generated normal distribution data using a linear relationship,
variations of the model predictions are generated. This set of solutions is used to obtain
correlated solutions.</p>
      <p>To generate a correlation of random predictions, we use the Cholesky
decomposition. If Corr is a correlation matrix, the Cholesky expansion has the form</p>
      <sec id="sec-6-1">
        <title>Accordingly, random variables can be generated.</title>
        <p>Here X is the base distribution, Y is the correlated distribution.</p>
        <p>When generating a correlation of two random variables</p>
        <p>LLT = Corr</p>
        <p>LX = Y
 1
L = 
</p>
        <p>0 
1 −  2 

Here  is the Pearson correlation coefficient.
(16)
(17)
(18)
(19)
Thus, using the set of Pearson correlation coefficient 1,  2 ,...,  n , i = 1,..., n , sets
of correlated model predictions S are formed.
Clustering is a big area of data analysis methods that are used to obtain information
about the data structure. This allows determining the aggregation of data and the
identification of these groups. The data in each cluster are most similar in terms of the
measure of similarity of the correlation distance, as one of the criteria. Cluster
analysis will be made based on the decisions made by the models regarding the data
element. In our case, clustering is one of the stages of classifications. Of the various
clustering methods, we will use the simplest one in order to simplify the
understanding of the proposed approach. The method should use the definitions of grouping data
into a predetermined number of clusters. In experimental studies, we will use two
clusters. For this purpose, the k-means algorithm is chosen. This is an iterative
algorithm that splits data into disjoint groups. Predictions P = p1, p2 ,..., pn  , pi  R d ,
i = 1,..., n , within a group are most similar to the breakdown criteria, and clusters
k  N are most distant from one another. In terms of usability, the K-means
approach defines the centers  i , i = 1,..., k of cluster sets Si , i = 1,..., k ,
k
 S = P ,
Si  S j = , i  j, i, j  k . Here t - iteration index.</p>
        <p>Algoritm: K-means
set  i , i = 1,..., k ; t  1
while</p>
        <p>i 1, k :  it   i(t−1)
p  P : p  S it
if
p −  it
2</p>
        <p>2
 p −  j(t−1) ,j 1, k
i :  it = pSit
t++
 p</p>
        <p>Sit
K-means strives to create clusters in which the fewer the variations, the more uniform
the data points in one cluster.</p>
        <p>The cluster center is a common solution of the formed group. The dispersion of
prediction of a group of models is leveled and embodied in a point solution with a
possible bias at the next hierarchy level.
6.2</p>
        <p>Experiment Results
As the base we use a linear model, presented in the form
y = xT b + 
(20)
Here b are the model parameters,  is the random error of the model. Alternative
model predictions will be obtained by generating correlated sets with relation to the
distribution of the basic random error.</p>
        <p>Set 1
Set 2</p>
        <p>Set 4
Fig. 2. Distributions of predictions based on a linear model with error modeling and a
hierarchical cluster structure of an ensemble of models on given sets of correlations
Using a visual representation of the averaged cluster solutions, one can observe a
local amplification of weakly correlated model predictions. Moreover, the general
statistical estimates of the distribution of ensemble predictions remain in acceptable
values. Compared to the random error of the linear model, the hierarchical structure of
the ensemble gives a significant advantage from the point of view of classification.</p>
        <p>We determine the effect of the size of the set of models in the ensemble on the
change in the results of the ensemble. To do this, we fix the distribution of the
generated data based on the seed parameter, and change the set of the correlation parameter.</p>
        <p>Examples of sets of randomly generated normal distribution data visually
demonstrate clustering using. Cluster decisions are generalized by the centers of these
clusters, the values of which are determined using the K-means algorithm.</p>
        <p>Cluster decisions are passed to the next level of the hierarchy. These experimental
studies applied two levels of ensemble hierarchy. When making an ensemble
decision, cluster predictions are averaged.</p>
        <p>Set 1</p>
        <p>Set 3
Fig. 3. Change of the ensemble predictions based on the structure of the set of correlation
parameter on a fixed data distribution
Variation of the predictions of ensemble models makes it possible to reduce the
influence of a random prediction error on the ensemble result. This confirms the
effectiveness of the use of group decision-making methods. At the same time, a prediction is
divided into localization levels depending on the architecture of the ensemble
construction. This indicates the fact that local model errors practically do not affect the
global characteristics of the ensemble. The hierarchical structure allows enhancing the
influence of the model with relative deviations from the agglomeration groups. Thus,
the redundancy of the models is leveled by reducing their influence.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>Building an ensemble on a hierarchy of clusters, based on the correlation of model
decisions, has several advantages.
─ The hierarchy of clusters creates localization of predictions at the localization level
with a constant global estimation.
─ The proposed ensemble structure allows managing decisions based on the level of
decision.
─ The introduction of decision levels allows creating a control mechanism for the
influence of local decisions on the global result. This allows to locally change
decision parameters without affecting global prediction.
─ A controlled mechanism of globalization of local predictions is being created.
─ The influence of model predictions is ranked depending on the correlation
properties of recognized information content.
─ Correlation of the solution is the main factor forming clusters.
─ The role of model predictions with a strong correlation of the solution decreases.
─ The influence of the identified low information content on the global ensemble
solution is increasing.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Carney</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>Diversity versus quality in classification ensembles based on feature selection</article-title>
          .
          <source>ECML</source>
          . Springer,
          <fpage>109</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Andiojaya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Demirhan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>A bagging algorithm for the imputation of missing values in time series</article-title>
          .
          <source>Expert Systems with Applications</source>
          ,
          <volume>129</volume>
          ,
          <fpage>10</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2016</year>
          ,
          <article-title>August)</article-title>
          .
          <article-title>Xgboost: A scalable tree boosting system</article-title>
          .
          <source>In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining</source>
          (pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pavlyshenko</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2018</year>
          ,
          <article-title>August)</article-title>
          .
          <article-title>Using Stacking Approaches for Machine Learning Models</article-title>
          .
          <source>In 2018 IEEE Second International Conference on Data Stream Mining &amp; Processing (DSMP)</source>
          (pp.
          <fpage>255</fpage>
          -
          <lpage>258</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cavalcanti</surname>
            ,
            <given-names>G. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>L. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moura</surname>
            ,
            <given-names>T. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Carvalho</surname>
            ,
            <given-names>G. V.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Combining diversity measures for ensemble pruning</article-title>
          .
          <source>Pattern Recognition Letters</source>
          ,
          <volume>74</volume>
          ,
          <fpage>38</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Shipp</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kuncheva</surname>
            ,
            <given-names>L. I.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Relationships between combination methods and measures of diversity in combining classifiers</article-title>
          .
          <source>Information fusion</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ),
          <fpage>135</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Brusco</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cradit</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Steinley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Combining diversity and dispersion criteria for anticlustering: A bicriterion approach</article-title>
          .
          <source>British Journal of Mathematical and Statistical Psychology.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>M. A. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Didaci</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Fumera</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Using diversity for classifier ensemble pruning: an empirical investigation</article-title>
          .
          <source>Theoretical and Applied Informatics</source>
          ,
          <volume>29</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bian</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Ensemble Pruning Based on Objection Maximization With a General Distributed Framework</article-title>
          .
          <article-title>IEEE transactions on neural networks and learning systems.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>A cluster-based semisupervised ensemble for multiclass classification</article-title>
          .
          <source>IEEE Transactions on Emerging Topics in Computational Intelligence</source>
          ,
          <volume>1</volume>
          (
          <issue>6</issue>
          ),
          <fpage>408</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>A survey on ensemble learning</article-title>
          .
          <source>Frontiers of Computer Science</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Krak</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barmak</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Manziuk</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Using visual analytics to develop human and machine-centric models: A review of approaches and proposed information technology</article-title>
          .
          <source>Computational Intelligence.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Manziuk</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barmak</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krak</surname>
            ,
            <given-names>Yu.V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kasianiuk</surname>
            ,
            <given-names>V.S.</given-names>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>Definition of information core for documents classification</article-title>
          .
          <source>Journal of Automation and Information Sciences</source>
          ,
          <volume>50</volume>
          (
          <issue>4</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Barmak</surname>
            ,
            <given-names>A. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krak</surname>
            ,
            <given-names>Y. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manziuk</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kasianiuk</surname>
            ,
            <given-names>V. S.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Information Technology of Separating Hyperplanes Synthesis for Linear Classifiers</article-title>
          .
          <source>Journal of Automation and Information Sciences</source>
          ,
          <volume>51</volume>
          (
          <issue>5</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Krak</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barmak</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manziuk</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kulias</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2019</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Data Classification Based on the Features Reduction and Piecewise Linear Separation</article-title>
          .
          <source>In International Conference on Intelligent Computing &amp; Optimization</source>
          (pp.
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          ). Springer, Cham.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>