<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Quality Indicators of Decision Tree and Forest Based Models</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>National University "Zaporizhzhia Polytechnic"</institution>
          ,
          <addr-line>Zhukovsky str., 64, Zaporizhzhia, 69063</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The problem of quality model creation for models based on decision trees and forests is considered. The set of indicators characterizing properties of decision trees and forests is proposed. It allows to quantitatively evaluate such properties as diversity, equivalence, retraining, confidence in a decisionmaking, hierarchy, equifinality, generalization, nonlinearity, robustness, homogeneity, sensitivity to the input signals, plasticity, variability adaptability, symmetry, asymmetry, emergence (integrity), interpretability (logical transparency), learnability, and autonomy as for individually considered single tree model, as for an ensemble of tree models (forest).</p>
      </abstract>
      <kwd-group>
        <kwd>decision tree</kwd>
        <kwd>forest</kwd>
        <kwd>quality</kwd>
        <kwd>property</kwd>
        <kwd>model selection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Automation of decision-making in applied tasks, as a rule, requires the construction of
a decision-making model. To solve the problem of decision-making model
constructing on the precedents, a wide class of computational intelligence methods has been
proposed, including neural networks [1-5], neuro-fuzzy networks [6-9], decision and
regression trees [10-17], forests of decision trees [18-22] and etc.</p>
      <p>Usually, the quality of such models is characterized by the error function [1, 2]. As
a result model is selected from several alternative obtained models, which has the
smallest error. Note that for each class of methods, even for the same given training
sample of observations, it is possible to obtain a wide range of different models with
acceptable accuracy. At the same time, achieving the maximum accuracy (the
smallest error) does not guarantee a high level of customer properties of the model.</p>
      <p>Earlier in [23-26] author has proposed a set of indicators, applicable to models based
on neural and neuro-fuzzy networks. However, most of these indicators are not
applicable to the models based on decision trees and forests due to their paradigm difference
from network models paradigm. Therefore, it is necessary to develop a quality model
for decision trees and forests providing comparability of its indicators with the quality
indicators of models based on neural and neuro-fuzzy networks proposed earlier in
[2325].</p>
      <p>The properties of models can be affected not only by the structural parameters, but
also by the properties of the training sample [27-31]. Therefore, it is necessary to take
into account information about the properties of the sample when determining the
quantitative indicators characterizing the properties of models.</p>
      <p>The aim of this work was to create a quality model for models based on decision
trees and forests as a set of quality indicators.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Formal problem statement</title>
      <p>Let we have a training sample of observations in the form &lt;x, y&gt;, where x = {xs},
xs={xsj}, y={ys}, s = 1, 2, ..., S, j = 1, 2, ...., N, xs is a set of input values (descriptive)
features of s-th sample instance, xsj is a value of j-th input feature of s-th sample
instance, y is an output feature value for the s-th instance in a sample, S is a number of
instances in the training sample, N is a number of input (descriptive) features
characterizing the instances of training sample.</p>
      <p>Then the problem of model building for the dependence y=f(w, x) on a sample of
observations &lt;x, y&gt; based on the decision tree tree consists in identifying the model
structure f and the values of its parameters w that provide an acceptable value of the
given quality functional of the model F(x, y, f, w) [17].</p>
      <p>The problem of model building based on a forest of decision trees, in turn, can be
represented as a problem of obtaining a set of models forest={treet}, treet =ft(wt, x),
that provide an acceptable value of a given quality functional F(x, y, {ft},{wt}), where
t is a number of a tree in the forest, t = 1, 2, ..., T, T is a number of trees in the forest, f
is a model structure of a t-th tree , wt is a set of model parameter values of a t-th tree
[20].</p>
      <p>It is obviously, that the problem of the quality functional creation of models based
on decision trees and forests requires the determination of a set of indicators {Ii} that
quantitatively describers the properties of the models.
3</p>
      <p>Primary model characteristics
Along with the sample parameters described above, we will use such notation for the
characteristics of the samples: &lt;xtest, ytest&gt; is a test sample, Stest is a number of
instances in the test sample; N, N' are, respectively, the number of signs in the original
set and in the reduced set of features; fmax , fmin are, respectively, the maximum and
minimum boundary values of a model output, y max , y min are, respectively, the
maximum and minimum boundary values of output feature.</p>
      <p>The basic properties characterizing the model structure are defined as: M is a
number of levels in a tree, Nn is a number of nodes in the model, N is a number of nodes
in the  -th layer of a tree, N nmax is a maximum possible number of nodes in a model,
fi is an i-th node function,  min is a smallest possible change of the real number,
taking into account the bit grid of the computer, o ( j) is a complexity of the j-th node,
which can be defined similarly to [250] in units of elementary operations of addition
and multiplication, ia*ut is a characteristic of autonomy of a formation of i-th element
of a model structure ( ia*ut = 0, if the inclusion (or exclusion) of i-th node to the model
determined only by the human; ia*ut = 1, if the inclusion (or exclusion) of i-th node to
the model is automatically defined by the training method; ia*ut = 0.5, if the inclusion
(or exclusion) of i-th node to the model can be determined by the human or training
method), np (i) is a characteristics of plasticity of i-th node of the tree, which is
equal to the number of possible states of a node i (for leaf nodes containing singleton
(not containing functions) np (i) = 1, for the nodes with functions the np (i) should
be taken equal to the number of different functions that may contained in the node, for
the rest nodes of the tree we should take np (i) as a number of branches of the node),
wij is a connection of i-th and j-th nodes ( wij = 0, if i-th and j-th nodes are not
connected, and wij = 1, if i-th and j-th nodes are connected).</p>
      <p>Let introduce the notation for the description of a model parameters: wmax , wmin
are, accordingly, the maximum and minimum possible values of a model parameters,
Nw is a number of adjustable parameters of model node, N max is a maximum possible
w
number of adjustable parameters of a model nodes, wim,jax , wim,jin are, respectively, the
maximum and minimum possible values of j-th parameter in i-th node, w
is a
smallest possible change in weights taking into account the size of the computer bit
grid, wj is a j-th model parameter, Nw(i) is a number of parameters of i-th model node,
wi, j is a minimum possible change of j-th parameter of i-th node taking into
account the bit grid size of a computer, sp(i) is a characteristics of a plasticity of
parameters of i-th node ( sp(i) = 0 if the node has no adjustable parameters, otherwise</p>
      <p>Nw (i)
set: sp (i)   round((wim,jax  wim,jin ) / wi, j ) ),</p>
      <p>j1</p>
      <p>We define the notation for describing the functioning of a model as: Etr, Еtest are,
respectively, model errors for the training and test samples, E(w) is a model error at a
set of weights w.</p>
      <p>The following notation will be used for the model training method parameters:
N met is a the number of training method parameters, Nmauett is a number of parameters
of a training method, which values of are determined automatically without a human
intervention, aut (wi ) is a characteristic of autonomy of i-th node parameter values
setting ( aut (wi ) = 0, if parameter values set only by a human; aut (wi ) = 1, if the
values of the node parameters are determined automatically by the training method;
aut (wi ) = 0.5, if the values of the node parameters can be determined by the human
and method).</p>
      <p>For tree models in a forest we denote: w f ,max , w f ,min , accordingly, maximum and
minimum possible values of parameters of a forest model, Nwf is a number of
parameters of a forest model without taking into account the number of parameters of its
trees, T max is a maximum possible number of trees in the forest.</p>
      <p>If necessary to distinguish the use of indicators for the decision tree we will use
notations in the form of I t or I tree (here I is an indicator, t is a tree number, tree is a
tree symbol), and for the forest we will use notation of the form of I forest (here I is
an indicator, forest is a forest symbol).
4</p>
    </sec>
    <sec id="sec-3">
      <title>Model Quality Indicators</title>
      <p>Diversity is defined by a number of different states of a system. In accordance with
the law of "Requisite variety" of W.R. Ashby, creating of a system able to decide a
problem, which has certain known diversity (complexity), it is necessary to provide
for a system an even greater diversity (knowledge of solving methods) than the
diversity of the problem being addressed, or to ensure the ability of the system to create
this diversity within itself (it would have methodology, could have developed a
methodic or proposed new methods for solving the problem) [32].</p>
      <p>
        The absolute indicator of the limiting diversity of the synthesized model based on
the decision tree tree, by analogy with [24], is defined as (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ):
      </p>
      <p>Idiv (tree)  round wmax  wmin Nw Nnmax</p>
      <p> w  i1 (np (i)).</p>
      <p>Idiv (tree,  x, y )  Idiv (tree) ;</p>
      <sec id="sec-3-1">
        <title>Idiv (x, y)</title>
        <sec id="sec-3-1-1">
          <title>Idiv ( forest,  x, y )  Idiv ( forest) ;</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Idiv (x, y)</title>
        <p>
          – in relation to the population universe as (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) and (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ):
        </p>
        <p>
          The absolute indicator of the limiting diversity of the synthesized model based on
the forest of decision trees forest is defined as (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ):
        </p>
        <p>I div ( forest )  round w f ,max  w f ,min  Nwf T max</p>
        <p> w  t1 (I div (treet )).</p>
        <p>The greater the value of limiting diversity, the wider the range of models we can
obtained on its basis.</p>
        <p>
          By analogy with [24], we define the diversity indicators of the tree model tree and
forest model forest:
– in relation to the training sample as (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) and (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ):
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
I div (tree, X ,Y )  I div (tree) ;
        </p>
        <p>I div ( X ,Y )</p>
        <sec id="sec-3-2-1">
          <title>I div ( forest, X ,Y )  I div ( forest) .</title>
          <p>
            I div ( X ,Y )
(
            <xref ref-type="bibr" rid="ref5">5</xref>
            )
(
            <xref ref-type="bibr" rid="ref6">6</xref>
            )
(
            <xref ref-type="bibr" rid="ref7">7</xref>
            )
(
            <xref ref-type="bibr" rid="ref8">8</xref>
            )
(
            <xref ref-type="bibr" rid="ref9">9</xref>
            )
          </p>
          <p>The more the value of Idiv(tree,&lt;x, y&gt;) for a single tree, and the value of
Idiv(forest,&lt;x, y&gt;) for the forest, the more will be the model potential for
approximating the relationship represented by the sample. The smaller the value of the
corresponding indicator for the model at an acceptable level of error E , the better the
approximation of the sample is.</p>
          <p>The more the value of Idiv(tree, X, Y) for a single tree, and the value of Idiv(forest, X,
Y) for a forest, the more the model will be able to solve the given problem. However,
if the relevant indicator is greater than one, or is near to one, then the model is too
excessive for the problem solving.</p>
          <p>The equivalence of models is determined as follows: two models are equivalent if
they have the same sets of answers (they respond equally to the same input stimuli)
[33].</p>
          <p>
            The equivalence coefficient of trained models based on decision trees t1 and t2 for
the sample &lt;x, y&gt; we defined as (
            <xref ref-type="bibr" rid="ref7">7</xref>
            ):
          </p>
          <p> 1 S</p>
          <p>Ieq (t1, t2 )  exp  S s1  ft1 (xs )  ft2 (xs )2 .</p>
          <p>The equivalence coefficient of trained forests models can be determined similarly
to the above, replacing the calculated tree outputs with the corresponding forest
outputs.</p>
          <p>The values of the equivalence indicator will be in the range from zero to one: the
more similar the responses of the models with the same input influences, the greater
the value of the equivalence coefficient.</p>
          <p>
            Retraining of a model for the training sample x relatively to the test sample
&lt;xtest, ytest&gt;  &lt;x, y&gt; may be defined as (in the given form with substitutions) [24]:
– for classification problems as (
            <xref ref-type="bibr" rid="ref8">8</xref>
            ):
(x, xtest ) 
1 S
          </p>
          <p> 1 f (xs )  y s 
S x1
1 Stest</p>
          <p> 1 f (xtsest )  ytsest ;</p>
          <p>
            Stest x1
– for the evaluation problems as (
            <xref ref-type="bibr" rid="ref9">9</xref>
            ):
(x, xtest) 
          </p>
          <p>1  f (xs )  ys 
1 S
S x1
1 Stest</p>
          <p>1  f (xtsest)  ytsest .</p>
          <p>Stest x1
where  is an error threshold.</p>
          <p>
            Since the error threshold for an instance in practice cannot always be set, as well as
for greater universality and uniformity in solving various problems we define it as
(
            <xref ref-type="bibr" rid="ref10">10</xref>
            ):
(
            <xref ref-type="bibr" rid="ref10">10</xref>
            )
(
            <xref ref-type="bibr" rid="ref11">11</xref>
            )
(x, xtest )  
1 S  
          </p>
          <p>exp 
S s1</p>
          <p>f (xs )  ys
 max ( ys )  min ( ys )   
  s1,2,...,S s1,2,...,S
2 

1 Stest  </p>
          <p>exp 
Stest s1</p>
          <p>f (xtsest )  ytsest.</p>
          <p> max ( ytsest ) 
  s1,2,...,S</p>
          <p>s  .
min ( ytest )  
s1,2,...,S
2 </p>
          <p>These indicators can be used not only for a model based on a single tree, but also
for a forest of decision trees, using as f the output value determined by the ensemble
of forest trees.</p>
          <p>The higher the value of the retraining indicator, the worse the approximating
properties of the model for data that did not used in the training.</p>
          <p>Confidence in a decision-making is a subjective assessment by the model of the
made decision [34].</p>
          <p>
            Regarding the value at the model output for the instance xs fed to its inputs, we
determine the confidence indicator of the model in the made decision (
            <xref ref-type="bibr" rid="ref11">11</xref>
            ):
 N 
Icert(xs )  exp ( f (xs )  ys )2 exp (Cuj(xs )  xsj )2  ,
 j1 
where Cqj is a value of the coordinate on j-th feature of q-th cluster center
corresponding to the node of the tree, to which the recognized instance xs hits. For
instances that are not included in the training set, instead of ys it is possible to substitute
the value of the output feature associated with the center of the corresponding cluster.
          </p>
          <p>It is possible to estimate the coordinates of the cluster centers for leaf nodes on the
basis of a training sample:</p>
          <p>– for the decision tree constructed on the basis of the training sample, determine
the belonging of the training sample instances to leaf nodes;</p>
          <p>– for every q-th leaf node uq form a cluster Cq  {Cqj} of instances fallen into this
node, q = 1, 2, ..., Q , where Q is a number of leaf nodes (clusters);</p>
          <p>
            – for each j-th feature as the coordinate of the cluster center take the arithmetic
average of the coordinates of the instances of the corresponding cluster according to the
corresponding feature (
            <xref ref-type="bibr" rid="ref12">12</xref>
            ):
С qj 
1 S
q {xsj | xs  uq} , j = 1, 2, ..., N; q = 1, 2, ..., Q,
S s1
(
            <xref ref-type="bibr" rid="ref12">12</xref>
            )
          </p>
          <p>Indicators of subjective confidence of model based on a decision tree will take
values in the range from zero to one: the higher the value, the closer the properties of a
recognized instance to formed cluster center templates, and the more the model is
confident in the made decision.</p>
          <p>
            The confidence of the forest of decision trees for the instance xs is defined as (
            <xref ref-type="bibr" rid="ref14">14</xref>
            ):
where
          </p>
          <p>Icfeorrtest (xs )  Icfeorrtest (xs , k ), k  1, 2, ..., K ,</p>
          <p>k
Icfeorrtest ( xs , k )  I t</p>
          <p>cert | ft (xs )  k,t  1, 2, ...,T ,
t
ft is an estimated value of the model output of t th tree, I forest is an indicator of
cert
confidence of forest models,  ,  are symbols of operators, defining the
confik t
dence of the forest in the decision for the k-th class and the t-th tree, respectively. As
such operators it is possible to use the minimum, maximum, arithmetic mean of the
set of arguments.</p>
          <p>
            The indicator of averaged confidence of forest for sample x is defined as (
            <xref ref-type="bibr" rid="ref16">16</xref>
            ):
(
            <xref ref-type="bibr" rid="ref13">13</xref>
            )
(
            <xref ref-type="bibr" rid="ref14">14</xref>
            )
(
            <xref ref-type="bibr" rid="ref15">15</xref>
            )
(
            <xref ref-type="bibr" rid="ref16">16</xref>
            )
where S q is a number of instances of the training sample that fell into the q-th node
(cluster).
          </p>
          <p>– determine the function u( xs ) that maps the recognized instance xs to the node
number of the tree into which it fell.</p>
          <p>
            The average confidence of the decision tree for a sample x is defined as (
            <xref ref-type="bibr" rid="ref13">13</xref>
            ):
Icert (x) 
1 S
          </p>
          <p> Icert (xs ).</p>
          <p>S s1
Icfeorrtest (x) 
1 S</p>
          <p> Icfeorrtest (xs ).</p>
          <p>S s1</p>
          <p>Indicators of subjective confidence of the forest will take values in the range from
zero to one: the higher the indicator value, the more the trees ensemble confident in
made decision.</p>
          <p>The hierarchical organization of the structure, the integrity and crushability of
elements allows to build models of complex objects from simpler ones; the work of the
hierarchical structure requires that the information element in each hierarchical level
behave as a whole, but when moving from level to level it must be fragmented, and
when moving from the upper hierarchical level to the lower, this fragmentation
corresponds to the allocation of its constituent elements, and when moving from the lower
level to the top, it corresponds to the inclusion of a certain part of this element in a
more complex object [33].</p>
          <p>
            The hierarchy of the model based on the decision tree is defined by analogy with
[25] as (
            <xref ref-type="bibr" rid="ref17">17</xref>
            ):
          </p>
          <p>Ih </p>
          <p>M
2 N</p>
          <p>1
M (M  1)Nn</p>
          <p>, M  1, Nn  1.</p>
          <p>The greater the Ih value , the greater the number of hierarchical levels in the model
with respect to maximum possible number of levels for a given number of nodes Nn.</p>
          <p>
            Estimate the maximum possible number of levels. Since the maximum number of
levels in the tree will be at the minimum number of outcomes from nodes, then at
each level there should be at least one node with two outcomes, and the rest of the
nodes should be leafy. Moreover, the greatest number of levels will be achieved for
the tree, where only one node at each level (except for the last) has two outcomes, and
the rest are leafy and contain only two nodes at each level. Thus, for the deepest tree,
the number of nodes of the highest level is 1 (root), for the lower level is 2 (leaves),
for the remaining layers is 2, i.e. 1+2(M–1) = Nn . From here we get: M =0.5(Nn–1)+1.
Therefore, for the decision tree we get (
            <xref ref-type="bibr" rid="ref18">18</xref>
            ):
          </p>
          <p>Ih </p>
          <p>M
2 N</p>
          <p>1
0,25Nn3  Nn2  0,75Nn
, M  1, Nn  1.</p>
          <p>
            The hierarchy of the model based on the forest of decision trees is defined as (
            <xref ref-type="bibr" rid="ref19">19</xref>
            ):
I forest 
h
max {Iht }.
          </p>
          <p>
            t1,2,...,T
The elasticity for a function y(x) on the variable xj in [35] is defined as (
            <xref ref-type="bibr" rid="ref20">20</xref>
            ):
y
Elx j ( y)  lim y   lim y  x j ,
x j 0 x j  x j 0 x j  y
          </p>
          <p>x j
where y  y(x j  x j )  y(x j ) , xj &gt; 0, y &gt; 0.</p>
          <p>y(x j )</p>
          <p>
            The relative elasticity indicator on the variable xj of approximating function y=f(x)
realized by the model at output y trained on a training sample &lt;x, y&gt; is defined
similarly to [24] as (
            <xref ref-type="bibr" rid="ref21">21</xref>
            ):
          </p>
          <p>
            El x j ( y ) 
1 S  ~x js  ~f( ~xjs x j )  ~f( ~xjs )  ,
2S s1  x j  ~f( ~xjs ) 2 
(
            <xref ref-type="bibr" rid="ref17">17</xref>
            )
(
            <xref ref-type="bibr" rid="ref18">18</xref>
            )
(
            <xref ref-type="bibr" rid="ref19">19</xref>
            )
(
            <xref ref-type="bibr" rid="ref20">20</xref>
            )
(
            <xref ref-type="bibr" rid="ref21">21</xref>
            )
~
where f( ~x js ) is a calculated value at the output of the model when applying
normalized values of the features of s-th instance to its inputs; ~f( ~x js  x j ) is a value of the
model output when applying normalized values of features of s-th instance to its
inputs, and corrected normalized by x j value of j-th feature of s -th instance to j-th
input.
          </p>
          <p>The larger the value of elasticity indicator, the more elastic is the model. This
indicator applies both to a single tree model and for a forest model.</p>
          <p>Equifinality is a regularity of functioning and development of the system,
characterizing its ultimate capabilities [36].</p>
          <p>
            Relative equifinality of a model based on a decision tree tree defined similarly to
[24] as (
            <xref ref-type="bibr" rid="ref22">22</xref>
            ):
          </p>
          <p>Ieqf (tree,  x, y )  NNmwax NNmnax exp  1 S ( f (w, xs )  y s )2 .</p>
          <p>
            w n  S s1 
(
            <xref ref-type="bibr" rid="ref22">22</xref>
            )
          </p>
          <p>Relative equifinality will receive the largest value (top-limited by one) for those
models which have reached the maximum possible size and the number of parameters
during synthesis as well as the smallest error (bottom limited by zero) in the learning
process.</p>
          <p>
            For the forest of decision trees, we determine the relative equifinality indicator as
(
            <xref ref-type="bibr" rid="ref23">23</xref>
            ):
          </p>
          <p>T T
Ieqf ( forest ,  x, y )  T max  Ieqf (treet ,  x, y ).</p>
          <p>t1</p>
          <p>Generalization is the model’s ability to integrate partial data to determine patterns
and prolongate results, that is, after training based on the training set, to give answers
for test sample instances similar to the training sample but not included in it [37, 38].</p>
          <p>
            The generalization indicator of the decision tree for the training &lt;x, y&gt; and test
&lt;xtest, ytest&gt; samples is determined by analogy with [37, 38] as (
            <xref ref-type="bibr" rid="ref24">24</xref>
            ):
          </p>
          <p> ( f (x p*)  ytpest)2 N(xsj  xjptest)2
IG  1 exp 1 Stest  j1
 Stest p1  N(ys  ytpest)2
 



(yis  yiptest)2  0,



s  arg min N(xtj  xjptest)2, j = 1, 2, ..., N, p = 1, 2, ..., Stest.</p>
          <p>
            t1,2,...,S j1
(
            <xref ref-type="bibr" rid="ref23">23</xref>
            )
(
            <xref ref-type="bibr" rid="ref24">24</xref>
            )
          </p>
          <p>In a similar way, the generalization indicator for the forest of decision trees will be
determined.</p>
          <p>Generalization indicator will take values in the range from zero to one, and will be
the greater, the smaller the error of a model at instance recognition, and difference of
recognized instance to nearest by features instance of a training sample is more.</p>
          <p>
            The generalization indicator of the trained model is defined as (
            <xref ref-type="bibr" rid="ref25">25</xref>
            ):
          </p>
          <p>If generalization indicator is significantly greater than one, then the model shows
great ability to generalization, if the generalization indicator much smaller than one,
then the model does not shows no generalizing properties.</p>
          <p>
            Generalization indicator for a forest of decision trees is defined as (
            <xref ref-type="bibr" rid="ref26">26</xref>
            ):
I gen 
          </p>
          <p>NS
N N
w n</p>
          <p>exp((Etr  Etest )2 ).</p>
          <p>NS
I gfeonrest  T
 N t N t</p>
          <p>w n
t 1</p>
          <p>exp((Etr  Etest )2 ).</p>
          <p>The errors here are defined for the ensemble of trees.</p>
          <p>Nonlinearity is a dependency that cannot be explained by a linear combination of
variable inputs [39, 40].</p>
          <p>
            The nonlinearity indicator for classification problems is defined similarly to [39] as
(
            <xref ref-type="bibr" rid="ref27">27</xref>
            ):
(
            <xref ref-type="bibr" rid="ref25">25</xref>
            )
(
            <xref ref-type="bibr" rid="ref26">26</xref>
            )
(
            <xref ref-type="bibr" rid="ref27">27</xref>
            )
(
            <xref ref-type="bibr" rid="ref28">28</xref>
            )
Inl 
          </p>
          <p> S1 f  xp 1  xs   f (xp )
2 S S  0   S  S   .</p>
          <p>
            S(S 1) s1 ps1 jN1 xsj  xjp 2 
The nonlinearity indicator for estimation problems is defined as (
            <xref ref-type="bibr" rid="ref28">28</xref>
            ):
Inl 
 S f (xpS1  (1 S1)xs )  f (xp) 
 
S(S21) sS1 pSs1 0 Nf(mxasjx  xfmjp)in2 .
          </p>
          <p> j1 
where Inl ( x, y ) is a nonlinearity indicator of the sample, determined according to
[41, 42].</p>
          <p>~
If the indicator I nl is equal to one, then we can conclude that the model
corre~
sponds to the sample in complexity. If the indicator I nl is less than one, then the
smaller its value, the greater the effect of retraining will be, and it would show
possi~
ble redundancy of a model. If the indicator I nl value exceeds one, then the model is
not sufficient for good approximation (require additional training or change the model
structure).</p>
          <p>Robustness is a model property to reliably solve a problem when receiving
incomplete and / or damaged data. In addition, the results must be consistent, even if some
part of the model is damaged [43, 44].</p>
          <p>
            The robustness of the model on the basis of a decision tree in relation to the input
signals is defined as (
            <xref ref-type="bibr" rid="ref30">30</xref>
            ):
(
            <xref ref-type="bibr" rid="ref29">29</xref>
            )
(
            <xref ref-type="bibr" rid="ref30">30</xref>
            )
where
          </p>
          <p>I Rxb </p>
          <p> min 1, 2  
j1m,2i, n..., N  max {x sj } 
 s1,2,...,S
min {x sj }
s1,2,...,S
min {x sj } ,
s1,2,...,S
2 </p>
          <p>min 2
1 S f  xs* xbs*xbs ,b j,b1,2,...,N yis  
S s1  xsj*xsj xj </p>
          <p>x j ,
x
j </p>
          <p>min
s 1,2,..., S
{ x sj }    max {x sj } 
x  s 1,2,..., S</p>
          <p>min
s 1,2,..., S</p>
          <p>{x sj } ,...,
max {x sj }    max {x sj } 
s 1,2,..., S x  s 1,2,..., S</p>
          <p>min
s 1,2,..., S
{x sj } ,
x 0,1 is a constant that regulates the accuracy of determining a robustness indicator
on model inputs.</p>
          <p>The indicator is normalized smallest change in the input signal, resulting in a
significant increase of model error.</p>
          <p>
            The robustness of a model based on a decision tree with respect to weights
(parameters) is defined as (
            <xref ref-type="bibr" rid="ref31">31</xref>
            ):
w j ,
w j ,
w j  wmin  wwmax  wmin ,...,wmax  wwmax  wmin,
where  w  0, 1 is a constant regulating accuracy of determination of the robustness
indicator on model parameters.
          </p>
          <p>The indicator is a normalized least change in weight values, leading to significantly
increase in model error.</p>
          <p>
            The integral robustness indicator for a decision tree can be defined as (
            <xref ref-type="bibr" rid="ref32">32</xref>
            ):
          </p>
          <p>I Rb  I RbxI Rwb.</p>
          <p>The indicator IRb will take values in the range [0, 1]. The closer its value to zero,
the lower the robustness of the model, the more sensitive the model to a change in
input signals or parameter values. The closer the value of the indicator IRb to one, the
greater the robustness of the model, the less sensitive the model to changes in input
signals or parameter values.</p>
          <p>In this way, robustness indicators for the forest of decision trees can be determined.</p>
          <p>The homogeneity of the elements lies in the fact that models are built from many
simple unified standard elements that perform elementary actions and are
interconnected by various connections [45].</p>
          <p>
            The homogeneity of the functions of the nodes of a decision tree is defined by
analogy with [25] as (
            <xref ref-type="bibr" rid="ref33">33</xref>
            ):
          </p>
          <p>I hn </p>
          <p>N n N n
2  {1 | fi  f j }
i 1 j i 1</p>
          <p>N n ( N n  1)</p>
          <p>.</p>
          <p>T
 Ihtn Nnt ( Nnt  1)
Ihfnorest  t 1</p>
          <p>T
 Nnt (Nnt  1)
t 1
.</p>
          <p>
            (
            <xref ref-type="bibr" rid="ref32">32</xref>
            )
(
            <xref ref-type="bibr" rid="ref33">33</xref>
            )
(
            <xref ref-type="bibr" rid="ref34">34</xref>
            )
          </p>
          <p>The homogeneity indicator will vary from zero to one: the more its value, the more
uniform the corresponding elements of the model.</p>
          <p>
            The homogeneity of the node functions of the of the forest trees is defined as (
            <xref ref-type="bibr" rid="ref34">34</xref>
            ):
The sensitivity to the input signals is characterized by calculating the partial
derivatives of the model error function [46–48]. However, this approach is
computationally hard.
          </p>
          <p>
            The averaged normalized indicator of the sensitivity of the output of the decision
tree to a change in the input signal is defined as (
            <xref ref-type="bibr" rid="ref35">35</xref>
            ):
          </p>
          <p>1 S N
I tol  SN y max  y min  s1 j1 max 1 ,  2 ,</p>
          <p>The value of the sensitivity indicator will be in the range [0, 1]. The higher the
value of the sensitivity indicator, the stronger the model reacts to changes in the input
signal, the greater are its categorization capabilities. However, too high sensitivity
may indicate a weak model resistance to noise and interference in the input signal.</p>
          <p>The average indicator of the sensitivity of the forest to changes in the input
signal is defined as (36):</p>
          <p>Itofolrest 
1 T</p>
          <p> Itol (treet ).</p>
          <p>T t 1</p>
          <p>Plasticity determines the complexity of the model’s behavior, which is considered
as a result of the interaction of many elements, each of which limits the action of
others and is limited by others on the way to the formation of global observable behavior
[49, 50]. As an analogue of neural plasticity, where neuron nodes are considered as
plastic elements for neural network models, with respect to decision trees, we will
consider the plasticity of nodes. As an analogue of synaptic plasticity (modification of
the strength of the synaptic connection between nodes, implemented by the scales in
neural network models) as applied to the decision tree, we will consider the plasticity
of tunable parameters of tree nodes.</p>
          <p>The relative indicator of plasticity of the nodes of the model by analogy with [25]
is defined as (37):</p>
          <p>Nn
 np (i)
Inp  i1</p>
          <p>Nnmaxnmpax .</p>
          <p>The indicator Inp will take values in the range from zero to one: the greater its
value, the higher the level of plasticity of the model nodes.</p>
          <p>The relative plasticity indicator of the adjustable model parameters, by analogy
with [25], is defined as (38):
(36)
(37)</p>
          <p>The coefficient Isp will take values in the range from zero to one: the bigger its
value, the higher the level of plasticity of the model parameters.</p>
          <p>The relative indicator of plasticity of a model is defined as (39) [25]:</p>
          <p>The relative indicator of plasticity will take values in the range from zero to one:
the greater its value, the higher the level of plasticity of the model and, therefore, it
has better adaptive abilities.</p>
          <p>For the forest of decision trees, we define the plasticity indicators as (40)–(42):
Isp </p>
          <p>Nn
 sp (i)
i1
Nn2round wmax  wmin </p>
          <p>
 w 
.</p>
          <p>Variability is an ability to obtain several different models for approximating
dependencies from the same data sample using the same method [51-53]. As applied to
decision trees, the variability of models is determined by the choice of a feature for
the root node and the order in which features are added for checks at other nodes, the
method of determining threshold values in nodes, etc.</p>
          <p>The absolute indicator of the variability of the model we define similarly to [25] as
(43):</p>
          <p>I pl  Inp Isp .</p>
          <p>Infporest 
I forest 
sp
I pfolrest 
1 T</p>
          <p> I ntp ,
T t 1
1 T</p>
          <p> Istp ,
T t 1
1 T</p>
          <p> I tpl .</p>
          <p>T t 1</p>
          <p>M N
Iv    v (n(, i))v (w(, i)),</p>
          <p>1 i1
where v (n(, i)) is a variability of the verification of the i-th node of  -th layer of
the model: v (n(, i)) = 1, if a non-random feature hit into the node; v (n(, i)) =N,
if a feature for checking in the node selected as random from all original feature
set v (n(, i)) = N* , if a feature for checking in the node is randomly selected from
the set of N* not yet considered features (44):
(38)
(39)
(40)
(41)
(42)
(43)
where v (w(, i)) is a variability of determining the values of the parameters of i-th
node of  -th model layer: v (w(, i)) is equal to the number of parameters in the
node that can be configured non-deterministically, if all parameters in the node
depend on previous nodes, then v (w(, i)) = 1.</p>
          <p>The more the Iv value, the more different models can be obtained based on the
corresponding paradigm.</p>
          <p>The absolute indicator of the variability of the forest of decision trees is defined as
(45):</p>
          <p>Ivforest  T Ivt .</p>
          <p>t1</p>
          <p>Noise resistance is the property of the model to provide the correct response to an
input signal containing noise [54].</p>
          <p>The indicator of the resistance of the trained model to additive noise in the input
signal at the j-th input is defined similarly to [24] as (46):</p>
          <p>
  exp 
Itol j

1 S 
2 s1 1  2 ,
1  ( f ( x s )  f (( j  ) x s ))2 , 2  ( f ( x sj )  f (( j ) x s ))2 ,
( j ) xgs  xgs   max x p </p>
          <p> p1,2,...,S g
xgs , g  j, g  1,2,..., N ;
( j ) xgs  xgs   max x p </p>
          <p> p1,2,...,S g
xgs , g  j, g  1,2,..., N ;
min x p , g  j,
p1,2,...,S g 
min x p , g  j,
p1,2,...,S g 
 </p>
          <p> sp1ms,21,i.,n....,.S,S;xsj  x jp y s  y p 
min  .</p>
          <p>j1,2,...,N  pm1,2a,.x..,Sx jp  pm1,2i,n...,Sx jp  
where  is a given noise level, 0 &lt;  &lt; 1. In order to automate the process of 
setting it is proposed to use such formula (47):</p>
          <p>The greater the value of the indicator of model resistance to noise in the input
signal on j-th input, the less important this input to the decision making.
(45)
(46)
(47)</p>
          <p>This indicator is applicable both to a model based on a single decision tree and to a
forest-based model.</p>
          <p>Adaptability is a property of structures to dynamically and independently change
their behavior in response to an input stimulus [55]. In relation to a model based on
decision trees, adaptability is determined, first of all, by plasticity, which determines
the resources for adaptation: the greater the plasticity, the more adaptive the
properties of the model.</p>
          <p>Plasticity is a necessary but insufficient prerequisite for adaptability. Along with
plasticity, the adaptive properties of the model are influenced by the sensitivity of the
model, which determines the strength of the reaction of the model to the minimum
change in the values of its parameters.</p>
          <p>The adaptability indicator of a model is defined as (48):</p>
          <p>I adapt  I pl  Itol .
(48)</p>
          <p>The larger the adaptability indicator value, the greater the possibility has models to
adapt to a given task.</p>
          <p>The adaptability indicator in this way can be determined for the forest of decision
trees.</p>
          <p>Symmetry reflects the proportionality in the arrangement of the parts of the whole
in a space, the complete correspondence (in location, size) of one half of the whole to
the other half [56]. In relation to decision trees, it is possible to determine the
indicators of symmetry and asymmetry.</p>
          <p>The indicator of symmetry of the structure of the decision tree is defined as (49):
n
Isym </p>
          <p>2 Nn Nn</p>
          <p>Nn2  Nn i1 ji11 fi  f j .</p>
          <p>The greater the Insym value, the greater the symmetry of the structure of the decision
tree.</p>
          <p>The asymmetry indicator of the model structure based on the decision tree as (50):
I asym  1  I snym.</p>
          <p>n
The greater the value of Inasym the greater the asymmetry of the model structure.
The symmetry indicator of the nodes of the decision tree is defined as (51):
The higher the Iwsym value the more the symmetry of model connections.</p>
          <p>The asymmetry indicator of model connections based on the decision tree is
defined as (52):</p>
          <p>w
Isym 
1 N n Nn
N 2  1wij  wji .</p>
          <p>n i1 j1
Iasym  1  Iswym.</p>
          <p>w
(49)
(50)
(51)
(52)
The larger the Iwasym value the greater the asymmetry of the model connections.</p>
          <p>The general indicator of the symmetry of a model based on a decision tree is
defined as (53):</p>
          <p>The higher the value of Isym the bigger the symmetry of the decision tree .</p>
          <p>The general indicator of the asymmetry of a model based on a decision tree is
defined as:</p>
          <p>I sym  I snym I swym.</p>
          <p>Iasym  1 Isym.</p>
          <p>The larger the Iasym the greater the asymmetry of the model.</p>
          <p>For the forest of decision trees, the symmetry and asymmetry indicators can be
determined as the average values of the corresponding indicators of the individual trees
included in the forest.</p>
          <p>Emergence (integrity) is a regularity that manifests itself in the system in the
appearance of new properties in it, which are absent in its elements. The integrity
property is associated with the purpose for which the system is created. Let Co is an
intrinsic complexity, which is the total complexity (content) of system elements without
interconnecting them (in the case of pragmatic information, the total complexity of
the elements that affect the achievement of the goal), Cv is a mutual complexity
characterizing the degree of interconnection of elements in the system (i.e. the complexity
of its schema or structure). The degree of system integrity in accordance with [24, 57,
58] is defined as (55):</p>
          <p>The emergence of a model based on a decision tree is defined similarly to [24] as
(56):</p>
          <p>I   Сv / Co .</p>
          <p>Nn Nn
  v (i, j)
I  i1 j1</p>
          <p>.</p>
          <p>Nn
 o ( j)
j1
(53)
(54)
(55)
(56)
(57)
The larger the I value, the more holistic is the model.</p>
          <p>For a forest-based model, emergence is defined as (57):</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>I forest </title>
        <p></p>
        <p>T 2
T
 I t
t 1
.</p>
        <p>Interpretability (logical transparency) is a model property to be understandable for
human perception and analysis [59, 60]. Obviously, a model is more interpretable if it
is hierarchical, and the average number of node connections does not exceed 5-7 (this
number is caused by the peculiarities of the human psyche). Since each node in the
decision trees has only one input, it is necessary to consider mainly the number of
outcomes from the node.</p>
        <p>Heuristically we define interpretability through hierarchy and the number of nodes
[25] as (58):</p>
        <p>Ih
Iinterp.  Nn Nn
 {wij | j  i}
i1 j 1
.</p>
        <p>The level of model interpretability increases with increasing of Iinterp. value.
For the forest of decision trees, we define the interpretability as (59):</p>
        <p>Iinfoterreps.t 
1  max {I ht}
t 1,2,...,T
2Nn2
.
where L is the Lipschitz constant for the training sample [62] as (61):
L(tree) is a Lipschitz constant (complexity) of the model [63, 64], which, as applied to
the binary decision tree, is estimated as (62):</p>
        <p>Learnability is the property of a model to improve its work (to learn or adapt),
using examples to turn it to solve a particular problem [61].</p>
        <p>The decision tree model learning indicator similarly to [25] is defined as (60):
(58)
(59)
(60)
(61)
(62)
(63)
Ilr </p>
        <p>I pl L(tree)</p>
        <p>NSL</p>
        <p>,
L </p>
        <p>max
s, p1,2,....,S;
s p
 y s  y p / xs  x p ,</p>
        <p>N '
L(tree)  10 K</p>
        <p>N '/ K N ' 
N ' K     .</p>
        <p>0  2 
T
 I tpl L(treet )</p>
      </sec>
      <sec id="sec-3-4">
        <title>Ilrforest  t1</title>
        <p>The greater the value of the learnability indicator, the model has bigger potential
for solving the problem of approximating dependence y = f(x) given in tabular form.</p>
        <p>For a forest of decision trees, we define the learnability indicator as (63):
I aut 
1  N met</p>
        <p>N n
  ia*ut  aut ( wi ) 
1  N mauett  i1</p>
        <p>.</p>
        <p>N n
I afuotrest </p>
        <p>min
t 1, 2,..., T</p>
        <p>{ I atut }.</p>
        <p>Autonomy is an agent's ability to act without direct human intervention by control
on its own actions and internal state. Autonomy also implies the possibility of
learning based on experience [65].</p>
        <p>Since the trained computational intelligence models, as a rule, in the process of
their functioning in decision-making does not require human intervention, they are
equally have a property of autonomous functioning. However, in the process of
training the level of autonomy for different models and different training methods may
vary considerably.</p>
        <p>Therefore, we will consider further characteristics of model autonomy only in
relation to the process of its learning. Since the ability to learn is determined by plasticity,
the autonomy of learning (self-adaptivity) will be characterized by an indicator that
depends on the plasticity characteristics of the model. On the other hand, the
dependence of model learning from the human may be characterized by its influence
(portion) on the formation of the structure and parameters of the model.</p>
        <p>Combining these considerations, we obtain the indicator of autonomy of model
training method (64):
(64)
(65)
As Iaut increases, the level of model autonomy in the training process increases.</p>
        <p>For a forest of decision trees, the indicator Iaut can be defined as a smallest of the
indicators of forest trees (65):
5</p>
        <p>Integral indicators of model quality
Information quality criteria is a family of integral indicators, depending on the model
error E , the training sample volume S and the number of adjusted model parameters
Nw . They include Hannan-Quinn Criterion [66], Bayesian Information Criterion [67],
Akaike's Information Criterion [68], Corrected AIC [69], and Unbiased AIC [69]. A
number of criteria in addition to the error, the sample size and the number of
adjustable parameters also take into account the maximum possible number of adjustable
parameters N max . They include a Minimum Description Length [70] , Shortest Data
w
Description [71], Consistent AIC [69] and Mallow Criterion [72].</p>
        <p>At model constructing and comparing it is usually assumed to be identical the
sample size. Therefore, it is advisable to exclude the sample size from the comparison
criteria. At the same time, various synthesized models may not use all of the features
of the sample. Therefore the number of features used in the models N' should be seen
as an important property of the models at their comparison.</p>
        <p>On the basis of these considerations we can define integral information criterion of
a model as (66):</p>
        <p>The IIC criterion will take values in the range from zero to one. The less its value
the worse the model, and the bigger its value the better the model. Here, for different
models, the error values E and the maximum number of adjustable parameters Nwmax ,</p>
        <p>Having resulted similar, we receive (69):</p>
        <p>I ef ' 
2 arcsec  NN ' NNwmwax N nmax NS  exp  E ,
 N n N w </p>
        <p>N '  1, N nmax  1, N nmax  1.</p>
        <p>I ef ' 
2 arcsec  N 2 SN wmax N nmax </p>
        <p> exp  E ,
  N ' N w 2 N n </p>
        <p>N '  1, N nmax  1, N nmax  1.</p>
        <p>An alternative generalized efficiency indicator Ief' may be used as for comparing
models and methods for their synthesis, as to optimize the process of model building.
(68)
(69)</p>
        <p>Ief ' indicator can be used also for models based on forest of decision-trees. For the
case of ensembles, N and S will remain unchanged, and the error E, the number of
selected features N', and the number of adjustable parameters N max will be
deterw
mined for the entire ensemble of trees.
6</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussion</title>
      <p>The set of indicators proposed above is extensive and for its application in practice it
is advisable to analyze the proposed indicators.</p>
      <p>The Fig. 1 presents the classification of a set of proposed indicators characterizing
the properties of models based on decision trees and forests.</p>
      <p>Horizontally at Fig. 1, indicators are divided into groups according to the
complexity of data compilation: sample (indicators characterize the properties of the sample
and are model independent), tree model (indicators are defined for a single decision
tree model), forest (indicators are defined for a set of decision tree models of a forest).</p>
      <p>The more to the right the indicator is located horizontally at Fig. 1, the higher the
level of complexity of data generalization by the model is required to determine it.</p>
      <p>Vertically at Fig. 1, indicators are divided into groups according to the level of
computational complexity relative to primary characteristics: basic properties (easily
identifiable characteristics of data and models), primary indicators (indicators
determined on the basis of basic properties), secondary indicators (indicators determined
on based on primary indicators), and integrative indicators (indicators determined on
the basis of indicators of previous levels).</p>
      <p>The higher the level of the indicator, the more difficult it is to calculate it with
respect to the primary properties of the data and models.
7</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The problem of creation of a quality model for models based on decision trees and
forests is solved.</p>
      <p>The set of indicators characterizing properties of decision trees and forests is
proposed. It allows to quantitatively evaluate such model properties as diversity,
equivalence, retraining, confidence in a decision-making, hierarchy, equifinality,
generalization, nonlinearity, robustness, homogeneity, sensitivity to the input signals, plasticity,
variability adaptability, symmetry, asymmetry, emergence (integrity), interpretability
(logical transparency), learnability, and autonomy as for individually considered
(single) tree model, as for a an ensemble of tree models (forest).</p>
      <p>The prospects of further study are to obtain estimates of the computational (time)
and spatial (memory) complexity of calculating the proposed indicators, to conduct an
experimental study of the set of the proposed indicators for assessing the properties of
models in solving practical problems of diagnosis and automatic classification on
features, to identify the relationships between different indicators of the properties of
models based on decision trees and forests.
I Rb
Iinterp.</p>
      <p>Icert (x)
Idiv(tree, x, y )
Idiv (tree, X ,Y )
I div (tree)
Icert(xs )
I
h
I
I
x
Rb
w
Rb
I tol</p>
      <p>
Itol j
max
n

N'
M
Nn
N
N
f
fi

f
max
fmin</p>
      <p>min
o ( j)
i*</p>
      <p>aut
 np (i)
sp(i)</p>
      <p>I '
ef
~
I nl
I
gen
I
lr
I
adapt
IG
I
nl
I aut
I

I
v
I
hn
Ieq (t1, t2 )</p>
      <p>w
w
w
max
max
w
Nw
N max
w
w
ij
max
wi, j
wim,jin
wj
Nw(i)
w</p>
      <p>i, j
Ieqf (tree,  x, y )</p>
      <p>L(tree)
Level</p>
      <p>Sample</p>
      <p>Tree model</p>
      <p>Forest model
Fig. 1. Analysis of decision tree and forest indicators</p>
      <p>Iasym
I
sym
n
Iasym
w
I asym
I pl</p>
      <p>E
I
n
sym
I
w
sym
I
I
np
sp
El x ( y )</p>
      <p>j
( x, xtest )
v (n(, i))
v (w(, i))</p>
      <p>N</p>
      <p>met
N aut</p>
      <p>met
aut (wi )</p>
      <p>Idiv ( forest,  x, y )</p>
      <p>I div ( forest, X ,Y )</p>
      <p>I forest ( x)
cert
I forest ( xs )
cert</p>
      <sec id="sec-5-1">
        <title>I pfolrest</title>
        <p>Idiv ( forest )</p>
        <p>Icfeorrtest ( xs , k )
Ieqf ( forest ,  x, y )
I forest
h</p>
      </sec>
      <sec id="sec-5-2">
        <title>I gfeonrest</title>
        <p>I afuotrest
I forest
lr
I forest
interp.</p>
        <p>I
forest

I forest
v</p>
      </sec>
      <sec id="sec-5-3">
        <title>Infporest</title>
        <p>I forest
sp</p>
      </sec>
      <sec id="sec-5-4">
        <title>Ihfnorest</title>
        <p>Itofolrest
w
w
treet
f ,max
f ,min
N</p>
        <p>wf
T max
36. Beven, K.J., Freer, J.: Equifinality, data assimilation, and uncertainty estimation in
mechanistic modelling of complex environmental systems. Journal of Hydrology, 249: 11–29
(2001).
37. Pedrycz, W.: Fuzzy modelling: paradigms and practice. Springer, Berlin (1996).
38. Hoekstra, A.: Generalisation in feed forward neural classifiers. Technische Universiteit</p>
        <p>
          Delft, Delft (1998).
39. Hoekstra A., Duin, R.: On the nonlinearity of pattern classifiers. In: Proceedings of 13
International conference on Pattern recognition, Vienna, 25-29 August 1996, vol. 4, pp. 271–
275. IEEE, Los Alamitos (1996)
40. Grossberg, S.: Nonlinear neural networks: principles, mechanisms, and architectures.
Neural Networks, 1(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ): 17-61 (1988).
41. Subbotin, S. A.: The training set quality measures for neural network learning. Optical
Memory and Neural Networks (Information Optics), 19 (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ): 126–139 (2010). doi:
10.3103/s1060992x10020037
42. Subbotin, S.A. The sample properties evaluation for pattern recognition and intelligent
diagnosis. In: 10th International Conference on Digital Technologies (DT 2014), Zilina, pp.
321-332. IEEE, Los Alamitos (2014). doi: 10.1109/dt.2014.6868734
43. Weng T.-W, Zhang, H., Chen, P.-Yu, Yi, J., Su, D., Gao, Y., Hsieh C.-J., Daniel, L.:
Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach.
https://arxiv.org/pdf/1801.10578
44. Carlini, N., Wagner, D.: Towards Evaluating the Robustness of Neural Networks. In: 2017
IEEE Symposium on Security and Privacy (SP), vol. 1, pp. 39-57. IEEE, Los Alamitos
(2017). DOI:10.1109/SP.2017.49
45. Krus, D.J., Blackman, H.S., Test reliability and homogeneity from perspective of the
ordinal test theory. Applied Measurement in Education, 1: 79–88 (1988).
46. Alippi, C., Piuri, V., Sami, M.: Sensitivity to errors in artificial neural networks: a
behavioral approach. IEEE transactions on сircuits and systems – I: Fundamental theory and
applications, 42 (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ): 358–361 (1995).
47. Hashem, S.: Sensitivity analysis for feedforward artificial neural networks with
differentiable activation functions. In: Proceedings of International Joint Conference on Neural
Networks, Baltimore, 7-11 June 1992, vol. I., pp. 419–424. IEEE, Los Alamitos (1992).
48. Tao C.-W., Nguyen H.T., Yao J.T., Kreinovich V.: Sensitivity analysis of neural control.
        </p>
        <p>
          In: Proceedings of Fourth International Conference on Intelligent Technologies, Chiang
Mai, 17-19 December 2003, pp. 478–482. Chiang Mai University, Chiang Mai (2003).
49. Gerrow, K. Synaptic stability and plasticity in a floating world. Current Opinion in
Neurobiology, 20 (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ): 631–639 (2010). doi:10.1016/j.conb.2010.06.010. PMID 20655734
50. Meyer, D., Bonhoeffer, T., Scheuss, V.: Balance and Stability of Synaptic Structures
during Synaptic Plasticity. Neuron, 82 (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ): 430–443 (2014).
51. Kick, D.R., Schulz, D.J.: Variability in neural networks. eLife, 7: e341532018 (2018). doi:
10.7554/eLife.34153
52. Norris, B.J., Wenning, A., Wright, T.M., Calabrese, R.L.: Constancy and variability in the
output of a central pattern generator. Journal of Neuroscience, 31:4663–4674 (2011). doi:
10.1523/JNEUROSCI.5072-10.2011
53. Masquelier, T.: Neural variability, or lack thereof. Front. Comput. Neurosci., 7: 7 (2013).
        </p>
        <p>
          doi: 10.3389/fncom.2013.00007
54. Rusiecki, A., Kordos, M., Kamiński, T., Greń, K.: Training Neural Networks on Noisy
Data. In: International Conference on Artificial Intelligence and Soft Computing (ICAISC
2014) pp. 131-142. Springer, Cham (2014).
55. Martín J.A., de Lope, J., Maravall, D.: Adaptation, Anticipation and Rationality in Natural
and Artificial Systems: Computational Paradigms Mimicking Nature. Natural Computing,
8(
          <xref ref-type="bibr" rid="ref4">4</xref>
          ): 757-775 (2009).
56. Mainzer, K.: Symmetry And Complexity: The Spirit and Beauty of Nonlinear Science.
        </p>
        <p>World Scientific, Singapore (2005).
57. Bunge, M. A.: Emergence and Convergence: Qualitiative Novelty and the Unity of</p>
        <p>
          Knowledge, University of Toronto Press, Toronto (2003).
58. Wan, P.: Emergence à la Systems Theory: Epistemological Totalausschluss or Ontological
Novelty? Philosophy of the Social Sciences, 41 (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ): 178–210 (2011).
doi:10.1177/0048393109350751
59. Japaridze, G., De Jongh, D. The logic of provability. in Buss, S., ed., Handbook of Proof
        </p>
        <p>
          Theory, pp. 476-546 North-Holland, Amsterdam (1998).
60. Molnar, Ch.: Interpretable Machine Learning. A Guide for Making Black Box Models
Explainable. https://christophm.github.io/interpretable-ml-book/
61. Valiant, L.: A theory of the learnable. Communications of the ACM, 27 (
          <xref ref-type="bibr" rid="ref11">11</xref>
          ): 1134–1142
(1984). doi:10.1145/1968.1972.
62. Iuliano, E.: A Comparative Evaluation of Surrogate Models for Transonic Wing Shape
Optimization. In: Andrés-Pérez, E., González, L.M., Periaux, J., Gauger, N.R.,
Quagliarella, D., Giannakoglou, K.C., eds., Evolutionary and Deterministic Methods for
Design Optimization and Control With Applications to Industrial and Societal Problems,
pp. 161-180. Springer, Cham (2019).
63. Scaman, K., Virmaux, A.: Lipschitz regularity of deep neural networks: analysis and
efficient estimation. In: 32nd Conference on Neural Information Processing Systems (NeurIPS
2018), Montréal, Canada.
https://papers.nips.cc/paper/7640-lipschitz-regularity-of-deepneural-networks-analysis-and-efficient-estimation.pdf
64. Lipschitz constant. Encyclopedia of Mathematics.
        </p>
        <p>http://www.encyclopediaofmath.org/index.php?title=Lipschitz_constant&amp;oldid=30687
65. Russell, S., Norvig, P.: Artificial intelligence: a modern approach. Prentice Hall, Upper</p>
        <p>
          Saddle River (2009).
66. Hannan, E. J., Quinn, B. G.: The determination of the order of an autoregression. Journal
of the Royal Statistical Society, Ser. B, 41(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ): 190–195 (1979).
67. Schwarz, G. E.: Estimating the dimension of a model. Annals of Statistics, 6 (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ): 461–464
(1978).
68. Akaike, H.: A new look at the statistical model identification. IEEE Transactions on
        </p>
        <p>
          Automatic Control, 19 (
          <xref ref-type="bibr" rid="ref6">6</xref>
          ): 716–723 (1974).
69. Gheissari, N., Bab-Hadiashar, A.: Model selection criteria in computer vision: are they
different? In: Proceedings of Digital Image Computing: Techniques and Applications,
Sydney, 10-12 December 2003, pp. 185-194. CSIRO, Collingwood (2003).
70. Grünwald, P., Myung J., Pitt, M.: Advances in Minimum Description Length: Theory and
        </p>
        <p>
          Applications. MIT Press, Cambridge (2005).
71. Rissanen, J.: Modeling by shortest data description. Automatica, 14 (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ): 465-471 (1978).
72. Mallows, C. L.: Some Comments on CP. Technometrics, 15 (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ): 661–675 (1973).
doi:10.2307/1267380
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <source>Neural Networks and Learning Machines. Pearson</source>
          , London (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bishop</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Neural networks for pattern recognition</article-title>
          . Oxford University Press, New York (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The special deep neural network for stationary signal spectra classification</article-title>
          .
          <source>In: Proceedings of 14th International Conference on Advanced Trends in Radioelectronics</source>
          , Telecommunications and Computer Engineering (TCSET
          <year>2018</year>
          ), Slavske, pp.
          <fpage>123</fpage>
          -
          <lpage>128</lpage>
          . IEEE, Los Alamitos (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1109/tcset.
          <year>2018</year>
          .8336170
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Neural network modeling of medications impact on the pressure of a patient with arterial hypertension</article-title>
          .
          <source>In: Proceedings of the International Conference on Information and Digital Technologies (IDT</source>
          <year>2016</year>
          ), Zilina, pp.
          <fpage>249</fpage>
          -
          <lpage>260</lpage>
          . IEEE, Los Alamitos (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1109/dt.
          <year>2016</year>
          .7557182
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          :
          <article-title>The neural network model synthesis based on the fractal analysis</article-title>
          .
          <source>Optical Memory and Neural Networks (Information Optics)</source>
          ,
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <fpage>257</fpage>
          -
          <lpage>273</lpage>
          (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .3103/s1060992x17040099
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>K. V.</given-names>
          </string-name>
          :
          <article-title>Neural networks and fuzzy logic</article-title>
          . S. K. Kataria &amp; Sons, New Delhi (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Rutkowska</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Neuro-Fuzzy Architectures and Hybrid Learning</article-title>
          .
          <source>Studies in Fuzziness and Soft Computing</source>
          , Springer, Berlin (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Oliinyk</surname>
            ,
            <given-names>A.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zayko</surname>
            ,
            <given-names>T.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.O.</given-names>
          </string-name>
          :
          <article-title>Synthesis of Neuro-Fuzzy Networks on the Basis of Association Rules</article-title>
          .
          <source>Cybernetics and Systems Analysis</source>
          ,
          <volume>50</volume>
          (
          <issue>3</issue>
          ):
          <fpage>348</fpage>
          -
          <lpage>357</lpage>
          (
          <year>2014</year>
          ).
          <source>doi: 10.1007/s10559-014-9623-7</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The neuro-fuzzy network synthesis and simplification on precedents in problems of diagnosis and pattern recognition</article-title>
          .
          <source>Optical Memory and Neural Networks (Information Optics)</source>
          ,
          <volume>22</volume>
          (
          <issue>2</issue>
          ):
          <fpage>97</fpage>
          -
          <lpage>103</lpage>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .3103/s1060992x13020082
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olshen</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stone</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          :
          <article-title>Classification and regression trees</article-title>
          .
          <source>Chapman and Hall</source>
          , Wadsworth, New York (
          <year>1984</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Rabcan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rusnak</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Classification by fuzzy decision trees inducted based on Cumulative Mutual Information</article-title>
          .
          <source>In: Proceedings of 14th International Conference on Advanced Trends in Radioelectronics</source>
          , Telecommunications and Computer Engineering (TCSET
          <year>2018</year>
          ), pp.
          <fpage>208</fpage>
          -
          <lpage>212</lpage>
          . IEEE, Los Alamitos (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1109/tcset.
          <year>2018</year>
          .8336188
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Geurts</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Irrthum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wehenkel</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Supervised learning with decision tree-based methods in computational and systems biology</article-title>
          .
          <source>Molecular Biosystems</source>
          ,
          <volume>5</volume>
          (
          <issue>12</issue>
          ):
          <fpage>1593</fpage>
          -
          <lpage>1605</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rabcan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levashenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaitseva</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kvassay</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Application of Fuzzy Decision Tree for Signal Classification</article-title>
          .
          <source>IEEE Transactions on Industrial Informatics</source>
          ,
          <volume>15</volume>
          (
          <issue>10</issue>
          ):
          <fpage>5425</fpage>
          -
          <lpage>5434</lpage>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1109/tii.
          <year>2019</year>
          .2904845
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kamiński</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakubczyk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szufel</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A framework for sensitivity analysis of decision trees</article-title>
          .
          <source>Central European Journal of Operations Research</source>
          ,
          <volume>26</volume>
          (
          <issue>1</issue>
          ):
          <fpage>135</fpage>
          -
          <lpage>159</lpage>
          (
          <year>2017</year>
          ).
          <source>doi:10.1007/s10100-017-0479-6</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Rabcan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Levashenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaitseva</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kvassay</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Non-destructive diagnostic of aircraft engine blades by Fuzzy Decision Tree</article-title>
          . Engineering Structures,
          <volume>197</volume>
          (
          <year>2019</year>
          ). doi:
          <volume>10</volume>
          .1016/j.engstruct.
          <year>2019</year>
          .109396
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Quinlan</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          :
          <article-title>Induction of decision trees</article-title>
          .
          <source>Machine learning</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>81</fpage>
          -
          <lpage>106</lpage>
          (
          <year>1986</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirsanova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The regression tree model building based on a clusterregression approximation for data-driven medicine</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          ,
          <volume>2255</volume>
          :
          <fpage>155</fpage>
          -
          <lpage>169</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.: Random</given-names>
          </string-name>
          <string-name>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Machine Learning</source>
          .
          <volume>45</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          (
          <year>2001</year>
          ). doi:
          <volume>10</volume>
          .1023/A:1010933404324
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Denisko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoffman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Classification and interaction in random forests</article-title>
          .
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          ,
          <volume>115</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1690</fpage>
          -
          <lpage>1692</lpage>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1073/pnas.1800256115.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A random forest model building using a priori information for diagnosis</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          ,
          <volume>2353</volume>
          :
          <fpage>962</fpage>
          -
          <lpage>973</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Boulesteix</surname>
            ,
            <given-names>A.-L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janitza</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kruppa</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>König</surname>
            ,
            <given-names>I. R.</given-names>
          </string-name>
          :
          <article-title>Overview of random forest methodology and practical guidance with emphasis on computational biology and bioinformatics</article-title>
          .
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          ,
          <volume>2</volume>
          (
          <issue>6</issue>
          ):
          <fpage>493</fpage>
          -
          <lpage>507</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yongho</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>Random forests and adaptive nearest neighbors</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          .
          <volume>101</volume>
          (
          <issue>474</issue>
          ):
          <fpage>578</fpage>
          -
          <lpage>590</lpage>
          (
          <year>2006</year>
          ).
          <source>doi:10.1198/016214505000001230</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Methods of data sample metrics evaluation based on fractal dimension for computational intelligence model buiding</article-title>
          .
          <source>In: Proceedings of 4th International ScientificPractical Conference Problems of Infocommunications Science and Technology (PICS and T 2017)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1109/infocommst.
          <year>2017</year>
          .8246136
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          :
          <article-title>Models of criterions of comparison of neural networks and neuro-fuzzy networks in the problems of diagnosis and pattern classification</article-title>
          .
          <source>Scientific Reports of Donetsk National Technical University. Serie "Informatics, Cybernetics and Computers"</source>
          ,
          <volume>12</volume>
          (
          <issue>165</issue>
          ):
          <fpage>148</fpage>
          -
          <lpage>151</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          :
          <article-title>Analysis of properties and criterions of comparison of neural network models for solving diagnostics anf pattern recognition problems</article-title>
          .
          <source>Data Registration, Saving and Processing</source>
          ,
          <volume>11</volume>
          (
          <issue>3</issue>
          ):
          <fpage>42</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          :
          <article-title>Metodics and criterions of comparison of models and algorithms of artificial neural network synthesis</article-title>
          .
          <source>Radio Electronics</source>
          , Computer Science, Control,
          <volume>2</volume>
          :
          <fpage>109</fpage>
          -
          <lpage>114</lpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliinyk</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>The dimensionality reduction methods based on computational intelligence in problems of object classification and diagnosis</article-title>
          .
          <source>Advances in Intelligent Systems and Computing</source>
          ,
          <volume>543</volume>
          :
          <fpage>11</fpage>
          -
          <lpage>19</lpage>
          (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -48923-
          <issue>0</issue>
          _
          <fpage>2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Quasi-relief method of informative features selection for classification</article-title>
          .
          <source>In: Proceedings of IEEE 13th International Scientific and Technical Conference on Computer Sciences and Information Technologies (CSIT)</source>
          , pp.
          <fpage>318</fpage>
          -
          <lpage>321</lpage>
          . IEEE, Los Alamitos (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1109/stc-csit.
          <year>2018</year>
          .8526627
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          :
          <article-title>Methods of sampling based on exhaustive and evolutionary search</article-title>
          .
          <source>Automatic Control and Computer Sciences</source>
          ,
          <volume>47</volume>
          (
          <issue>3</issue>
          ):
          <fpage>113</fpage>
          -
          <lpage>121</lpage>
          (
          <year>2013</year>
          ). doi:
          <volume>10</volume>
          .3103/s0146411613030073
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliinyk</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The sample and instance selection for data dimensionality reduction</article-title>
          .
          <source>Advances in Intelligent Systems and Computing</source>
          ,
          <volume>543</volume>
          :
          <fpage>97</fpage>
          -
          <lpage>103</lpage>
          (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -48923-0_
          <fpage>13</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Subbotin</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The instance and feature selection for neural network based diagnosis of chronic obstructive bronchitis</article-title>
          .
          <source>Studies in Computational Intelligence</source>
          ,
          <volume>606</volume>
          :
          <fpage>215</fpage>
          -
          <lpage>228</lpage>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -19147-8_
          <fpage>13</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Ashby</surname>
            ,
            <given-names>W. R.:</given-names>
          </string-name>
          <article-title>An Introduction to Cybernetics</article-title>
          . Martino Fine Books,
          <string-name>
            <surname>Eastford</surname>
          </string-name>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Dopico</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          , de la Calle,
          <string-name>
            <given-names>J. D.</given-names>
            ,
            <surname>Sierra</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. P.</surname>
          </string-name>
          :
          <source>Encyclopedia of artificial intelligence. Information Science Reference</source>
          , New York (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Luger</surname>
            ,
            <given-names>G. F.: Artificial</given-names>
          </string-name>
          <string-name>
            <surname>Intelligence</surname>
          </string-name>
          .
          <article-title>Structures and Strategies for Complex Problem Solving</article-title>
          .
          <source>Pearson Education</source>
          , London (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Nievergelt</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>The Concept of Elasticity in Economics</article-title>
          .
          <source>SIAM Review</source>
          ,
          <volume>25</volume>
          (
          <issue>2</issue>
          ):
          <fpage>261</fpage>
          -
          <lpage>265</lpage>
          (
          <year>1983</year>
          ). doi:
          <volume>10</volume>
          .1137/1025049.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>