<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Synthesis of multithreshold neural network classifier</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladyslav Kotsovsky</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>State University “Uzhhorod National University”</institution>
          ,
          <addr-line>Narodna Square 3, Uzhhorod, 88000</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Issues related to the application of multithreshold neural units in machine learning are considered in the paper. A labeled modification of the model of binary-valued multithreshold neuron is proposed. The 2-layer feedforward network architecture with the hidden layer consisting of multithreshold neurons is treated. The multicategory classifier is designed on the base of such architecture. The supervised algorithm is proposed for the synthesis of multithreshold neural networks. Simulation results are presented related to the performance of multithreshold classifier on real-world dataset compared to some popular classifiers. The impact of algorithm hyperparameters on the classifier performance, as well as on its representational capacity is discussed.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Multithreshold neural unit</kwd>
        <kwd>artificial neural network</kwd>
        <kwd>classification</kwd>
        <kwd>machine learning 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Neural networks play extremely important role in machine learning [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and intelligent data
proceeding [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. They are used as efficient components of modern smart hardware [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and software
solutions [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Powerful capabilities of artificial neural network (NN) are provided by the
possibility of the use of problem-oriented network architecture [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], appropriate supervised or
unsupervised learning techniques [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ] and the proper choice of hyperparameters of the learning
algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        During few past decades we were witnesses of numerous rises and falls of the interest in neural
computations. Early neural-like models (e.g., a model of McCulloch and Pitts artificial neuron)
used linear threshold units in which binary-valued Heaviside step function [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] was used. These
models were inspired by the brain study of the structure of biological neurons [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. But the limited
capacity of a single threshold unit [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] along with difficulty in the design of learning algorithms
for threshold networks even in case of the relatively small number of units were stumbling blocks
for an efficient application of threshold NN in machine learning and caused the decline of interest
in them [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        In order to overcome these drawbacks of early models of neurons, many improvements were
proposed [
        <xref ref-type="bibr" rid="ref12 ref2">2, 12</xref>
        ] intended to boost the recognition capability of a single threshold unit and make
easier the process of the network learning [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. They can be divided into two big categories. The
first contains models using the more complicated approach to the aggregation of the input signals
of the neural unit instead of the weighted sum of inputs in linear threshold units, e.g., polynomial
threshold units [
        <xref ref-type="bibr" rid="ref11 ref13">11, 13</xref>
        ]. The second category contains models in which the basic threshold
function is replaced by its more sophisticated modifications [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This category is much more
useful because it allows to use diverse learning techniques adapted to the particular problem or
application [
        <xref ref-type="bibr" rid="ref15">15, 16</xref>
        ].
      </p>
      <p>The models with multithreshold activation function were one of first in the second category
[17]. A multithreshold neuron has weight vector w   w1, , wn   Rn associated with inputs
x1, , xn and threshold vector t  t1, ,tk   Rk instead of a single threshold used in McCulloch
and Pitts neuron. The use of multiple thresholds allows to operate with neuron in two modes:
binary-valued and multi-valued, respectively [18]. The first mode uses the following activation:
(1)
(2)
1, if t2 j1  x  t2 j , j 1, k / 2,
ft2  x  </p>
      <p>0, if t2 j  x  t2 j1, j 0,1, k / 2.</p>
      <p>Note that bipolar output values are also acceptable for such function, if 0 is replaced by –1 in the
last row in braces in equation (1).</p>
      <p>The second mode can be preferable for application in multiclass classification [19, 20]. In this
mode the following (k 1) -valued counterpart of (1) is used
0, if x  t1,
1, if t1  x  t2 ,

ftk 1  x  ........................</p>
      <p>k  1, if tk 1  x  tk ,
k, if tk  x.</p>
      <p>Notice that both activations (1) and (2) can be used in multiclass classification [18]. But in the
case of the binary-valued activation (1) the output layer of special shape is required as well as the
employing of the “one versus all” approach for the learning of classifier [21].</p>
      <p>The main goal of the paper is the study of the model of 2-layer neural network whose hidden
layer consists of binary-valued multithreshold neurons and examination whether this model is
suitable for the multiclass classification, as well as the development of the synthesis algorithm
for such networks.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        As it was mentioned in introduction, multithreshold neural units were proposed in the early
studies in artificial intelligence [17, 21, 22]. The first models employed the binary-valued
activation in the form (1). The reason for the use of additional thresholds was the assumption that it
can improve the capability of basic (single-threshold) linear threshold units [23, 24]. The
theoretical reasons were studied in [17, 21, 23, 25] confirming this assumption and giving quantitative
expressions for it. These works dealt with the estimation of capability of multithreshold elements
to produce dichotomies of finite sets in n-dimensional real space [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. These studies confirmed the
power of multithreshold approach and its advantage over single-threshold one. But the ways of
the practical realization of this advantage remained almost unclear, because few applications of
multithreshold-based models and networks were proposed [26] despite their better potential
capacities compared to the basic linear threshold units [18, 24]. The hardness results stated in [
        <xref ref-type="bibr" rid="ref5">5,
18</xref>
        ] for 2-valued multithreshold neurons and networks shed some light on these difficulties related
to the design of multithreshold systems in the case of two or more thresholds. Nevertheless, the
lack of learning techniques for multithreshold systems implies the reduction of the interest in their
development [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7, 27</xref>
        ].
      </p>
      <p>
        Papers [
        <xref ref-type="bibr" rid="ref5">5, 28</xref>
        ] stated some recent advances in the study of bithreshold neural units and NN
that changed the situation somewhat. It was caused by new approaches in the synthesis of NN
whose hidden layer consists of bithreshold neurons, which were proposed in [
        <xref ref-type="bibr" rid="ref5">5, 24</xref>
        ] using the
generalization of the Baum’s idea [29] for threshold networks combined (see [24]) with the power
of the bithreshold activation function, which is the partial case of (1) obtained for k  2.
      </p>
      <p>
        The above observation was confirmed in [
        <xref ref-type="bibr" rid="ref1 ref3">1, 3</xref>
        ], where bithreshold neural networks were
considered as effective tools for intellectual data processing and machine learning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The drawbacks
and limitations of the basic bithreshold NN were stated in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. They were explained as the
consequence of the non-local nature of bithreshold activation function in [24]. Hybrid models of the
multiclass classifier with additional hidden layers were proposed in [29] in order to reduce the
mentioned downside of the bithreshold paradigm. It should be noted that bithreshold NN can be
useful not only for solution of classification problems. The scope of their application can be
considerably extended [
        <xref ref-type="bibr" rid="ref14 ref8">8, 14</xref>
        ] to new scopes. E.g., they were mentioned in design of powerful
deep NN providing the exponential improvement of the memorization capacity [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Moreover,
the bithreshold approach is not limited only by real-valued domain [16]. The generalization of
multithreshold paradigm is known for systems proceeding with complex-valued data [16]. The
complex bithreshold-like activations were introduced in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] in order to increase the capability of
complex-valued neural units.
      </p>
      <p>
        Notice that the mentioned advances in the use of multithreshold activation are actually related
to only networks that use bithreshold nodes in their hidden layers [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. There are applications of
networks including nodes neither with multithreshold activation (1) nor (2) in the case where the
number of thresholds k  2 [
        <xref ref-type="bibr" rid="ref5">5, 18</xref>
        ]. Papers [18, 28] clarify the reasons causing difficulties related
to the application of multithreshold approach. The learning of a single multithreshold unit with
activation (1), as well as the learning of NN consisting of neurons with activation (2) proved to
be NP-hard even in the case k  2 . Therefore, the intractability is the internal property of
multithreshold models [30].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Models and methods</title>
      <sec id="sec-3-1">
        <title>3.1. Model of multithreshold neural unit</title>
        <p>Let us consider again a model of k-threshold binary-valued neuron with the weight vector w and
(ordered) threshold vector t, which output is given by (1). It is possible to enhance this model a
bit by using a binary starting label. The performance of such unit can be described as follows:
 , if t2 j1  w  x  t2 j , j 1, k / 2,
y  
 , if t2 j  w  x  t2 j1, j 0,1, k / 2,
(3)
where x   x1, , xn  is an input vector, w  x  w1x1  ... wn xn ,  0,1 is a binary label of the
neuron,  is its negation, and y is the output value of the neuron. Further, we shall consider only
binary-valued neurons that obey (3) and their attribute “binary-valued” will be omitted. The
corresponding labeled activation function will be denoted as ft,2 .</p>
        <p>
          The labeled k-threshold neuron is completely defined by the triplet  w, t,  . As well as in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
we call this triplet the structure vector of labeled multithreshold neuron.
        </p>
        <p>Two sets A  Rn and A  Rn are called k-separable [18] if there exists a k-threshold
neuron with structure  w, t,  such that</p>
        <p>for all x  A ft,2 w  x  1 and for every x  A ft,2 w  x  0 .</p>
        <p>
          In this case labeled k-threshold neuron with the structure  w, t,  produces the k-dichotomy
 A , A  of the set A  A  A in two disjoint sets A and A . Therefore, k-threshold neurons
can be used as binary classifiers. But, as stated in [18], the computational power provided by a
single multithreshold unit is insufficient for the design of efficient multiclass classifiers [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>Model (3) has a simple geometrical interpretation [18]. The family of parallel hyperplanes
H j : w  x  t j , j 1, ..., k divides the space Rn by k  1 parts, which can be successively
labeled by numbers 0, 1, …, k. In the case   1 all points belonging to parts with even indices are
attributed as “negative”, whereas “odd” parts are considered as “positive” [18]. If   0 , then the
order of parts should be changed to opposite. The illustration is shown in Figure 1, where the case
n  2, k  4,   1 is considered, and the sets of blue (negative) and yellow (positive) points are
4-separable.</p>
        <p>Figure 1 can also illustrate the nature of difficulties related to the application of a
binaryvalued k-threshold neuron. Its value can alternate k times. It ensures the greater capability of the
neuron on the one hand, but results in the hardness of its training on the other hand.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model of 2-layer multithreshold neural network</title>
        <p>We can overcome the limitation of a single multithreshold neuron by using a NN in which
multithreshold neurons are connected layer-wise. The simplest architecture has only two computation
layers and its hidden layer consists of labeled multithreshold neurons, whereas outputs layer
contains single-threshold nodes each of them corresponds to a separate class of patterns. The
architecture of a 2-layer feedforward fully-connected multithreshold neural network is shown in
Figure 2.</p>
        <p>The NN has n nodes in the input layer, h multithreshold neurons in the hidden layer and l
threshold neurons in the output layer, where l is a predefined number of classes. Note that the
hidden layer can be a bit heterogeneous in the sense of the number of thresholds and labels of
neurons. Namely, the same number of thresholds is not required for neurons in this layer. This
assumption can make network synthesis easier and may reduce the size of memory that is
necessary to store all network parameters.</p>
        <p>Consider how multithreshold NN can be used to perform the desired partition of patterns in
ndimensional space. The example of the network performance in two dimensions is presented in
Figure 3.</p>
        <p>Circles are members of the first class C1. The second class C2 is formed by triangles, and
squares belong to the third class C3. Note that both C1 and C3 are formed by 2 compact clusters,
which are denoted using multiple occurrences of class names. The presented partition can be
produced by 2-layer multithreshold NN, which have 2-4-3 topology. The hidden layer of the
network consists of 4 multithreshold neurons each of them is responsible of the recognition of
patterns belonging to one of 3 classes. For the proper classification of patterns from class C1 a
single neuron is sufficient. The same is true for C2. The first class C1 requires the use of
multithreshold neuron with 3 thresholds, whereas a single bithreshold neuron suffices for the second
class. The proper classification of patterns belonging to the third class C3 requires the use of 2
bithreshold neurons (one of them can be even replaced by a single-threshold neuron). The
hyperplanes that form bounding surfaces of multithreshold neurons are drawn in solid lines.</p>
        <p>Each hidden neuron makes its own primary “slice” whose width depends on neuron two
thresholds, which in turn are defined by the relative position of patterns. Every multithreshold neuron
recognizes at least two patterns of corresponding class, with possibly some additional points. All
primary slices are equidistant from their pivot lines depicted in Figure 3 in dashed line. For the
neuron corresponding to the first class one more (unbounded) slice is required, which is
performed using an additional third threshold. The corresponding line is marked in Figure 3 using
an arrow. The label 0 should be assigned to this neuron, whereas all other neurons in the hidden
layer would be labeled with 1. Every single-threshold neuron in the output level can be connected
only with neurons corresponding to its own class and performs only the disjunction of its inputs.</p>
        <p>Consider the network performance in the prediction mode. Let P1 and P2 are new instances
presented to the network. Using (3) and the rule of the action of the output neurons we can
conclude that network classifies P1 and P2 as members of C1 and C3, respectively. It should be
noted that classification of patterns P3 and P4 seems less clear. We shall return to this issue later.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Network synthesis algorithm</title>
        <p>Consider the task of the synthesis of multithreshold NN. Essentially, multithreshold NN with l
outputs computes a vector mapping F : Rn  0,1l . The question arises how to design of a
computationally efficient algorithm for synthesis a network, which is able to produce a given partition
C1, ,Cl  of the finite set C into l disjoint nonempty sets C , ,Cl , i.e., for all x C
1
Fi x  1  x Ci , where i 1, ,l .</p>
        <p>
          In the algorithm design the approach to NN synthesis shall be used that is the generalization
of similar techniques used in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] for the synthesis of bithreshold NN. This approach extends an
idea of the synthesis of single-threshold NN [29] to the case of neurons with several thresholds.
The necessary number of such neurons would be inserted in the hidden layer that each of them is
capable to recognize its own portion of patterns belonging to a particular class by separating them
from patterns belonging to other classes using 2 or more parallel hyperplanes.
        </p>
        <p>Let a collection of pairs x1, y1 , ,xm , ym  be our training set, where m  m  n is its size,
xi  Rn is a training pattern provided with its label yi 1, ,l and yi  j implies that the
corresponding pattern xi belongs to the class C j , i 1, , m . Suppose that all patterns in the
sequence X  x1, , xm are in general position in n-dimensional space—that is, every subset of
n or fewer patterns consists of linearly independent elements. Assume also that
X i  X  Ci  , i  1, , m, y   y1,..., ym  . Consider the following pseudocode of the network
synthesis algorithm:</p>
        <p>MultithresholdSynthesis(X, y, α, p)
Stage 1. Synthesis of the hidden layer
1.1 Z  X1, , X l 
1.2 h  0
1.3 while Z   :
1.4 select a random set X i from Z
1.5 h  h  1
1.6 ch  i
1.7 r  min n, Xi 
1.8 Move r randomly chosen patterns from Xi into the matrix A
1.9 Solve linear system w  AT  1
1.10   min | w  x 1| x  X \ Xi 
1.11
1.12
1.13
1.14
1.15
1.16
1.17
1.18
1.19
1.20
1.21
1.22
1.23</p>
        <p>Xi  Xi \  x  Xi | w  x 1|  
t  1  , 1   
  1
foreach a  X i :
s  w  a
1  min s  w  x x  X \ Xi , w  x  s
 2  min w  x  s x  X \ Xi , w  x  s
B  x x  Xi , 1  w  x  s   2
if B  p :</p>
        <p>X i  X i \ B
if  1 is well-defined: insert s   1 in the ordered vector t
else:   0
if  2 is well-defined: insert s   2 in the ordered vector t</p>
        <sec id="sec-3-3-1">
          <title>Add multithreshold neuron  w, t,  in the hidden layer</title>
          <p>if X i   :
remove Xi from S
if</p>
          <p>ch  i : vij  2
else: vij  0</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Add threshold neuron  vi ,1 in the output layer</title>
          <p>The presented algorithm has 2 parameters X—the sequence of training patterns, y—the target
vector of class labels for all patterns (is used only in step 1.1), and two hyperparameters α and p.
The synthesis process is divided by 2 stages. The main first stage defines the creation of the
network hidden layer consisting of multithreshold neurons with variable number of thresholds.
The second stage is very simple and defines the output layer consisting of single-threshold
neurons.</p>
          <p>In the above algorithm h is the counter responsible for the current number of hidden
multithreshold neurons. After the completion of the first stage h equals the size of the hidden layer.
Inside a while loop the portion of unhandled patterns is selected, which belong to the same class.
Then r patterns are moved into rows of auxiliary matrix A. In step 1.9 the vector of length r is
denoted by 1, which consists solely of 1’s. Steps 1.8–1.12 are intended to determine the parameter
of the “main slice” of multithreshold neuron which is defined by the weight vector w from step
1.9—a normal vector of the two parallel hyperplanes that perform this slice using thresholds 1 
and 1   , respectively. Such slices are crucial for MultithresholdSynthesis, because each of them
properly classifies at least n patterns of ith class (when patterns are in general position). The
nested loop in steps 1.14–1.23 is intended to try candidates for secondary slices. Each such slice
is taken into account only if it detects at least p new properly classified patterns, where p is the
hyperparameter of the synthesis algorithm.</p>
          <p>Note also that 1 and  2 determine the width of corresponding secondary slice and can be
undefined. In this case at most a single threshold will be added into the threshold vector. It can
result in multithreshold neurons with odd number of thresholds. If 1 is undefined, then 0 is
assigned to the neuron label. It is useful in the case when all remaining class patterns are only training
instances lying in the “negative” half-space produced by corresponding hyperplane w  x  s .</p>
          <p>
            As mentioned in [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ], the hyperparameter α can be considered as a tolerance measure of
multithreshold NN. If  (0,1] , then the network performs well on training set (training error is almost
0) but it can be “overfitted” and less effective outside the training data. If α is slightly greater than
1, then the network can have the better generalization ability. Moreover, the use of α &gt; 1 can
result in the reduction of the number of neurons in the hidden layer.
          </p>
          <p>In step 2.5 we denote by (vi, 1) the ith single-threshold neuron with the weight vector vi 
vi1, , vih  and the unit threshold.</p>
          <p>The following proposition estimates the complexity of the considered algorithm.</p>
          <p>Proposition. An arbitrary partition of m-set X into l disjoint nonempty sets X1, …, Xl of
ndimensional vectors can be performed by a 2-layer multithreshold NN with at most m / n  l
nodes in the hidden layer and l neurons in the output layer, whereas the network can be
synthesized using MultithresholdSynthesis(X, y, α, p) in O  m3  time.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment</title>
      <p>Let us consider the capability of MultithresholdSynthesis to synthesis a multithreshold multiclass
classifier to solve the classification problems compared to classical classifiers, such as k-nearest
neighbor, decision tree, random forest, RBF SVC (support vector classifier with RBF kernel),
MLPClassifier (multilayer perceptron or fully-connected feed-forward NN, using softmax
function for estimates the probability of the target class over all possible target classes and
crossentropy loss function). Classifiers were tested on the “Wine Quality” dataset [31], related to red
and white variants of the Portuguese "Vinho Verde" wine provided by “UC Irvine Machine
Learning Repository” [32], as well as other popular repositories, e.g., [33]. This medium-sized
dataset contains 6497 instances, 4898 of which are related to white wine, and other 1599—to red
one [31].</p>
      <p>The data includes only 12 features, 11 physicochemical properties as inputs (see [34]), wine
color and one sensory quality score as the output from 0 to 10 (actually, for red wine the quality
is between 3 and 8, and for white wine—between 3 and 9). The distribution of qualities is shown
in Figure 4.
Note that only 9 instances have the quality value 9 and they are almost invisible in Figure 4. The
most frequent quality value is 6 with 2836 representatives. Thus, the classes are very imbalanced.
Moreover, not all input features are supposed to be quite relevant [31]. Thus, the training task for
this dataset can be considered hard enough to estimate classifier ability [34].</p>
      <p>
        Dataset contains only numerical data (except the color value). Thus, the feature extraction was
very simple. Data were standardized by removing the mean and scaling to unit variance using
standard scaler from Scikit Learn [
        <xref ref-type="bibr" rid="ref2">2, 35</xref>
        ]. 80% of instances of every dataset were used as the
training set (namely, 5198 instances), and the rest 20%—as the test set for the final evaluation of
classifier results. In order to avoid the dependence of the selection of patterns for the training,
random sub-sampling validation was used [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], i.e., randomly chosen testing instances keeping the
same proportion 1:3 between instances concerning red and white wines, respectively.
      </p>
      <p>For every classifier stratified 5-fold cross-validation was used in order to evaluate the classifier
performance and to tune its hyperparameters. The average performance measure over all 5 splits
and grid search was employed for this purpose.</p>
      <p>For k-nearest neighbor classifier only the number of neighbors was explored. For multilayer
perceptron Adam optimizer was used and the number of hidden layer as well as their sizes, the
kind of activation function and batch size were investigated. For support vector classifier with
RBF kernel different combinations of values were tried of regularization parameter C, degree of
the polynomial kernel function and kernel coefficient gamma [35]. In the case of random forest
classification, the number of trees in the forest and the maximum depth of the tree were chosen.
The last parameter also was studied for decision tree classifier.</p>
      <p>Both hyperparameters α and p were explored for MultithresholdSynthesis(X, y, α, p). {0.7,
0.8, 0.85, 0.9, 0.95, 1, 1.05, 1.1, 1.15, 1.2} and {2, 3, 4, 5, 6, 7} ranges were used for α and p,
respectively. During the experiments we used implementations of 5 above-mentioned popular
classifiers provided by Scikit-Learn library [35].</p>
      <p>Two classification metrics were used during simulation. The first was the accuracy, the
second—top-2 score (successful when the target class is within the top 2 predictions) [35].</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>The following two tables contain results of experiments in the case of first and second metrics,
respectively. For first 5 classifiers best between all specified combinations of hyperparameter
values are not shown for sake of brevity (and because their study is beyond the main scope of the
paper). Note that all classifiers were retrained one more time on full training set after the finish
of cross-validation and grid search. That is why the values of metrics on the test set are often
higher than average across all splits during cross-validation.</p>
      <p>Results for two versions of multithreshold 2-layer NN are presented in last two rows in both
tables. In the synthesis of NN1 α = 1 was used. NN2 was synthesized using the best value
α* = 1.05 found by the grid search. In both networks the best value p* = 4 was used.</p>
      <p>By analyzing the performance results shown in Table 1 and Table 2, we can conclude that:









</p>
      <p>The basic version of multithreshold network NN1 was good memorizer and performed
perfectly on the training set.</p>
      <p>The generalization ability of NN1 was very poor. Its performance on the test set was worst
for both considered metrics.</p>
      <p>The optimized version NN2 lost a bit in the accuracy on the training set, but performed
considerable better on new patterns.</p>
      <p>The proper choice of the α allow us to avoid the great overfitting.</p>
      <p>The test error during cross-validation varies significantly.</p>
      <p>The use of value of  lying inside the half-interval (0, 1] results in the degradation of the
generalization capability of the classifier.</p>
      <p>The most typical number of thresholds of the neuron in the hidden layer of NN2 was 3.
The growth of the value of hyperparameter p results in almost bithreshold hidden layer.
Using the modified software implementation and additional memory it is possible to
decrease the time complexity of MultithresholdSynthesis(X, y, α, p) in Proposition from
O  m3  to O m2 n  log m .</p>
      <p>The estimated hidden layer size m / n  l from Proposition is too pessimistic. The actual
size in the experiment was 231. But it seems plausible that the number of nodes in the
hidden layer growths as m / n .</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussions</title>
      <p>
        The proposed model of 2-layer multithreshold NN and corresponding synthesis algorithm can be
considered as the first step towards an application of multilayer NN based on the multithreshold
paradigm in machine learning in the case of an arbitrary number of thresholds. Even the first
applications of such networks in classification on the real-word datasets showed that they have
the similar pros and cons that bithreshold neural networks considered in [
        <xref ref-type="bibr" rid="ref5">5, 24</xref>
        ].
      </p>
      <p>The principal consequences that follow from concrete results of simulation concerning the
pattern classification on synthetic [35] as well as real-world datasets [32, 33] using multithreshold
2-layer NN suggest that networks yielded by MultithresholdSynthesis or its more complicated
modifications have good performance on the design set but are rather inefficient on new instances.
More precisely, as it was observed that for “Wine Quality” and other datasets:</p>
      <p>Making prediction, classifier can attribute new patterns to a class whose all representatives in
the training set are very distant from this pattern (“nonlocality” of classifier [24]).</p>
      <p>The size of the hidden layer of the network strongly depends on the order in which the training
patterns were presented to learner. Moreover, during the cross-validation the NN synthesized on
different splits demonstrated the significant diversity of the size of the hidden layer.</p>
      <p>The size of the hidden layer is too large in the case of the larger datasets.</p>
      <p>This downsides of the basic multithreshold 2-layer NN classifier is illustrated in Figure 3. The
network classifies pattern P3 as a member of first class despite it looks more likely be a member
of one of two other classes (first drawback). Moreover, the network is undecisive when it predicts
a class for new pattern P4, despite it is situated in the proximity of a compact cluster of members
of the first class (second drawback).</p>
      <p>
        A conjecture was stated in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that the reason for such two downsides is the nature of the
multithreshold activation function. These downsides can be eliminated by using a localized
modifications of activation function (1), e.g., the multithreshold modifications of smoothed local
activations, which were proposed in [24]. Another solution consists in the use of modified
network architectures with additional layers comprising nodes of different kinds, as it was proposed
in [30] for bithreshold neural networks. The last drawback can be partially reduced using some
additional hyperparameters in the MultithresholdSynthesis, e.g., the limitation of the hidden layer
size.
      </p>
      <p>It should be noted that the more expensive part of MultithresholdSynthesis algorithm is
computations in step 1.10, which have the O(m2n) contribution in its time complexity. These
computations can be performed simultaneously with solely synchronization of the shared variable
ε. Therefore, MultithresholdSynthesis has massively parallel structure and can be effectively
realized using multicores processors and graphical accelerators.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>The ways of the use of multithreshold paradigm in neural computation were treated in the paper.
The new modification of the model of binary-valued multithreshold neural unit was proposed,
which allows often to reduce number of thresholds. The 2-layer neural network architecture was
studied whose hidden layer consists of multithreshold neurons with different number of
thresholds. The synthesis algorithm was developed for such multithreshold NN. The multiclass
classifier was designed on the base of such architecture.</p>
      <p>Experiment results of the performance of multithreshold classifier on “Wine Quality”
realworld dataset were presented. These results clarify pros and cons of the proposed model compared
to some popular classifiers. The impact of algorithm hyperparameters on the classifier ability to
make a precise prediction, as well as and the ways to improve it were discussed. The main
conclusion is that there exists a need in the development of localized modifications of the
multithreshold activation function and additional hyperparameters in synthesis algorithm in order
to regulate the network performance more precisely.</p>
    </sec>
    <sec id="sec-8">
      <title>8. References</title>
      <p>[16] Z.-G. Zhang, Y.-L. Xiao, J. Zhong, Unitary learning in conditional models for deep optics
neural networks, in: Proceedings of SPIE – The International Society for Optical
Engineering, volume 12565, 2023, no. 1256543.
[17] R. Takiyama, Multiple threshold perceptron, Pattern Recognition 10.1 (1978): 27–30.
[18] V. Kotsovsky, A. Batyuk, Multithreshold neural units and networks, in: Proceedings of IEEE
18th International Conference on Computer Sciences and Information Technologies, CSIT
2023, Lviv, Ukraine, 2023, pp. 1-5, doi: 10.1109/CSIT61576.2023.10324129.
[19] N. Jiang, Z. Zhang, X. Ma, J. Wang, Y. Yang, Analysis of nonseparable property of
multivalued multi-threshold neuron, in: Proceedings of 2008 IEEE International Joint Conference
on Neural Networks (IEEE World Congress on Computational Intelligence), Hong Kong,
China, 2008, pp. 413-419, doi: 10.1109/IJCNN.2008.4633825.
[20] I. Prokíc, Characterization of multiple-valued threshold functions in the
Vilenkin-Chrestenson basis, Journal of Multiple-Valued Logic and Soft Computing 34.3-4 (2020): 223–238.
[21] R. Takiyama, The separating capacity of a multithreshold threshold element, IEEE
Transactions on Pattern Analysis and Machine Intelligence. PAMI-7.1 (1985): 112–116.
[22] K. Ashenayi, J. Vogh, M.R. Sayeh, B. Karimi, T. Baradaran, Multiple threshold perceptron
using sinusoidal function, International Journal of Modelling and Simulation 12.1 (1992):
22–26.
[23] S. Olafsson, Y. S. Abu-Mostafa, The capacity of multilevel threshold function, IEEE
Transactions on Pattern Analysis and Machine Intelligence 10.2 (1988): 277–281.
[24] V. Kotsovsky, A. Batyuk, Feed-forward neural network classifiers with bithreshold-like
activations, in: Proceedings of IEEE 17th International Scientific and Technical Conference on
Computer Sciences and Information Technologies, CSIT 2022, Lviv, Ukraine, 2022, pp. 9–12.
[25] N. Jiang, Y. X. Yang, X. M. Ma, and Z. Z. Zhang, Using three layer neural network to
compute multi-valued functions, in 2007 Fourth International Symposium on Neural
Networks, June 3-7, 2007, Nanjing, P.R. China, Part III, LNCS 4493, 2007, pp. 1-8.
[26] V.K. Venkatesan, I. Izonin, J. Periyasamy, A. Indirajithu, A. Batyuk, M.T. Ramakrishna,
Incorporation of energy efficient computational strategies for clustering and routing in
heterogeneous networks of smart city, Energies 15.20 (2022): 7524.
[27] Y. Andrashko et al., A method for assessing the productivity trends of collective scientific
subjects based on the modified PageRank algorithm, Eastern-European Journal of Enterprise
Technologies, 1.4 (121) (2023): 41–47.
[28] V. Kotsovsky, A. Batyuk, V. Voityshyn, On the size of weights for bithreshold neurons and
networks, in: Proceedings of IEEE 16th International Conference on Computer Sciences and
Information Technologies, CSIT 2021, Lviv, Ukrain, 2021, volume 1, pp. 13–16.
[29] E. B. Baum, On the capabilities of multilayer perceptrons, Journal of Complexity 4.3 (1988):
193–215.
[30] V. Kotsovsky, “Hybrid 4-layer bithreshold neural network for multiclass classification,” in</p>
      <p>CEUR Workshop Proceedings, volume 3387, 2023, pp. 212–223.
[31] P. Cortez, A.L. Cerdeira, F. Almeida, T. Matos, J. Reis., Modeling wine preferences by data
mining from physicochemical properties. Decision Support Systems 47 (2009): 547–553.
[32] M. Kelly, R. Longjohn, K. Nottingham, The UCI machine learning repository, 2023. URL:
http://archive.ics.uci.edu.
[33] OpenML: A worldwide machine learning lab, 2024. URL: https://openml.org.
[34] D. C. Angus, Modeling wine quality from physicochemical properties, (2019). URL:
https://api.semanticscholar.org/CorpusID:209521363.
[35] Scikit-learn: Machine Learning in Python, (2024). URL: https://scikit-learn.org.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.H.</given-names>
            <surname>Houssein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.E.</given-names>
            <surname>Hosney</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.M. Emam</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          <string-name>
            <surname>Younis</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          <string-name>
            <surname>Ali</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          <string-name>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <article-title>Soft computing techniques for biomedical data analysis: open issues and challenges</article-title>
          ,
          <source>Artificial Intelligence Review</source>
          <volume>56</volume>
          (
          <year>2023</year>
          ):
          <fpage>2599</fpage>
          -
          <lpage>2649</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Géron</surname>
          </string-name>
          ,
          <article-title>Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems</article-title>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly Media</surname>
          </string-name>
          , Sebastopol, CA,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.K.</given-names>
            <surname>Venkatesan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.T.</given-names>
            <surname>Ramakrishna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Batyuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Barna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Havrysh</surname>
          </string-name>
          ,
          <article-title>High-Performance artificial intelligence recommendation of quality research papers using effective collaborative approach</article-title>
          ,
          <source>Systems 11.2</source>
          (
          <year>2023</year>
          ):
          <fpage>81</fpage>
          . doi:
          <volume>10</volume>
          .3390/systems11020081.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>I. Izonin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tkachenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Mitoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Faramarzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Tsmots</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mashtalir</surname>
          </string-name>
          ,
          <article-title>Machine learning for predicting energy efficiency of buildings: a small data approach</article-title>
          , in: Procedia Computer Science, volume
          <volume>231</volume>
          ,
          <year>2024</year>
          , pp.
          <fpage>72</fpage>
          -
          <lpage>77</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.procs.
          <year>2023</year>
          .
          <volume>12</volume>
          .173.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kotsovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Batyuk</surname>
          </string-name>
          ,
          <article-title>Representational capabilities and learning of bithreshold neural networks</article-title>
          , in: S. Babichev et al. (Eds),
          <source>Advances in Intelligent Systems and Computing</source>
          , volume
          <volume>1246</volume>
          , Springer, Cham,
          <year>2021</year>
          , pp.
          <fpage>499</fpage>
          -
          <lpage>514</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tkachenko</surname>
          </string-name>
          ,
          <article-title>An integral software solution of the SGTM neural-like structures implementation for solving different Data Mining tasks</article-title>
          , in: S.
          <string-name>
            <surname>Babichev</surname>
          </string-name>
          , V. Lytvynenko (Eds.),
          <source>Lecture Notes on Data Engineering and Communications Technologies</source>
          , volume
          <volume>77</volume>
          , Springer, Cham,
          <year>2022</year>
          , pp.
          <fpage>696</fpage>
          -
          <lpage>713</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Havryliuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hovdysh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tolstyak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chopyak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kustra</surname>
          </string-name>
          ,
          <article-title>Investigation of PNN optimization methods to improve classification performance in transplantation medicine</article-title>
          ,
          <source>in: CEUR Workshop Proceedings</source>
          , volume
          <volume>3609</volume>
          ,
          <year>2023</year>
          , pp.
          <fpage>338</fpage>
          -
          <lpage>345</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mitsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sharkan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vargha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lupei</surname>
          </string-name>
          ,
          <article-title>Analyzing Ukrainian media texts by means of support vector machines: aspects of language and copyright</article-title>
          , in: Z. Hu.,
          <string-name>
            <surname>I. Dychka</surname>
          </string-name>
          , M. He (Eds.),
          <article-title>Advances in Computer Science for Engineering and Education VI</article-title>
          .
          <source>ICCSEEA 2023, Lecture Notes on Data Engineering and Communications Technologies</source>
          , volume
          <volume>181</volume>
          , Springer, Cham,
          <year>2023</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Haykin</surname>
          </string-name>
          ,
          <source>Neural Networks and Learning Machines</source>
          , 3rd ed, Pearson Education, Upper Saddle River, NJ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>O.</given-names>
            <surname>Kuchanskyi</surname>
          </string-name>
          et al.,
          <article-title>Gender-related differences in the citation impact of scientific publications and improving the authors' productivity</article-title>
          ,
          <source>Publications</source>
          <volume>11</volume>
          .3 (
          <year>2023</year>
          ):
          <fpage>37</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Anthony</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ratsaby</surname>
          </string-name>
          ,
          <article-title>Large-width machine learning algorithm</article-title>
          ,
          <source>Progress in Artificial Intelligence</source>
          <volume>9</volume>
          .3 (
          <year>2020</year>
          ):
          <fpage>275</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Kotsovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Geche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Batyuk</surname>
          </string-name>
          ,
          <article-title>Artificial complex neurons with half-plane-like and angle-like activation function</article-title>
          ,
          <source>in: Proceedings of the International Conference on Computer Sciences and Information Technologies</source>
          ,
          <string-name>
            <surname>CSIT</surname>
          </string-name>
          <year>2015</year>
          , Lviv, Ukraine,
          <year>2015</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Amirgaliyev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kuchanskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Andrashko</surname>
          </string-name>
          ,
          <article-title>Building a dynamic model of profit maximization for a carsharing system accounting for the region's geographical and economic features</article-title>
          ,
          <source>Eastern-European Journal of Enterprise Technologies</source>
          ,
          <volume>2</volume>
          .
          <fpage>4</fpage>
          -
          <lpage>116</lpage>
          (
          <year>2022</year>
          ):
          <fpage>22</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shlahta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mitsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Horoshko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tsybko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gorbachuk</surname>
          </string-name>
          ,
          <article-title>Development of an interactive map within the implementation of actual state and public directions</article-title>
          ,
          <source>in: Proceedings of the 12th International Conference on Advanced Computer Information Technologies, ACIT</source>
          <year>2022</year>
          , Ruzomberok, Slovakia,
          <year>2022</year>
          , pp.
          <fpage>384</fpage>
          -
          <lpage>387</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rajput</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sreenivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Papailiopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karbasi</surname>
          </string-name>
          ,
          <article-title>An exponential improvement on the memorization capacity of deep threshold networks</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          , volume
          <volume>16</volume>
          ,
          <year>2021</year>
          , pp.
          <fpage>12674</fpage>
          -
          <lpage>12685</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>