<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AppAlicpaptliicoantioonf oRf eRlealtaitvivee DDeerriivvaattioionnTeTremrsmbsy by PoPlyolnyonmomi aiall NNeeuurraallNNetewtowrokrsks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ladislav Zjavka Ladislav Zjavka</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>of Ostrava, IT4innovations O</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>98</fpage>
      <lpage>108</lpage>
      <abstract>
        <p>A lot of problems involve unknown data relations, which can define a derivative based model of dependent variables generalization. Standard softcomputing methods (as artificial neural networks or fuzzy rules) apply usual absolute interval values of input variables. The new proposed differential polynomial neural network makes use of relative data, which can better describe the character regarding a wider range of input values. It constructs and resolves an unknown partial differential equation, using fractional polynomial sum derivative terms of relative data changes. This method might be applied to solve problems concerned a visual pattern generalization or complex system modeling.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
m – number of variables
      </p>
      <p>A(a1, a2, ... , am), ... - vectors of parameters X(x1, x2, ... , xm) - input vector</p>
      <p>
        D-PNN resulted from the GMDH polynomial neural network (Fig.1.), which
was created by a Ukrainian scientist Aleksey Ivakhnenko in 1968 [3]. When the
backpropagation technique was not known yet a technique called Group Method of Data
Handling (GMDH) was developed for neural network structure design and
parameters of polynomials adjustment. General connection between input and output
variables is expressed by the Volterra functional series, a discrete analogue of which
is Kolmogorov-Gabor polynomial (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ). This polynomial can approximate any
stationary random sequence of observations and can be computed by either adaptive
methods or system of Gaussian normal equations.
      </p>
      <p>
        y’ = a0 + a1xi + a2xj + a3xixj + a4xi2 + a5xj2
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>
        GMDH decomposes the complexity of a process into many simpler relationships
each described by low order polynomials (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) for every pair of the input values.
Typical GMDH network maps a vector input x to a scalar output y', which is an
estimate of the true function f(x) = y. Each neuron of the polynomial network fits its
output to the desired value y for each input vector x from the training set. It defines an
optimal structure of complex system model with identifying non-linear relations
between input and output variables [5].
2
      </p>
      <p>
        Differential polynomial neural network
The basic idea of the D-PNN is to create and replace a partial differential equation
(DE) (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ), which is not known in advance and is able to describe a system of
dependent variables, with a sum of fractional multi-parametric polynomial derivative
terms (
        <xref ref-type="bibr" rid="ref4">4</xref>
        )[2].
      </p>
      <p>
        a + ∑i=n1 bi ∂∂xui + ∑i=n1 ∑jn=1 cij ∂x∂i2∂ux j + ... = 0 u = ∑k∞=1 uk
u(x1, x2,, … , xn) - searched function of all input variables (dependent variable)
a, B(b1, b2,, ..., bn), C(c11, c12, ,... ) - parameters
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
      </p>
      <p>The applied method of integral analogues replaces math operators and symbols of
a DE by ratio of corresponding variables. Derivatives are replaced by the integral
analogues, i.e. derivative and all operators are replaced by analogous or proportion
marks in equations [4].</p>
      <p>u i =
(a0 + a1 x1 + a 2 x2 + a3 x1 x2 + a 4 x12 + a5 x22 + ...)m n =</p>
      <p>
        b0 + b1 x1 + ... ∂x1∂x2 ...∂xm
n – combination degree of a complete polynomial of n-variables
m – combination degree of denominator
∂ m f ( x1 ,..., xn )
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
      </p>
      <p>
        The numerator of a term (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) is a polynomial of all n-input variables of a single
neuron and partly defines an unknown function u of eq. (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )(
        <xref ref-type="bibr" rid="ref4">4</xref>
        ). The denominator is a
derivative part of a DE term (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ), which arose from the partial derivation of the
complete n-variable polynomial by competent variable(s). The root function of
numerator takes the polynomial into competent combination degree but needn’t be
used at all if not necessary.
      </p>
      <p>
        A block of the D-PNN (Fig.1.) consists of basic neurons, one for each fractional
polynomial (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ), defining a sum partial derivative term of the DE (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) solution. Blocks
of higher layers are additionally extended with compound neurons of composite
functions, which apply previous layer block outputs and inputs. Each block contains a
single output polynomial (without derivative part), thus the block skeleton of the
DPNN is formed by the GMDH network. Neurons don’t affect the block output but are
applied directly in the sum of a total output calculation of a PDE composition (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ).
Each block has 1 and neuron 2 vectors of adjustable parameters a, resp. a, b.
      </p>
      <p>Combination
degree n=3
x1
x2
x3</p>
      <p>
        Input
variables
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
Π
/
      </p>
      <p>/
Polynomial
Block output</p>
      <p>Neurons
Compound
neurons</p>
      <p>
        Root mean square (RMS) error method (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) was applied for polynomial parameter
optimization and PDE term selection.
      </p>
      <p>M
∑ (y d − yi )
i=1</p>
      <p>2</p>
      <p>M
E =
→ min
3</p>
    </sec>
    <sec id="sec-2">
      <title>Multi-layered backward D-PNN</title>
      <p>
        Multi-layered D-PNN forms composite polynomial functions (Fig.2.). Compound
DE terms, i.e. derivatives in respect to variables of previous layers, are calculated
according to the composite function partial derivation rules (
        <xref ref-type="bibr" rid="ref7">7</xref>
        )(
        <xref ref-type="bibr" rid="ref8">8</xref>
        ). They are formed
by products of partial derivatives of external and internal functions.
      </p>
      <p>F(x1, x2, … , xn) = f(y1, y2, … , ym) = f(φ1(X), φ2(X),..., φm(X))
i = 1, … , m
∂F
∂xk</p>
      <p>m ∂f ( y1, y2 ,..., ym ) ⋅ ∂φ i ( X )
= ∑
i=1 ∂yi ∂xk
k=1, … , n</p>
      <p>
        Each block of the D-PNN involves basic neurons e.g. (
        <xref ref-type="bibr" rid="ref9">9</xref>
        ), at first of only linear
regression. Additionally blocks of the 2nd and following hidden layers are also
extended with neurons, which form composite derivatives utilizing outputs and inputs
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
(
        <xref ref-type="bibr" rid="ref8">8</xref>
        )
of back connected previous layer blocks, e.g. the 1st block of the last (3rd) hidden
layer (10)(11) [7].
1
y1 = ∂f ( x∂2x12,1x22 ) = w1 (a0 + a1 x221 ⋅+(ba02 +x2b21+x2a1)3 x21 x22 ) 2
1
y2 = ∂f (x∂2x11,1x22 ) = w2 (a0 + a1x21 + a22⋅xx2222 + a3 x21x22 ) 2 ⋅
      </p>
      <p>The best-fit neuron selection is the initial phase of the DE composition and may
apply a proper genetic algorithm (GA). Parameters of polynomials might be adjusted
by means of difference evolution algorithm (EA), supplied with sufficient random
mutations [1]. The parameter optimization is performed simultaneously with the GA
term combination search, where may arise a quantity of local and global error
solutions. There would be welcome to apply an adequate gradient descent method
too, which parameter updates result from partial derivatives of polynomial DE terms
in respect with the single parameters [6]. The number of network hidden layers
coincides with a total amount of input variables.</p>
      <p>k
∑ yi
Y = i=1
k
k = amount of active DE terms
(11)</p>
      <p>Only some of all potential combination DE terms (neurons) may participate in the
DE composition, in despite of they have an adjustable term weight (wi). D-PNN’s
total output Y is the sum of all active neuron outputs, divided by their amount k (11).</p>
    </sec>
    <sec id="sec-3">
      <title>4 Identification of data relations</title>
      <p>Consider first only a linear simplification of data relations, thus only linear
polynomials of neurons and blocks might be applied. D-PNN consisting of only 1
block of 2 neurons, all terms of the DE (12), is able to identify simple linear
2variable dependence (function), e.g. x1 = 2x2.</p>
      <p>y = w1
(a0 + a1 x1 + a2 x2 + a3 x1 x2 ) 12
b0 + b1 x1
+ w2
(a0 + a1 x1 + a2 x2 + a3 x1 x2 ) 12</p>
      <p>More complicated dependence, where 2 variables depend on a 3rd (e.g. x1 + x2 =
x3) may be resolved again D-PNN with one 3-variable combination block. The
complete DE (of 1 and 2-combination derivatives) consists of 6 sum terms (neurons)
but only 3 may be employed, derivative terms for x3 (13), x1x3 (14), x2x3 (15). Some
neurons must be inactivated, having an undesirable effect on the network correct
operation. The applied 2-variable combination block D-PNN has 3 hidden layers
(Fig.2.).</p>
      <p>y1 = w1 (a0 + a1x1 + a2x2 + a3bx03 ++ba14xx31x2 + ...+ a7x1x2x3 )13
y2 = w2
y3 = w3
(a0 + a1x1 + a2x2 + a3x3 + a4x1x2 + ... + a7x1x2x3 )23
b0 + b1x1 + b2x3 + b3x1x3</p>
      <p>2
( a0 + a1x1 + a2x2 + a3x3 + a4x1x2 + ...+ a7x1x2x3) 3</p>
      <p>D-PNN can indicate the learned dependence of 3 variables (function) by the
output value 1.0 (or any desired). It was trained with only 6 data samples (Tab.1),
which were selected to involve proportionally the whole training data interval values
&lt;0,500&gt;. However the output function values x3=x1+x2 (x-axis) of the test random
input vectors can exceed the maximal trained sum value 500, while the response is
(12)
(13)
(14)
(15)
kept (Fig.3.). Output errors can result from very disproportional random vector
values, which D-PNN was not trained to, e. g. 360 = 358 + 2.</p>
      <p>The identification of data relations might be applied to a generalization of
fragmented visual patterns into some characteristic dependent elements, which shape
assume moved or sized form in the input matrix and where ANN applications fail [7].
The outcomes of 1-block and multi-layered D-PNN are comparable, however the 2nd
type is able to involve far larger amount of DE terms and so form a more accurately
description of a model.</p>
    </sec>
    <sec id="sec-4">
      <title>5 Function approximations</title>
      <p>D-PNN can approximate a multi-parametric function, analogously to the ANN
approach. Consider the sum function again yt = x1 + x2 + x3, however it could be any
linear function. The network with 3 input variables, forming 1 output y = f(x1, x2, x3)
should approximate the true function yt by means of sum derivative terms of the
partial DE solution. The training data was necessary to be doubled into 12 samples
(Tab.2). The D-PNN and ANN approximation is co-equal on the trained interval
values &lt;6, 520&gt;, however the ANN approximation ability rapidly falls outside of this
range (Fig.4.). The type and operating principle of the D-PNN is the same with
applied the dependence identification (Fig.2.), though requiring more time-consuming
adjustment.</p>
      <p>⎛ ∂u ∂u ∂2u ∂2u ∂2u ⎞</p>
      <p>F⎜⎜⎝ x, y,u, ∂x , ∂y , ∂x2 , ∂x∂y , ∂y 2 ⎠⎟⎟ = 0
where F(x, y, u, p, q, r, s, t) is a function of 8 variables
(16)</p>
      <p>
        In the case of a real-data application D-PNN processes 2-combination square
polynomials of blocks and neurons (DE terms), the same as applied by the GMDH
algorithm (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). This simple polynomial type proves to yield best results besides an
easy use and improves also the linear function approximation (which is notable).
Thus each block includes 5 basic neurons of derivatives x1, x2, x1x2, x12, x22 of the 2nd
order partial DE (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) of an unknown 2-variable function u, which might be transferred
into form of eq. (16). Without this extension only a linear regression of the training
data set would be applied. The square and combination derivative terms are also
calculated according to the composite function derivation rules (17)(18). However
they don’t apply the complete sum of the formulas but only 2 simple terms with 1st
order external function derivatives e.g. (19).
      </p>
      <p>F(x, y) = f (u, v) = f [ϕ (x, y),ψ (x, y)]</p>
      <p>∂∂v2 2f ⎛⎝⎜ ∂∂ψx ⎞⎟⎠ 2 + ∂∂uf ⋅ ∂∂x2ϕ2 + ∂∂fv ⋅ ∂∂x2ψ2
y4 = ∂2 f (∂xx21211, x22) = w4 (a0 + a1x21 + a2x22 +3a⋅3xx2212x22 + a4x221 + a5x222)12 ⋅
(b)
(c)
(d)</p>
      <p>
        The 3-variable D-PNN applying extended polynomials (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) of blocks, neurons and
square, combination DE terms, tried to predict the static pressure of 1 site locality
time-series (Fig.5a-d). It was trained along with the 1-layer ANN the previous day
hourly pressure data series (24 or 48 hours, i.e. data samples), which are free on-line
available [9]. Meteorological forecasts require as a rule high amount of state input
variables to define a complex model, however some tendencies of the progress curves
are notable. The D-PNN 2 models double the applied amount of neurons compared
with D-PNN 1 (Fig.5.). The DE composition based predictions of the D-PNN seems
to succeed any better. The more varied models are formed than ANN (applying 4 or 5
input variables), which is induced by a different neuron combination selection.
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>D-PNN is a new neural network type, which identification and function
approximation is based on generalization of data relations. The relative data
processing is contrary to common soft-computing method approachs (e.g. ANN),
which applications are subjected to a fixed interval of absolute values. This handicap
disallows to use various learning and testing data range values (Fig.4.), which may
involve real data applications. Thus D-PNN’s non-linear regression can cover a
generalization of wider interval values. It forms and resolves a DE, composed of sum
fractional derivative terms, defining a system model of dependent variables. It is
trained only with a small set of input-output data samples, likewise the GMDH
algorithm does [1]. The inaccuracies of presented experiments can result from applied
incomplete rough training and selective methods, requiring large improvements.
Behavior of the presented method differs essentially from other common neural
network techniques.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgement</title>
      <p>This work has been elaborated in the framework of the IT4Innovations Centre of
Excellence project, reg. no. CZ.1.05/1.1.00/02.0070 supported by Operational
Programme ’Research and Development for Innovations’ funded by Structural Funds
of the European Union and by the Ministry of Industry and Trade of the Czech
Republic, under the grant no. FR-TI1/420 , and by SGS, VŠB – Technical University
of Ostrava, Czech Republic, under the grant No. SP2012/58.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abraham</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Particle swarm optimization and Differential evolution algorithms</article-title>
          .
          <source>Studies in Computational Intelligence (SCI) 116</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          ,
          <year>2008</year>
          . Springer-Verlag Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Iba</surname>
          </string-name>
          , H.:
          <article-title>Inference of differential equation models by genetic programming</article-title>
          .
          <source>Information Sciences</source>
          , Volume
          <volume>178</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>23</given-names>
          </string-name>
          <source>, 1 December</source>
          <year>2008</year>
          , Pages
          <fpage>4453</fpage>
          -
          <lpage>4468</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ivakhnenko</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          :
          <article-title>Polynomial theory of complex systems</article-title>
          .
          <source>IEEE Transactions on systems</source>
          ,
          <source>Vol. SMC-1, No.4</source>
          .
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kuneš</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vavroch</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Franta</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Essentials of modeling</article-title>
          .
          <source>SNTL Praha</source>
          <year>1989</year>
          (in Czech).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Nikolaev</surname>
            ,
            <given-names>N.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iba</surname>
          </string-name>
          , H.:
          <article-title>Adaptive Learning of Polynomial Networks</article-title>
          . Springer, New York
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Nikolaev</surname>
            ,
            <given-names>N. Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iba</surname>
          </string-name>
          , H.:
          <article-title>Polynomial harmonic GMDH learning networks for time series modelling</article-title>
          .
          <source>Neural Networks</source>
          <volume>16</volume>
          (
          <year>2003</year>
          ),
          <fpage>1527</fpage>
          -
          <lpage>1540</lpage>
          . Science Direct.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Zjavka</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Generalization of patterns by identification with polynomial neural network</article-title>
          .
          <source>Journal of Electrical</source>
          Engineering Vol.
          <volume>61</volume>
          , No. 2/2010, p.
          <fpage>120</fpage>
          -
          <lpage>124</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zjavka</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Recognition of Generalized Patterns by a Differential Polynomial Neural Network</article-title>
          . Engineering, Technology &amp; Applied Science Research Vol.
          <volume>2</volume>
          , No 1 (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>National</given-names>
            <surname>Climatic</surname>
          </string-name>
          <article-title>Data Center of National Oceanic and Atmospheric Administration (NOAA</article-title>
          ) http://cdo.ncdc.noaa.gov/cdo/3505dat.txt
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>