<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Surrogate solutions of Fredholm equations ⋆ by feedforward networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vˇera K˚urkov´a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computer Science, Academy of Sciences of the Czech Republic</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
      </contrib-group>
      <fpage>49</fpage>
      <lpage>54</lpage>
      <abstract>
        <p>Surrogate solutions of Fredholm integral equa- be compared with various surrogate models. One can tions by feedforward neural networks are investigated theo- investigate mathematical properties of these functions retically. Convergence of surrogate solutions computable by as well as properties of their surrogate models aimnetworks with increasing numbers of computational units to ing to estimate speed of convergence of approximatheoretically optimal solutions is proven and upper bounds tions computable by surrogate models with increasing eaoxnavmararptieeletssyooofffcpocenorcmveepprgturetoanntciseonaanarlde uGdneairutisvs,seidta.hneTyrhaeadrieraelsiuullnlutissttsrh.aotledd fboyr pmliocdaetledcofmorpmleuxlaitsy. Mtoaftuhnecmtiaotnicsadletshceroirbyedofbayppthroexciomma-tion of functions by neural networks offers some tools for derivation of such estimates.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <sec id="sec-1-1">
        <title>A large class of functions described by mathemati</title>
        <p>
          cal formulas, numerical calculations of which are
diffiSurrogate modeling is one of successful applications of cult, is formed by solutions of Fredholm integral
equaneural networks. Often it has been used for empirical tions. These equations play an important role in many
functions, i.e., functions for which no mathematical problems in applied science and engineering. They
formulas are known and thus their values can only be arise in image restoration, differential problems with
gained experimentally. When such experimental eval- auxiliary boundary conditions, potential theory and
uations are too expensive or time consuming, it can elasticity, etc. (see, e.g., [
          <xref ref-type="bibr" rid="ref22 ref23">23, 22, 24</xref>
          ]). Mathematical
debe useful to perform them merely for some samples scriptions of solutions of Fredholm equations following
of points of the domains of the empirical functions from classical Fredholm theorem [27, p.499] involve
and the obtained values use as training data for neu- complicated expressions in terms of infinite
Liouvilleral networks. The networks trained on such data can Neumann series with coefficients in the forms of
inteplay roles of surrogate models of these empirical func- grals. Thus numerical calculations of these expressions
tions. For example, input-output functions of feedfor- are time consuming.
ward networks have been used in chemistry as
surrogate models of empirical functions assigning to compo- Recently, several authors [
          <xref ref-type="bibr" rid="ref13 ref6">13, 6</xref>
          ] explored
experisitions of chemicals measures of quality of catalyzers mentally possibilities of surrogate modeling of
soluproduced by reactions of these chemicals, in biology tions of Fredholm equations by perceptron and kernel
as models of empirical functions classifying structures networks. Motivated by these experimental studies,
of RNA, and in economy as models of functions as- in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] we initiated a theoretical analysis of
approxisigning credit ratings to companies [
          <xref ref-type="bibr" rid="ref2 ref7">7, 2</xref>
          ]. However, it mation of solutions of Fredholm equtions by neural
should be emphasized that results obtained by surro- networks. In [
          <xref ref-type="bibr" rid="ref12 ref20 ref9">9, 12, 20</xref>
          ], estimates of rates of
approxgate modeling of empirical functions can only be used imation were derived for surrogate modeling by
netas suggestions to be confirmed by additional exper- works with kernel units induced by the same kernels
iments as no other than empirical knowledge of the as the kernels defining the equations and extended to
functions is available. Moreover, no methodology for certain smooth kernels.
choice of suitable network architectures, type of com- In this paper, we investigate surrogate solutions
putational units and their number has been developed. of Fredholm integral equations by networks with
gen
        </p>
        <p>In contrast to the case of empirical func- eral computational units. Taking advantage of results
tions, analytically described functions, which are sub- from nonlinear approximation theory and suitable
injects of surrogate modeling due to their complicated tegral representations of functions in the form of
“inand time consuming numerical calculations, provide finite” networks, we estimate how well surrogate
soa potential for theoretical analysis of quality of their lutions computable by feedforward networks can
apsurrogate models. Available analytic expressions can proximate exact solutions of Fredholm equations. We
⋆ This work was partially supported by MSˇMT grant apply general results to perceptron and Gaussian
ra</p>
        <p>COST INTELLI OC10047 and RVO 67985807. dial networks.</p>
        <p>The paper is organized as follows. In section 2, we This number can be interpreted as a measure of model
describe surrogate modeling of functions by feedfor- complexity of the network. In contrast to linear
apward neural networks and in section 3, we introduce proximation, the dictionary G has no fixed ordering.
Fredholm integral equations and theoretical approach Often, dictionaries are parameterized families of
to their solutions. In section 4, we recall some results functions modeling computational units, i.e., they are
from nonlinear approximation theory and apply them of the form
to approximation of solutions of Fredholm equations
by feedorward networks. We illustrate our results by GF (X, Y ) := {F (·, y) : X → R | y ∈ Y } , (3)
an example of approximation of Fredholm equations
with the Gaussian kernel by networks with
perceptrons and Gaussian radial units.
where F:X × Y →R is a function of two variables, an
input vector x ∈ X ⊆ Rd and a parameter y ∈ Y ⊆ Rs.</p>
        <p>When X = Y , we write briefly GF (X ). So
one-hiddenlayer networks with n units from a dictionary GF(X,Y )
compute functions from the set
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Surrogate modeling by neural networks</title>
      <p>A traditional approach to surrogate modeling of
functions has employed linear methods such as
polynomial interpolation. For suitable points x1, . . . , xm from
a domain X ⊂ Rd, empirically or numerically obtained
approximations φ¯(x1), . . . , φ¯(xm) of values φ(x1), . . . ,
φ(xm) of a function φ are interpolated by functions
from n-dimensional function spaces. These spaces are
obtained as linear spans
span{g1, . . . , gn} :=
( n</p>
      <p>X wigi | wi ∈ R
i=1
)
,
(1)
where the functions g1, . . . gn are first n elements from
a set G = {gn | n ∈ N+} with a fixed linear ordering
(we use the standard notation := meaning a
definition). Typical examples of linear approximators are
algebraic or trigonometric polynomials. They are
obtained by linear combinations of powers of increasing
degrees or trigonometric functions with increasing
frequencies, resp.</p>
      <p>
        Feedforward neural networks have more adjustable
parameters than linear models as in addition to
coefficients of linear combinations of basis functions, also
inner coefficients of computational units are optimized
during learning. Thus they are sometimes called
variable-basis approximation schemas in contrast to
traditional linear approximators which are called fixed
basis approximation schemas. In some cases, especially
in approximation of functions of large numbers of
variables, it was proven that neural networks achieve
better approximation rates than linear models with much
smaller model complexity [
        <xref ref-type="bibr" rid="ref10 ref11">11, 10</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>One-hidden-layer networks with one linear output</title>
        <p>unit compute input-output functions from sets of the
form
spann G :=
( n</p>
        <p>X wigi | wi ∈ R, gi ∈ G
i=1
)</p>
        <sec id="sec-2-1-1">
          <title>In some contexts, F is called a kernel. However, the</title>
          <p>above-described computational scheme includes fairly
general computational models, such as functions
computable by perceptrons, radial or kernel units, Hermite
functions, trigonometric polynomials, and splines. For
example, with</p>
          <p>F (x, y) = F (x, (v, b)) := σ(hv, xi + b)
and σ : R → R a sigmoidal function, the
computational scheme (2) describes one-hidden-layer
perceptron networks. Radial (RBF) units with an activation
function β : R → R are modeled by the kernel</p>
          <p>F (x, y) = F (x, (v, b)) := β(vkx − bk).</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Typical choice of β is the Gaussian function. Kernel</title>
          <p>units used in support vector machine (SVM) have the
form F (x, y) where F : X × X → R is a symmetric
positive semidefinite function [27].</p>
          <p>Various learning algorithms optimize parameters
y1, . . . , yn of computational units as well as coefficients
w1, . . . , wn of their linear combinations so that
network input-output functions
n
X wi F (., yi)
i=1
from the set spann GF (X, Y ) fit well to training
samples {(xi, φ¯(xi) |i = 1, . . . , m}.
3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Solutions of Fredholm integral equations</title>
      <sec id="sec-3-1">
        <title>Solving an inhomogeneous Fredholm integral equation</title>
        <p>
          where the set G is sometimes called a dictionary [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] of the second kind on a domain X ⊆ Rd for a given
and n is the number of hidden computational units. λ ∈ R \ {0}, K : X × X → R, and f : X → R is
a task of finding a function φ : X → R such that for
all x ∈ X
        </p>
        <sec id="sec-3-1-1">
          <title>The following proposition gives conditions guar</title>
          <p>anteeing compactness of operators TK in spaces
(C(X ), k.ksup), where X ⊆ Rd, of bounded
continuφ(x) − λ Z φ(y)K(x, y) dy = f (x). (4) ous functions on X with the supremum norm</p>
          <p>X kf ksup = supx∈X |f (x)| and to spaces (L2(X ), k.kL2)
The function φ is called solution, f data, K kernel, of square integrable functions with the norm kf kL2 =
and λ parameter of the equation (4). RX f (x)2 dx 1/2. The proof is well-known and easy to</p>
          <p>Fredholm equations can be described in terms of check (see, e.g., [26, p. 112]).
theory of inverse problems. Formally, an inverse
problem is defined by a linear operator A : X → Y between
two function spaces. It is a task of finding for f ∈ Y
(called data) some φ ∈ X (called solution) such that
Proposition 1. (i) If X ⊂ Rd is compact and K :
X × X → R is continuous, then TK : (C(X ), k.ksup) →
(C(X ), k.ksup) is a compact operator.
(ii) If X ⊂ Rd and K ∈ L2(X × X ), then TK :
(L2(X ), k.kL2) → (L2(X ), k.kL2 ) is a compact
operator.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Let TK denotes the integral operator with a kernel</title>
          <p>K : X × X → R defined for every φ in a suitable
function space as</p>
          <p>Z</p>
          <p>X
TK (φ)(x) :=
φ(y) K(x, y) dy</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>So by Corollary 1, when the assumptions of the</title>
          <p>Proposition 1 (i) or (ii) are satisfied and 1/λ is not
an eigenvalue of TK , then for every f in C(X ) or</p>
          <p>
            L2(X ), resp., there exists unique solution φ of the
(5) equation (4). It is known (see, e.g, [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]) that the
solution φ can be expressed as
φ(x) = f (x) − λ
          </p>
          <p>Z</p>
          <p>X
and IX denotes the identity operator. Then the
Fredholm equation (4) can be represented as an inverse
problem defined by the linear operator IX − λ TK . So
it is a problem of finding for a given data f a solution φ
such that
where RKλ : X × X → R is called a resolvent kernel .</p>
          <p>However, the formula expressing the resolvent kernel
(IX − λ TK )(φ) = f. (6) is not suitable for efficient computation as it is
expressed as an infinite Neumann series in powers of λ</p>
          <p>
            The classical Fredholm alternative theorem with coefficients in the form of integrals with iterated
from 1903 proved existence and uniqueness of solu- kernels [5, p.140]. So numerical calculations of values
tions of Fredholm equations for continuous one vari- of solutions of Fredholm equations based on (7) are
able functions on intervals. A modern version hold- quite computationally demanding. Thus various
mething for general Banach spaces is stated in the ods of finding surrogate solutions of (4) have been
exnext theorem from [27, p.499]. Recall that an operator plored [
            <xref ref-type="bibr" rid="ref13 ref6">13, 6</xref>
            ]. Some of these methods utilized
feedforT : (X , k.kX ) → (Y, k.kY ) between two Banach spaces ward networks. Such networks were trained on samples
is called compact if it maps bounded sets to precom- of input-output pairs {(x1, φ¯(x1)), . . . , (xm, φ¯(xm)}
pact sets (i.e., sets whose closures are compact). where {x1, . . . , xm} are selected points from the
domain X and {φ¯(x1), . . . , φ¯(xm)} are numerically
comTheorem 1. Let (X , k.kX ) be a Banach space, T : puted approximations of values {φ(x1), . . . , φ(xm)} of
(X , k.kX ) → (X , k.kX ) be a compact operator, and IX the solution φ. In these experiments, one-hidden-layer
be the identity operator. Then the operator IX + T : networks with perceptrons and Gaussian radial units
(X , k.kX ) → (X , k.kX ) is one-to-one if and only if it were used. However, without a theoretical analysis, it
is onto. is not clear how to choose a proper number n of
network units to guarantee that input-output functions
approximate well the solution and the networks are
not too large to make their implementation
unfeasible.
          </p>
          <p>A straightforward corollary of Theorem 1
guarantees existence and uniqueness of solutions of the
inverse problem (6) when T is a compact operator and
1/λ is not its eigenvalue (i.e., there is no φ ∈ X for
which T (φ) = λφ ).
f (y) RKλ (x, y) dy ,
(7)
Corollary 1. Let (X , k.kX ) be a Banach space,
T : (X , k.kX ) → (X , k.kX ) be a compact operator,
IX be the identity operator, and λ 6= 0 be such that 1/λ
is not an eigenvalue of T . Then the operator IX − λT
is invertible (one-to-one and onto).
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Rates of convergence of surrogate solutions</title>
      <sec id="sec-4-1">
        <title>Estimates of model complexity of one-hidden-layer networks approximating solutions of Fredholm equations follow from inspection of upper bounds on rates</title>
        <p>
          of decrease of errors in approximation of solutions of a network with n units computing functions from the
the equation (4) by sets spannG with n increasing. dictionary G approximates the function h within ε. So
Approximation properties of sets of the form spann G the size of G-variation of the function h to be
approxihave been studied in mathematical theory of neuro- mated is a critical factor influencing model complexity
computing for various types of dictionaries G of networks approximating h within a required
accuand norms measuring approximation errors such as racy. Generally, it is not easy to estimate G-variation.
Hilbert-space norms and the supremum norm (see, However, the following theorem from [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] shows that
e.g., [
          <xref ref-type="bibr" rid="ref4 ref8">4, 8</xref>
          ]). Some such bounds have the form ξ√(hn) , for the special case of functions with integral
represenwhere n is the number of network units and ξ(h) de- tations in the form of “infinite networks”, variational
pends on a certain norm of the function h to be ap- norms are bounded from above by the L1-norms of
proximated. “output-weight” functions of these networks.
        </p>
        <p>This norm is tailored to the dictionary of compu- Theorem 3. Let X ⊆ Rd, Y ⊆ Rs, w ∈ L1(Y ), K :
tation units and can be estimated for functions sat- X ×Y → R be such that GK (X, Y ) = {K(., y) | y ∈ Y }
isfying suitable integral equations. The norm is de- is a bounded subset of (L2(X), k.kL2), and h ∈ L2(X)
fined quite generally for any bounded nonempty sub- be such that for all x ∈ X, h(x) = RY w(y) K(x, y) dy.
set G of a normed linear space (X , k.kX ). It is called Then
G-variation, denoted k.kG, and defined for all f ∈ X khkGK(X,Y ) ≤ kwkL1.
as
kf kG,X := inf {c &gt; 0 | f /c ∈ clX conv (G ∪ −G)} ,</p>
      </sec>
      <sec id="sec-4-2">
        <title>To apply Theorem 2 to approximation of solutions</title>
        <p>where the closure clX is taken with respect to the of Fredholm equations by surrogate models formed by
topology generated by the norm k.kX and conv de- networks with units from a general dictionary G, we
notes the convex hull. So G-variation depends on the need upper bounds on G-variation. The next
proposiambient space norm, but when it is clear from the tion describes a relationship between variations with
context, we write merely kf kG instead of kf kG,X . respect to two sets, G and F ; its proof follows easily</p>
        <p>
          The concept of variational norm was introduced by from the definition of variational norm.
Barron [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] for sets of characteristic functions. Among
them, the set of characteristic functions of half-spaces Proposition 2. Let (X , k.kX ) be a normed linear
forming the dictionary of functions computable by space, F and G its bounded subsets such that cG,F :=
Heaviside perceptrons. Barron’s concept was general- supg∈GkgkF &lt; ∞. Then for all h ∈ X , khkG ≤
ized in [
          <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
          ] to variation with respect to an arbitrary cG,F khkF .
bounded set of functions and applied to various dictio- Combining Theorems 2, 3, and Proposition 2, we
naries of computational units such as Gaussian RBF obtain the next theorem on rates of approximation of
units or kernel units [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. functions which can be expressed as h = TK (w) by
        </p>
        <p>
          The following theorem on rates of approximation networks with units from a dictionary G.
by sets of the form spannG is a reformulation from [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]
of results by Maurey [25], Jones [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], Barron [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] in Theorem 4. Let X ⊆ Rd, K : X × Y → R be
terms of G-variation. For a normed space (X , k.kX ), a bounded kernel, and h ∈ L2(X) such that
g ∈ X and A ⊂ X , we denote by h = TK(w) = RY w(y)K(., y) dy for some w ∈ L1(Y ),
where GK (X, Y ) is a bounded subset of L2(X). Let G
kg − AkX := fi∈nAf kg − f kX be a bounded subset of L2(X) with sG = supg∈GkgkL2
such that cG,K = supg∈GkgkGK(X,Y ) is finite. Then
the distance of g from A. for all n &gt; 0,
Theorem 2. Let (X , k.kX ) be a Hilbert space, G its
bounded nonempty subset, sG = supg∈GkgkX , f ∈ X ,
and n be a positive integer. Then
kh − spannGk2X ≤
s2Gkhk2G − khk2X .
        </p>
        <p>n</p>
        <p>Theorem 2 guarantees that for every ε &gt; 0 and n
satisfying
n ≥
sG khkG
ε
2
kh − spann GkL2 ≤
sG cG,K kwkL1 .</p>
        <p>√n</p>
        <p>A critical factor in the estimate given in
Theorem 4 is the L1-norm of the “output-weight function”
w in the representation of the function h to be
approximated an “infinite network” with units
computing K(., y) in the form
h(x) = TK(w) =</p>
        <p>w(y) K(x, y) dy.</p>
        <p>Z</p>
        <p>The solution φ of the Fredholm equation minus the with the width b by surrogate solutions in the form
function f representing the data, φ − f , is the image of input-output functions of networks with two types
of λ φ mapped by the integral operator TK , i.e., of popular units: sigmoidal perceptrons and Gaussian
radial units. Note that Fredholm equations with
Gausφ − f = TK (λ φ) = λ Z φ(y)K(x, y) dy . sian kernels arise, e.g., in image restoration problems</p>
        <p>X [24]. By μ is denoted the Lebesgue measure on Rd and
Thus to apply Theorem 4 to approximation of a so- by Pdσ(X) the dictionary of functions on X computable
lution of Fredholm equation, we need to estimate the by sigmoidal perceptrons.</p>
        <p>L1-norm of the solution φ itself as λ φ plays the role
of the “output-weight” function in the infinite network
RX λ φ(y) K(x, y) dy.</p>
        <p>Corollary 2. Let X ⊂ Rd be compact, b &gt; 0,
Kb(x, y) = e−bkx−yk2, λ 6= 0 be such that λ1 is not an
eigenvalue of TKb and |λ| &lt; 1. Then the solution φ
Theorem 5. Let X ⊂ Rd be compact, K : X × X → of the equation (4) with f continuous satisfies for all
R be a bounded kernel such that K ∈ L2(X × X), n &gt; 0
ρK := RX supy∈X|K(x, y)|dx be finite, G be a bounded μ(X) |λ| kf kL1
subset of L2(X) with sG = supg∈GkgkL2 such that kφ − f − spann GKb (X)kL2 ≤ (1 − |λ| μ(X) ) √n
cG,K = supg∈GkgkGK(X) is finite, and λ 6= 0 be such
that λ1 is not an eigenvalue of TK and |λ| ρK &lt; 1. and</p>
        <sec id="sec-4-2-1">
          <title>Then the solution φ of the equation (4) satisfies for all n &gt; 0,</title>
          <p>kφ − f − spann Pdσ(X)kL2 ≤ (μ1(−X )|λ2|dμ|(λX|k)f)k√Ln1 .
sG cG,K |λ| kf kL1 .
kφ − f − spann GkL2 ≤ (1 − |λ| ρK ) √n</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Proof. It was shown in [17] that variation of the d</title>
        <p>Proof. As φ − f satisfies the Fredholm equation (4), dimensional Gaussian with respect to the dictionary
we have for every x ∈ X, formed by sigmoidal perceptrons is bounded from
above by 2d and thus by Proposition 2, cPdσ,Kb ≤ 2d.</p>
        <p>|φ(x)| ≤ |λ| kφkL1 supy∈X |K(x, y)| + |f (x)|. The statement then follows by Theorem 5, an estimate
Integrating over X we get sρGKKbb=≤μ(Xμ()X.) and equalities sPdσ = μ(X) an2d
kφkL1 ≤ |λ| ρK kφkL1 + kf kL1
and so kφkL1 (1 − |λ| ρK ) ≤ kf kL1. This inequality is
non trivial only when |λ| &lt; ρ1K . Thus we get kwkL1 =
|λ|kφkL1 ≤ |1λ−| |kλf|kρLK1 . The statement then follows from
Theorem 4. 2</p>
      </sec>
      <sec id="sec-4-4">
        <title>Theorem 5 estimates rates of approximation of the</title>
        <p>function φ − f = λ RX f (y) RKλ (x, y) dy by functions
computable by networks with units from dictionary G
formed by functions with GK -variations bounded
by cG,K . Numerical computations of values of the
function λ RX f (y) RKλ (x, y) dy are time consuming.
For |λ| &lt; ρ1K and any bounded dictionary G with
finite bound cG,K on GK(X)-variations on its elements,
input-output functions of networks with
increasing numbers of units from G converge to the function
φ − f . When for a reasonable size of the network
measured by the number n of units, the upper bound from
Theorem 5 is sufficiently small, the network can serve
as a good surrogate model of the solution of Fredholm
equation.</p>
        <p>To illustrate our results, consider approximation of
Fredholm equations with the Gaussian kernel
Kb(x, y) = e−bkx−yk</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>K.</given-names>
            <surname>Atkinson:</surname>
          </string-name>
          <article-title>The numerical solution of integral equations of the second kind</article-title>
          . Cambridge University Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Baerns</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Holenˇa: combinatorial development of solid catalytic materials</article-title>
          . Imperial College Press, London,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Barron</surname>
          </string-name>
          :
          <article-title>Neural net approximation</article-title>
          . In K. Narendra, (Ed.),
          <source>Proc. 7th Yale Workshop on Adaptive and Learning Systems</source>
          , Yale University Press,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Barron</surname>
          </string-name>
          <article-title>: Universal approximation bounds for superpositions of a sigmoidal function</article-title>
          .
          <source>IEEE Transactions on Information Theory 39</source>
          ,
          <year>1993</year>
          ,
          <fpage>930</fpage>
          -
          <lpage>945</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>R.</given-names>
            <surname>Courant</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Hilbert: Methods of mathematical physic</article-title>
          , volume I. Wiley, New York,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>S.</given-names>
            <surname>Effati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Buzhabadi</surname>
          </string-name>
          :
          <article-title>A neural network approach for solving Fredholm integral equations of the second kind</article-title>
          .
          <source>Neural Computing and Applications</source>
          . doi:
          <volume>10</volume>
          .1007/s00521-010-0489-y.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>A.</given-names>
            <surname>Forrester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sobester</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Keane: Engineering design via surrogate modelling: A practical guide</article-title>
          . Wiley,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>F.</given-names>
            <surname>Girosi</surname>
          </string-name>
          <article-title>: Approximation error bounds that use VCbounds</article-title>
          .
          <source>In Proceedings of ICANN 1995</source>
          , Paris,
          <year>1995</year>
          ,
          <fpage>295</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>G.</given-names>
            <surname>Gnecco</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>K˚urkov´a, M. Sanguineti: Bounds for 24</article-title>
          . Y.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <article-title>Xu: Integral equation models for approximate solutions of Fredholm integral equations image restoration: high accuracy methods and fast alusing kernel networks</article-title>
          .
          <source>In T. Honkela</source>
          and et al., (Eds), gorithms.
          <source>Inverse Problems</source>
          ,
          <volume>26</volume>
          .
          <source>doi: 10.1088/0266- Lecture Notes in Computer Science (Proceedings of 5611/26/4/045006. ICANN</source>
          <year>2011</year>
          ), vol.
          <volume>6791</volume>
          , Springer, Heidelberg,
          <year>2011</year>
          , 25. G. Pisier:
          <article-title>Remarques sur un r´esultat non publi</article-title>
          ´e de 126-
          <fpage>133</fpage>
          . B.
          <string-name>
            <surname>Maurey</surname>
          </string-name>
          . In S´eminaire d'Analyse Fonctionnelle
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>G.</given-names>
            <surname>Gnecco</surname>
          </string-name>
          , V. K˚urkov´a, M.
          <source>Sanguineti: Can</source>
          <year>1980</year>
          -
          <volume>81</volume>
          , vol. I, no. 12, E´cole Polytechnique,
          <article-title>Centre dictionary-based computational models outperform</article-title>
          the de Math´ematiques, Palaiseau, France,
          <year>1981</year>
          . best linear ones?
          <source>Neural Networks</source>
          <volume>24</volume>
          ,
          <year>2011</year>
          ,
          <fpage>881</fpage>
          -
          <lpage>887</lpage>
          . 26. W. Rudin:
          <article-title>Functional analysis</article-title>
          .
          <source>McGraw-Hill</source>
          , Boston,
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>G.</given-names>
            <surname>Gnecco</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>K˚urkov´a, M. Sanguineti: Some compar- 1991. isons of complexity in dictionary-based and linear com- 27. I. Steinwart, A. Christmann: Support vector machines</article-title>
          .
          <source>putational models. Neural Networks</source>
          ,
          <volume>24</volume>
          ,
          <year>2011</year>
          ,
          <fpage>171</fpage>
          - Springer, New York,
          <year>2008</year>
          .
          <volume>182</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>G.</given-names>
            <surname>Gnecco</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>K˚urkov´a, M. Sanguineti: Accuracy of approximations of solutions to Fredholm equations by kernel methods</article-title>
          .
          <source>Applied Mathematics and Computation</source>
          <volume>218</volume>
          ,
          <year>2012</year>
          ,
          <fpage>7481</fpage>
          -
          <lpage>7497</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>A.</given-names>
            <surname>Golbabai</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Seifollahi: Numerical solution of the second kind integral equations using radial basis function networks</article-title>
          .
          <source>Applied Mathematics and Computation</source>
          <volume>174</volume>
          ,
          <year>2006</year>
          ,
          <fpage>877</fpage>
          -
          <lpage>883</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>R.</given-names>
            <surname>Gribonval</surname>
          </string-name>
          , P. Vandergheynst:
          <article-title>On the exponential convergence of matching pursuits in quasi-incoherent dictionaries</article-title>
          .
          <source>IEEE Transactions on Information Theory 52</source>
          ,
          <year>2006</year>
          ,
          <fpage>255</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Jones</surname>
          </string-name>
          :
          <article-title>A simple lemma on greedy approximation in Hilbert space and convergence rates for projection pursuit regression and neural network training</article-title>
          .
          <source>Annals of Statistics</source>
          ,
          <volume>20</volume>
          :
          <fpage>608</fpage>
          -
          <lpage>613</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>P. C. Kainen</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>K˚urkov´a, M. Sanguineti: Complexity of Gaussian radial-basis networks approximating smooth functions</article-title>
          .
          <source>Journal of Complexity 25</source>
          ,
          <year>2009</year>
          ,
          <fpage>63</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>P. C. Kainen</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>K˚urkov´a, A. Vogt: A Sobolev-type upper bound for rates of approximation by linear combinations of Heaviside plane waves</article-title>
          .
          <source>Journal of Approximation Theory</source>
          <volume>147</volume>
          ,
          <year>2007</year>
          ,
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. V.
          <article-title>K˚urkov´a: Dimension-independent rates of approximation by neural networks</article-title>
          . In K. Warwick and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>K´arny</article-title>
          ´, (Eds),
          <source>Computer-Intensive Methods in Control and Signal Processing. The Curse of Dimensionality</source>
          , Birkh¨auser, Boston, MA,
          <year>1997</year>
          ,
          <fpage>261</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. V.
          <article-title>K˚urkov´a: High-dimensional approximation and optimization by neural networks</article-title>
          . In J. Suykens, G. Horv´ath, S. Basu,
          <string-name>
            <given-names>C.</given-names>
            <surname>Micchelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Vandewalle</surname>
          </string-name>
          , (Eds),
          <source>Advances in Learning Theory: Methods, Models and Applications</source>
          ,
          <source>(Chapter 4)</source>
          . IOS Press, Amsterdam,
          <year>2003</year>
          ,
          <fpage>69</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. V.
          <article-title>K˚urkov´a: Accuracy estimates for surrogate solutions of integral equations by neural networks</article-title>
          . In M. Bielikov´a, G. Fridrich, G. Gottlob,
          <string-name>
            <given-names>S.</given-names>
            <surname>Katzenbeisser</surname>
          </string-name>
          , R. Sˇp´anek, G. Tur´an, (Eds),
          <source>SOFSEM 2012: Theory and Practice of Computer Science</source>
          , vol. II,. Institute of Computer Science, Prague,
          <year>2012</year>
          ,
          <fpage>95</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. V.
          <article-title>K˚urkov´a: Complexity estimates based on integral transforms induced by computational units</article-title>
          .
          <source>Neural Networks</source>
          ,
          <volume>30</volume>
          :
          <fpage>160</fpage>
          -
          <lpage>167</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. A. T. Lonseth:
          <article-title>Sources and applications of integral equations</article-title>
          .
          <source>SIAM Review 19</source>
          ,
          <year>1977</year>
          ,
          <fpage>241</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. W. V. Lovitt:
          <article-title>Linear integral equations</article-title>
          .
          <source>Dover</source>
          , New York,
          <year>1950</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>