<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Combining Gaussian Processes and Neural Networks in Surrogate Modeling for Covariance Matrix Adaptation Evolution Strategy</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Koza</string-name>
          <email>kozajan@fit.cvut.cz</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jirˇí Tumpach</string-name>
          <email>tumpach@cs.cas.cz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zbyneˇk Pitra</string-name>
          <email>z.pitra@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Holenˇa</string-name>
          <email>martin@cs.cas.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academy of Sciences, Institute of Computer Science</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Charles University, Faculty of Mathematics and Physics</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Czech Technical University, Faculty of Information Technology</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Czech Technical University, Faculty of Nuclear Sciences and Physical Engineering</institution>
          ,
          <addr-line>Prague</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper focuses on surrogate models for Covariance Matrix Adaptation Evolution Strategy (CMA-ES) in continuous black-box optimization. Surrogate modeling has proven to be able to decrease the number of evaluations of the objective function, which is an important requirement in some real-world applications where the evaluation can be costly or time-demanding. Surrogate models achieve this by providing an approximation instead of the evaluation of the true objective function. One of the stateof-the-art models for this task is the Gaussian process. We present an approach to combining Gaussian processes with artificial neural networks, which was previously successfully applied to other machine learning domains. The experimental part employs data recorded from previous CMA-ES runs, allowing us to assess different settings of surrogate models without running the whole CMA-ES algorithm. The data were collected using 24 noiseless benchmark functions of the platform for comparing continuous optimizers COCO in 5 different dimensions. Overall, we used data samples from over 2.8 million generations of CMA-ES runs. The results examine and statistically compare six covariance functions of Gaussian processes with the neural network extension. So far, the combined model did not show up to outperform the Gaussian process alone. Therefore, in conclusion, we discuss possible reasons for this and ideas for future research.</p>
      </abstract>
      <kwd-group>
        <kwd>black-box optimization</kwd>
        <kwd>surrogate modeling</kwd>
        <kwd>artificial neural networks</kwd>
        <kwd>Gaussian processes</kwd>
        <kwd>covariance functions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In evolutionary black-box optimization, an increasing
attention is paid to tasks with expensive evaluation of the
black-box objective function. It is immaterial whether
that expensiveness is due to time-consuming
computation like in long simulations [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], or due to evaluation
through costly experiments like in some areas of science
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. To deal with such expensive evaluations, black-box
optimization has in the late 1990s and early 2000s adopted
an approach called surrogate modeling or metamodeling
[
        <xref ref-type="bibr" rid="ref13 ref15 ref34 ref42 ref45 ref48 ref7">7, 13, 15, 34, 42, 45, 48</xref>
        ].
      </p>
      <p>
        Basically, a surrogate model in continuous black-box
optimization is any regression model that with a sufficient
fidelity approximates the true black-box objective function
and replaces it in some of its evaluations. Among the
regression models most frequently used to this end, we are
most interested in Gaussian processes (GPs) [
        <xref ref-type="bibr" rid="ref47">47</xref>
        ], due to
the fact that they estimate not only the expected value of
the true objective function, but the whole distribution of its
values. Apart from regression, they are also encountered
in classification, and play a key role in Bayesian
optimization [
        <xref ref-type="bibr" rid="ref22 ref28">22, 28</xref>
        ], where the distribution of values provides
selection criteria that are alternatives to the objective
function value, such as the probability of improvement or the
expected improvement.
      </p>
      <p>
        The importance of GPs in machine learning incited
attempts to integrate them with the leading learning
paradigm of the last decades – neural learning, including deep
learning. The attractiveness of this research direction is
further supported by recent theoretical results concerning
relationships of asymptotic properties of important kinds
of artificial neural networks (ANNs) to properties of GPs
[
        <xref ref-type="bibr" rid="ref35 ref39 ref41">35,39,41</xref>
        ]. The integration of GP with neural learning has
been proposed on two different levels:
(i) Proper integration of an ANN with a GP, in which the
      </p>
      <p>
        GP forms the final output layer of the ANN [
        <xref ref-type="bibr" rid="ref52 ref9">9, 52</xref>
        ].
(ii) Only a transfer of the layered structure, which is a
crucial feature, to the GP context, leading to the
concept of deep GPs (DGPs) [
        <xref ref-type="bibr" rid="ref12 ref22 ref23 ref8">8, 12, 22, 23</xref>
        ].
      </p>
      <p>
        The recalled research into the integration of GPs with
neural learning has used regression data [
        <xref ref-type="bibr" rid="ref12 ref23 ref52 ref8 ref9">8, 9, 12, 23, 52</xref>
        ],
mostly from the UCI Machine Learning Repository [
        <xref ref-type="bibr" rid="ref50">50</xref>
        ],
but also data concerning locomotion of walking bipedal
robots [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and face patches and hadwritten digits [
        <xref ref-type="bibr" rid="ref52">52</xref>
        ].
In [
        <xref ref-type="bibr" rid="ref12 ref23">12,23</xref>
        ], also classification data were used. However, we
are aware of only one application of such an integration, in
particular of DGPs, to two very easy 1- and 2-dimensional
optimization problems [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Hence, there is a gap between
the importance of GPs in Bayesian optimization and
missing investigations of the suitability of integrated GP-ANN
models for surrogate modeling in black-box optimization.
That gap motivated the research reported in this paper.
      </p>
      <p>We have performed this novel kind of investigation of
GP-ANN integration using data from more than 5000 runs
of using GPs as Bayesian optimizers in black-box
optimization, i.e. optimization of pointwise observable
functions for which no analytic expression is available and
the function values have to be obtained either empirically,
e.g. through measuring or experiments, or through
numerical simulations. Based on that data, our investigation
addressed two research questions:
1. Does the integration of a GP with neural learning has
an added value compared with employing a GP alone
or an ANN alone?
2. What is the impact of each of the considered GP
covariance functions on the result of GP-ANN
integration?</p>
      <p>The next section introduces the principles of
surrogate modeling as well as of the evolutionary optimization
method Covariance Matrix Adaptation Evolution Strategy
(CMA-ES), in the context of which our research has been
performed. In Section 3, the fundamentals of GPs are
recalled and the method we have used for their integration
with neural networks is explained. The core part of the
paper is Section 4, which attempts to provide experimental
answers to the above research questions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Surrogate modeling in Black-Box</title>
    </sec>
    <sec id="sec-3">
      <title>Optimization</title>
      <p>
        Basically, the purpose of surrogate modeling – to
approximate an unknown functional dependence – coincides with
the purpose of response surface modeling in the design of
experiments [
        <xref ref-type="bibr" rid="ref24 ref40">24, 40</xref>
        ]. Therefore, it is not surprising that
typical response surface models, i.e., low order
polynomials, belong also to the most traditional and most
successful surrogate models [
        <xref ref-type="bibr" rid="ref1 ref2 ref20 ref29 ref45">1, 2, 20, 29, 45</xref>
        ]. Other frequently
used kinds of surrogate models are artificial neural
networks of the kind multilayer perceptron (MLP) or radial
basis function network [
        <xref ref-type="bibr" rid="ref18 ref25 ref26 ref27 ref42 ref53">18, 25–27, 42, 53</xref>
        ], and Gaussian
processes, in surrogate modeling also known as kriging
[
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref33 ref34 ref37 ref49 ref51 ref6 ref7">6, 7, 13–15, 33, 34, 37, 49, 51</xref>
        ]. Occasinally encountered
are support vector regression [
        <xref ref-type="bibr" rid="ref10 ref36">10, 36</xref>
        ] and random forests
[
        <xref ref-type="bibr" rid="ref44 ref5">5, 44</xref>
        ].
2.1
      </p>
      <sec id="sec-3-1">
        <title>CMA-ES and Its Surrogate-Assisted Variant DTS-CMA-ES</title>
        <p>
          The CMA-ES algorithm performs unconstrained
optimization on Rd , by means of iterative sampling of
populations sized l from a d-dimensional Gaussian
distribution N(m; s 2C), and uses a given parent number m among
the sampled points corresponding to the lowest objective
function values, to update the parameters of that
distribution. Hence, it updates the expected value m, which is used
as the current point estimate of the function optimum, the
matrix C and the step-size s . The CMA-ES is invariant
with respect to strictly increasing transformations of the
objective function. Hence, to make use of the evaluations
of the objective function in a set of points, it needs to know
only the ordering of those evaluations. Details of the
algorithm can be found in [
          <xref ref-type="bibr" rid="ref19 ref21">19, 21</xref>
          ].
        </p>
        <p>
          During the more than 20 years of CMA-ES existence,
a number of surrogate-assisted variants of this algorithm
have been proposed, a survey can be found in [
          <xref ref-type="bibr" rid="ref43 ref6">6, 43</xref>
          ].
Here, we pay attention only to the most recent GP-based
among them, the Doubly Trained Surrogate CMA-ES
(DTS-CMA-ES) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], surrogate-assisted variant of
CMAES. It employs two GPs f1 and f2, trained consecutively,
to find an evaluation of the population x1; : : : ; xl , with f1
used for active learning of training data for f2. Due to
the CMA-ES invariance with respect to strictly
increasing transformations, it evaluates the difference between
predictions only according to the difference in the
ordering of those predictions, using the ranking difference error
(RDE). The RDE of y 2 Rl with respect to y0 2 Rl
considering k best components is defined:
        </p>
        <p>RDE(y; y0) =
k
åi;(r(y0))i k(jr(y0))i (r(y))ij
maxp2P(l ) åik=1 ji p 1(i)j
;
(1)
where P(l ) denotes the set of permutaions of f1; : : : ; l g
and r(y) denotes the ordering of the components of y, i.e.,
(8y 2 Rl ) r(y) 2 P(l ) &amp; (r(y))i &lt; (r(y)) j ) yi y j.
Because the CMA-ES algorithm uses the surrogate model
to select the most promising candidates for true evaluation,
the metric considers only k best samples. The range of
RDE metric is [0; 1], it equals 0 for the exact ordering and
1 for thereverse order.</p>
        <p>The algorithm DTS-CMA-ES is described in
Algorithm 1, using the following notation:
• A for an archive – a set of points that have already
been evaluated by the true black-box objective
function BB;
• ds2C for the Mahalanobis distance given by s 2C:
ds2C(x; x0) =
q
(x
x0)&gt;s 2C 1(x
x0);
(2)
• Nk(x; A) for the set of k of ds2C-nearest neighbours of
x 2 Rd with respect to the archive A;
• fi(x1; : : : ; xl ) = ( fi(x1); : : : ; fi(xl )), for i = 1; 2;
• Th for the training set selection:</p>
        <p>l
Th = [ fx 2 Nh(x j; A)jds2C(x; x j) &lt; rmaxg
j=1
(3)
with rmax &gt; 0 for h = 1; : : : ; jAj;
• k(A) = maxfhjjThj</p>
        <p>Nmaxg, with Nmax 2 N;
• rPoI for decreasing ordering of f1(x1); : : : ; f1(xl )
according to the probability of improvement with
respect to the lowest BB value found so far,
i &lt; j ) PoI(rPoI( f1(x1; : : : ; xl )))i;V )</p>
        <p>PoI(rPoI( f1(x1; : : : ; xl ))) j;V ); (4)
where V = minx2A BB(x).
Algorithm 1 Algorithm DTS-CMA-ES
Require: x1; : : : ; xl 2 Rd , m, A, s and C – step
size and matrix from the CMA-ES
distribution, Nmax 2 N such that Nmax l , rmax &gt; 0,
b ; emin; emax; einitial; amin; amax; ainitial 2 (0; 1)
1: if this is the 1st call of the algorithm in the current
CMA-ES run then
a = ainitial; e = einitial
take over the returned values of a; e from its
previous call in the run
5: end if
6: Train a Gaussian process f1 on Tk(A), estimating mGP,
sn; s f ; ` through maximization of the likelihood (7)
7: Evaluate BB(x j) for x j not yet BB-evaluated and such
that (rPoI( f1(x1; : : : ; xl ))) j dal e
8: Update A to A [ f(x jj(rPoI( f1(x1); : : : ; f1(xl ))) j
dal eg
9: Train a Gaussian process f2 on Tk(A), estimating mGP,
sn, s f , ` through maximization of the likelihood (7)
10: For x j such that (rPoI( f1(x1; : : : ; xl ))) j dal e,
update f2(x j) = BB(x j)
11: Update e to b RDEm ( f1(x1; : : : ; xl ); ( f2(x1; : : : ; xl )) +
(1 b )e and a to amin + max(0; min(1; emeax emeimnin ))
12: For j fulfilling (rPoI( f1(x1; : : : ; xl )) j &gt; dal e, update
f2(x j) to f2(x j) minf f2(x j0)j(rPoI( f1(x1; : : : ; xl)) j0 &gt;
dal eg + minf f2(x j0 )j(rPoI( f1(x1; : : : ; xl )) j0 dal eg
13: Return f2(x1); : : : ; f2(xl ); e; a
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Gaussian Processes and Their</title>
    </sec>
    <sec id="sec-5">
      <title>Integration with Neural Networks</title>
      <p>3.1</p>
      <sec id="sec-5-1">
        <title>Gaussian processes</title>
        <p>
          A Gaussian process on a set X Rd ; d 2 N is a collection
of random variables ( f (x))x2X , any finite number of which
has a joint Gaussian distribution [
          <xref ref-type="bibr" rid="ref47">47</xref>
          ]. It is completely
specified by a mean function m : X ! R, typically assumed
constant, and by a covariance function k : X X ! R such
that for x; x0 2 X ,
        </p>
        <p>E f (x) = m; cov( f (x); f (x0)) = k(x; x0):
(5)
Therefore, a GP is often denoted GP(m; k(x; x0)) or
GP(m; k).</p>
        <p>The value of f (x) is typically accessible only as a noisy
observation y = f (x) + e, where e is a zero-mean
Gausssian noise with a variance sn &gt; 0. Then
cov(y; y0) = k(x; x0) + sn2I(x = x0);
(6)
where I(proposition) equals for a true proposition 1, for a
false proposition 0.</p>
        <p>Consider now the prediction of the random variable
f (x?) in a point x? 2 X if we already know the
observations y1; : : : ; yn in points x1; : : : ; xn. Introduce the
vectors x = (x1; : : : ; xn)&gt;, y = (y1; : : : ; yn)&gt; = ( f (x1) +
e; : : : f (xn) + e)&gt;, k? = (k(x1; x?); : : : ; k(xn; x?))&gt; and the
matrix K 2 Rn n such that (K)i; j = k(xi; x j) + sn2I(i = j).
Then the probability density of the vector y of observations
is
p(y; m; k; sn2) =
exp
12 (y
m)&gt;K 1(y</p>
        <p>m)
p(2p)2 det(K + sn2In)
;
(7)
where det(M) denotes the determinant of a matrix M.
Furthermore, as a consequence of the assumption of Gaussian
joint distribution, also the conditional distribution of f (x?)
conditioned on y is Gaussian, namely</p>
        <p>N(m(x?) + k?K 1(y
m); k(x?; x?) k?&gt;K 1k?):
(8)</p>
        <p>According to (6), the relationship between the
observations y and y0 is determined by the covariance function k.
In the reported research, we have considered 6 kinds of
covariance functions, listed below. In their definitions, the
notation r = kx0 xk is used, and among the parameters
of k, aka hyperparameters of the GP, frequently
encountered are s 2f; ` &gt; 0, called signal variance and
characteristic length scale, respectively. Other parameters are
introduced for each covariance function separately.
(i) Linear: kLIN(x; x0) = s02 + x&gt;x0, with a bias s02.
(ii) Quadratic is the square of the linear covariance:
kQ(x; x0) = (s02 + x&gt;x0)2.
(iii) Rational quadratic: kRQ(x; x0) = s 2f 1 + 2ar2`2
with a &gt; 0.
(iv) Squared exponential: kSE(x; x0) = s 2f exp
(v) Matérn 52 :</p>
        <p>kM52 atérn(x; x0) = s 2f 1 + p`5r + 35`r22 exp p`5r .
(vi) One composite covariance function, namely the sum
of kSE and kQ:
kSE+Q(x; x0) = kSE(x; x0) + kQ(x; x0).
a
3.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>GP as the Output Layer of a Neural Network</title>
        <p>
          The approach integrating a GP into an ANN as its output
layer has been independently proposed in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and [
          <xref ref-type="bibr" rid="ref52">52</xref>
          ]. It
relies on the following two assumptions:
        </p>
        <p>1. If nI denotes the number of the ANN input
neurons, then the ANN computes a mapping net of nI
-dimensional input values into the set X on which is the GP.
Consequently, the number of neurons in the last hidden
layer equals the dimension d, and the ANN maps an
input v into a point x = net(v) 2 X , corresponding to an
observation f (x + e) governed by GP (Figure 1). From
the point of view of the ANN inputs, the GP is now
GP(m(net(v)); k(net(v); net(v0))).</p>
        <p>2. The GP mean m is assumed to be a known constant,
thus not contributing to the GP hyperparameters and
independent of net.</p>
        <p>Due to the assumption 2., the only hyperparameters of
the GP are the parameters q k of the covariance function.
outputGPlayer:y</p>
        <p>(+distributionofy)
lasthiddenlayer:x
.
.</p>
        <p>.</p>
        <p>1sthiddenlayer
inputlayer:inputvaluesv</p>
        <p>I</p>
        <p>I</p>
        <p>
          As to the ANN, it depends on the one hand on the vector
qW of its weights and biases, on the other hand on the
parameters q A of its architecture (such as the number of
hidden layers, the numbers of neurons in each of them, the
activation functions). Altogether, the integrated GP-ANN
model depends on the vector (q k ; qW ; q A), which we in
accordance with [
          <xref ref-type="bibr" rid="ref52 ref9">9, 52</xref>
          ] and with the terminology common
for GPs also call hyperparameters, although in the context
of ANNs, this term is typically used only for q A.
        </p>
        <p>Consider now n inputs to the neural network, v1; : : : ; vn,
mapped to the inputs x1 = net(v1); : : : ; xn = net(vn) of the
GP, corresponding to the observations y = (y1; : : : ; yn)&gt;.
Then the log-likelihood of q is
where m is the constant assumed in Assumption 2., and</p>
        <p>L(q ) = ln p(y; m; k; sn2);
(K)i; j = k(net(vi); net(v j)):
(9)
(10)
To find the combination of values of q with the
maximal likelihood, methods for smooth optimization can be
applied only with respect to q k and qW , with respect to
which is L continuous. With respect to q A, we have two
possibilities. First, to fix q A in advance, thus effectively
restricting the hyperparameters of the integrated model to
(q k ; qW ). Second, to perform the optimization with
respect to q A in an outer loop (using, e.g. grid search
combined with cross-validation) and for each considered
combination of values of q A perform the optimization with
respect to (q k ; qW ) in an inner loop.</p>
        <p>Finally, let the smooth optimization be performed in
the most simple but, in the context of neural networks,
also most frequent way – as gradient descent. The partial
derivatives forming Ñ(qk ;qW )L can be computed as:
¶ L n ¶ L</p>
        <p>i; j=1 ¶¶KLi; j ¶¶Kqi`k;j ; ¶ q`W =
¶ q`k = å</p>
        <p>
          n
= å
i; j;k=1 ¶ Ki; j ¶ xk
¶ L ¶ Ki; j ¶ net(vk) :
¶ q`W
(11)
In (11), the partial derivatives ¶¶KLi; j , i; j = 1; : : : ; n, are
components of the matrix derivative ¶¶KL , for which the
calculations of matrix differential calculus [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ] together with (7)
and (9) yield
¶ L
¶ K
=
As a dataset to compare different configurations of the
combined GP-ANN model, we used recorded
DTS-CMAES runs from previous experiments. This allowed us to
effectively evaluate the surrogate model on its own
without having to perform the whole optimization. The
opensource Matlab implementation of DTS-CMA-ES, used to
obtain the data, is available at [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The underlying
surrogate models were Gaussian process with 8 different
covariance functions implemented using the GPML
Toolbox [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ]. The optimization runs were collected on the
platform for comparing continuous optimizers COCO. The
employed black-box optimization benchmark set contains
24 noiseless functions scalable to any input dimension. We
used dimensions 2, 3, 5, 10, and 20 in 5 different instances,
which consist in random rotation and translation in the
input space. Therefore, for each combination of benchmark
function and dimension, 40 runs are available. More
important than the number of runs, however, is the number of
their generations because data from each generation apart
from the first can be used for testing all those surrogate
models that could be trained with data from the previous
generations. The number of generations in a particular run
of CMA-ES algorithm is unknown before the run and
depends on the objective function and its particular instance.
The benchmarks and their total numbers of available
generations for individual dimensions are listed in Table 1.
We skipped the linear function (benchmark function
number 5), which is easy to optimize, and the recorded runs
did not provide enough data samples to train and evaluate
the surrogate models.
4.2
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Experimental Setup</title>
        <p>We compare the combinations of neural networks with
Gaussian processes using six different covariance
functions listed in Subsection 3.1. All models were trained on
the same set of training data Tk(A) as was used in steps 6
and 9 of Algorithm 1. Due to the condition (3), this set is
rather restricted and allows training only a restricted ANN
to prevent overfitting. Therefore, we decided to use a
multilayer perceptron with a single hidden layer, thus a
topology (nI ; nH ; nO), where nI is the dimension of the training
data, i.e. nI 2 f2; 3; 5; 10; 20g, and
nH = nO =
82 if nI = 2;
&gt;
&lt;</p>
        <p>
          3 if nI = 3; 5;
&gt;:5 if nI = 10; 20:
(13)
As the activation function for both the hidden and
output layer, we chose the logistic sigmoid. We trained the
weights and biases of the neural network together with the
parameters of the Gaussian process as proposed in [
          <xref ref-type="bibr" rid="ref52">52</xref>
          ]
and outlined in Subsection 3.2. As a loss function, we used
the Gaussian log-likelihood and optimized the parameters
with Adam [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] for a maximum of 1000 iterations. We
also kept a 10 % validation set out of the training data to
monitor overfitting, and we selected the model with the
lowest L2 validation error during the training.
4.3
        </p>
      </sec>
      <sec id="sec-5-4">
        <title>Results</title>
        <p>To evaluate different covariance functions, we used the
recorded data in a similar way as it would be used by
DTSCMA-ES algorithm. We took each generation of points
except the first one and we used it as testing data. Every
point evaluated by the true objective function in all
previous generations is then available as a training sample. We
filter these samples using the training set selection method
Tk(A) and trained the surrogate model on it.</p>
        <p>This way, we evaluated six different models on every
generation of samples listed in Table 1. The results are
presented first in a function-method view in Table 4, then
in a dimension-method view in Table 5. The metric used
in both tables is RDE, which shows how precise the model
is in ordering of the predicted values, therefore the lower
is the error value, the better. The first table contains
results for every benchmark function separately, averaged
over every input dimension, function instance and
DTSCMA-ES generation as well as the average over the whole
group of functions. The second table provides a view on
how the dimension of the input space affects the
approximation error. The results for a particular dimension are
averaged over all functions in a specific group, instances, and
generations. We also visualized with boxplots the
summarized RDE values for the function groups in Figure 2.</p>
        <p>
          Moreover, we verified the results by performing
multiple comparison tests. A Non-parametric Friedman test was
conducted on RDEs across all results for particular
functions and function types in Table 4, and for a particular
combination of dimensions and function types in Table 5.
If the null hypothesis of the equality of all six considered
methods was declined, the Wilcoxon signed-rank test was
performed for all pairs of covariance functions, and its
results were corrected for multiple hypotheses testing using
the Holm method. We summarized the results of statistical
testing in Tables 2 and 3.
The results show that the combined model performs best
with the simplest linear kernel. We compared the
GPANN with a linear kernel with pure Gaussian processes in
paper [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. We found out that if we compare the combined
GP-ANN with GP both with linear kernels, the neural
extensions can bring better results in some cases. However,
using more complex covariance functions with GP is still
better. Therefore, we tried to apply it also to the GP-ANN
combination. Unfortunately, the results with other kernels
ended up much worse.
        </p>
        <p>
          In our future research, we would like to try to overcome
this. We want to systematically investigate different ANN
topologies, including the direction of deep Gaussian
processes [
          <xref ref-type="bibr" rid="ref12 ref22 ref23 ref8">8, 12, 22, 23</xref>
          ], in which only the topology is used
from an ANN, but all neurons are replaced by GPs.
Moreover, in addition to the selection of the training set used
in DTS-CMA-ES and described in Algorithm 1, we want
to consider also alternative ways of training set selection,
allowing to train larger networks. Finally, we intend to
perform research into transfer learning of surrogate
models: An ANN-GP model with a deep neural network will
be trained on data from many optimization runs, such as
those employed in this paper, and then the model used in
a new run of the same optimizer will be obtained through
additional fine tuning restricted only to the GP and last 1-2
layers of the ANN.
        </p>
        <p>
          We would also like to try to change the way how the
model is trained. In paper [
          <xref ref-type="bibr" rid="ref52">52</xref>
          ] the parameters of
Gaussian process are learned together with the parameters of
the neural network. We think that this might not work
well in the domain of surrogate modeling in black-box
optimization. Therefore, we will try to train the network and
Gaussian process separately.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we examined an extension of Gaussian
processes used in surrogate modeling. At the beginning,
we described the CMA-ES algorithm for black-box
optimization and its surrogate-assisted variant DTS-CMA-ES.
Then we outlined Gaussian processes, various covariance
functions, and their neural network extension. The
presented research is our first attempt in the application of the
combination of Gaussian processes with neural networks
as a surrogate model in black-box optimization. We
implemented the combined model and compared the results
obtained with six different covariance functions. By looking
at the results presented in the previous section, we can
conclude that the combined model yields the best results with
the simplest linear kernel. The other covariance functions
produce much larger errors and the worst performing one
seems to be the composite kernel kSE+Q. Unfortunatelly
the results showed up to be worse than using Gaussian
processes alone, therefore we discuss possible ideas of further
improvements at the end of the paper.</p>
      <sec id="sec-6-1">
        <title>Acknowledgement</title>
        <p>The research reported in this paper has been supported
by SVV project number 260 575 and was also partially
supported by Czech Science Foundation (GA CˇR) grant
18-18080S. Computational resources were supplied by
the project "e-Infrastruktura CZ" (e-INFRA LM2018140)
provided within the program Projects of Large Research,
Development and Innovations Infrastructures.
n d u</p>
        <p>o
fu e r</p>
        <p>c g
e a r
c p a
n s l
ia t cu
r u i
a p tr
v n a
o i
x io un
is t f</p>
        <p>c t
r n n
o u e
f f r
s e</p>
        <p>k f
eu ra i</p>
        <p>f
l d
a</p>
        <p>m r
v h e
R b d</p>
        <p>e
e f g
g o a
a r
r e e
e p v
v y a
a t</p>
        <p>e
f e r
o h a
n t s
E c v</p>
        <p>n o
D e
p
c e
t h in
n t
e s
r d n
e n o
f
f a it
i
d n c
Moderately Ill-Conditioned
κ LIN
κ Q
κ RQ
κ SE
κ Mat
κ SE+Q
Highly Ill-Conditioned
κ LIN
κ Q
κ RQ
κ SE
κ Mat
κ SE+Q</p>
        <p>Multi-modal with Global Structure</p>
        <p>Multi-modal Weakly Structured</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Auger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brockhoff</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          .
          <article-title>Benchmarking the local metamodel cma-es on the noiseless BBOB'2013 test bed</article-title>
          .
          <source>In GECCO'13</source>
          , pages
          <fpage>1225</fpage>
          -
          <lpage>1232</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Auger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schoenauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Vanhaecke</surname>
          </string-name>
          .
          <article-title>LS-CMAES: A second-order algorithm for covariance matrix adaptation</article-title>
          .
          <source>In Parallel Problem Solving from Nature - PPSN VIII</source>
          , pages
          <fpage>182</fpage>
          -
          <lpage>191</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Baerns</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa. Combinatorial</surname>
          </string-name>
          <article-title>Development of Solid Catalytic Materials</article-title>
          .
          <article-title>Design of High-Throughput Experiments, Data Analysis, Data Mining</article-title>
          . Imperial College Press / World Scientific, London,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bajer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra. Surrogate</surname>
          </string-name>
          CMA-ES. https:// github.com/bajeluk/surrogate-cmaes,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bajer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa. Benchmarking</surname>
          </string-name>
          <article-title>Gaussian processes and random forests surrogate models on the BBOB noiseless testbed</article-title>
          .
          <source>In GECCO'15 Companion</source>
          , pages
          <fpage>1143</fpage>
          -
          <lpage>1150</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bajer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Repický</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa</surname>
          </string-name>
          .
          <article-title>Gaussian process surrogate models for the CMA evolution strategy</article-title>
          .
          <source>Evolutionary Computation</source>
          ,
          <volume>27</volume>
          :
          <fpage>665</fpage>
          -
          <lpage>697</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Booker</surname>
          </string-name>
          , J. Dennis,
          <string-name>
            <given-names>P.D.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.B.</given-names>
            <surname>Serafini</surname>
          </string-name>
          , Torczon V., and
          <string-name>
            <given-names>M.</given-names>
            <surname>Trosset</surname>
          </string-name>
          .
          <article-title>A rigorous framework for optimization by surrogates</article-title>
          .
          <source>Structural and Multidisciplinary Optimization</source>
          ,
          <volume>17</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hernandez-Lobato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hernandez-Lobato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Turner</surname>
          </string-name>
          .
          <article-title>Deep Gaussian processes for regression using approximate expectation propagation</article-title>
          .
          <source>In ICML</source>
          , pages
          <fpage>1472</fpage>
          -
          <lpage>1481</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Calandra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.E.</given-names>
            <surname>Rasmussen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.P.</given-names>
            <surname>Deisenroth</surname>
          </string-name>
          .
          <article-title>Manifold Gaussian processes for regression</article-title>
          .
          <source>In IJCNN</source>
          , pages
          <fpage>3338</fpage>
          -
          <lpage>3345</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.M.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.H.</given-names>
            <surname>Griebisch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.W.</given-names>
            <surname>Simpson</surname>
          </string-name>
          .
          <article-title>Analysis of support vector regression for approximation of complex engineering analyses</article-title>
          .
          <source>Journal of Mechanical Design</source>
          ,
          <volume>127</volume>
          :
          <fpage>1077</fpage>
          -
          <lpage>1087</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>The</surname>
            <given-names>COCO platform</given-names>
          </string-name>
          ,
          <year>2016</year>
          . http://coco.gforge.inria.fr.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cutajar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.V.</given-names>
            <surname>Bonilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Michiardi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Filippone</surname>
          </string-name>
          .
          <article-title>Random feature expansions for deep Gaussian processes</article-title>
          . In ICML, pages
          <fpage>884</fpage>
          -
          <lpage>893</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.A.</given-names>
            <surname>El-Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.B.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Keane</surname>
          </string-name>
          .
          <article-title>Metamodeling techniques for evolutionary optimization of computaitonally expensive problems: Promises and limitations</article-title>
          .
          <source>In Proceedings of the Genetic and Evolutionary Computation Conference</source>
          , pages
          <fpage>196</fpage>
          -
          <lpage>203</lpage>
          . Morgan Kaufmann Publishers,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Emmerich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Giannakoglou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Naujoks</surname>
          </string-name>
          .
          <article-title>Singleand multi-objective evolutionary optimization assisted by Gaussian random field metamodels,</article-title>
          .
          <source>IEEE Transactions on Evolutionary Computation</source>
          ,
          <volume>10</volume>
          :
          <fpage>421</fpage>
          -
          <lpage>439</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Emmerich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Giotis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Özdemir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bäck</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Giannakoglou</surname>
          </string-name>
          .
          <article-title>Metamodel-assisted evolution strategies</article-title>
          .
          <source>In PPSN VII</source>
          , pages
          <fpage>361</fpage>
          -
          <lpage>370</lpage>
          . ACM,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Finck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ros</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Auger</surname>
          </string-name>
          .
          <article-title>Real-parameter black-box optimization benchmarking 2010: Presentation of the noisy functions</article-title>
          .
          <source>Technical report, INRIA</source>
          , Paris Saclay,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Forrester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sobester</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Keane</surname>
          </string-name>
          . Engineering Design via Surrogate Modelling:
          <string-name>
            <given-names>A Practical</given-names>
            <surname>Guide</surname>
          </string-name>
          . John Wiley and Sons, New York,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>H.M.</given-names>
            <surname>Gutmann</surname>
          </string-name>
          .
          <article-title>A radial basis function method for global optimization</article-title>
          .
          <source>Journal of Global Optimization</source>
          ,
          <volume>19</volume>
          :
          <fpage>201</fpage>
          -
          <lpage>227</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          .
          <article-title>The CMA evolution strategy: A comparing review</article-title>
          .
          <source>In Towards a New Evolutionary Computation</source>
          , pages
          <fpage>75</fpage>
          -
          <lpage>102</lpage>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          .
          <article-title>A global surrogate assisted CMA-ES</article-title>
          .
          <source>In GECCO'19</source>
          , pages
          <fpage>664</fpage>
          -
          <lpage>672</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Ostermaier</surname>
          </string-name>
          .
          <article-title>Completely derandomized self-adaptation in evolution strategies</article-title>
          .
          <source>Evolutionary Computation</source>
          ,
          <volume>9</volume>
          :
          <fpage>159</fpage>
          -
          <lpage>195</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Hebbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Brevault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Balesdent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.G.</given-names>
            <surname>Talbi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Melab</surname>
          </string-name>
          .
          <article-title>Efficient global optimization using deep Gaussian processes</article-title>
          .
          <source>In IEEE CEC</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>G.</given-names>
            <surname>Hernández-Muñoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Villacampa-Calvo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. Hernández</given-names>
            <surname>Lobato</surname>
          </string-name>
          .
          <article-title>Deep Gaussian processes using expectation propagation and Monte Carlo methods</article-title>
          .
          <source>In ECML PKDD</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          , paper no.
          <issue>128</issue>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hosder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Watson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Grossman</surname>
          </string-name>
          .
          <article-title>Polynomial response surface approximations for the multidisciplinary design optimization of a high speed civil transport</article-title>
          .
          <source>Optimization and Engineering</source>
          ,
          <volume>2</volume>
          :
          <fpage>431</fpage>
          -
          <lpage>452</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hüsken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Olhofer</surname>
          </string-name>
          , and Sendhoff B.
          <article-title>Neural networks for fitness approximation in evolutionary optimization</article-title>
          . In Y. Jin, editor,
          <source>Knowledge Incorporation in Evolutionary Computation</source>
          , pages
          <fpage>281</fpage>
          -
          <lpage>306</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Olhofer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sendhoff</surname>
          </string-name>
          .
          <article-title>Managing approximate models in evolutionary aerodynamic design optimization</article-title>
          .
          <source>In CEC 2001</source>
          , pages
          <fpage>592</fpage>
          -
          <lpage>599</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Olhofer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Sendhoff</surname>
          </string-name>
          .
          <article-title>A framework for evolutionary optimization with approximate fitness functions</article-title>
          .
          <source>IEEE Transactions on Evolutionary Computation</source>
          ,
          <volume>6</volume>
          :
          <fpage>481</fpage>
          -
          <lpage>494</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>D.R.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schonlau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.J.</given-names>
            <surname>Welch</surname>
          </string-name>
          .
          <article-title>Efficient global optimization of expensive black-box functions</article-title>
          .
          <source>Journal of Global Optimization</source>
          ,
          <volume>13</volume>
          :
          <fpage>455</fpage>
          -
          <lpage>492</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Koumoutsakos</surname>
          </string-name>
          .
          <article-title>Local metamodels for optimization using evolution strategies</article-title>
          .
          <source>In PPSN IX</source>
          , pages
          <fpage>939</fpage>
          -
          <lpage>948</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>D.P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>Preprint arXiv:1412.6980</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Koza</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Tumpach</surname>
          </string-name>
          .
          <article-title>Surrogate networks</article-title>
          . https:// github.com/c0zzy/surrogate-networks,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>J.</given-names>
            <surname>Koza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tumpach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa</surname>
          </string-name>
          .
          <article-title>Using past experience for configuration of Gaussian processes in black-box optimization</article-title>
          .
          <source>In 15th Learning and Intelligent Optimization Conference, page accepted for publication</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.W.</given-names>
            <surname>Kruisselbrink</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.T.M. Emmerich</surname>
            ,
            <given-names>A.H.</given-names>
          </string-name>
          <string-name>
            <surname>Deutz</surname>
            , and
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Bäck</surname>
          </string-name>
          .
          <article-title>A robust optimization approach using kriging metamodels for robustness approximation in the CMA-ES</article-title>
          .
          <source>In IEEE CEC</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>S.J.</given-names>
            <surname>Leary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Keane</surname>
          </string-name>
          .
          <article-title>A derivative based surrogate model for approximating and optimizing the output of an expensive computer simulation</article-title>
          .
          <source>Journal of Global Optimization</source>
          ,
          <volume>30</volume>
          :
          <fpage>39</fpage>
          -
          <lpage>58</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bahri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Novak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.S.</given-names>
            <surname>Schoenholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          , et al.
          <article-title>Deep neural networks as Gaussian processes</article-title>
          .
          <source>In ICLR</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>I.</given-names>
            <surname>Loshchilov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schoenauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Sebag</surname>
          </string-name>
          .
          <article-title>Intensive surrogate model exploitation in self-adaptive surrogateassisted CMA-ES (saACM-ES)</article-title>
          .
          <source>In GECCO'13</source>
          , pages
          <fpage>439</fpage>
          -
          <lpage>446</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          .
          <article-title>An evolution strategy assisted by an ensemble of local Gaussian process models</article-title>
          .
          <source>In GECCO'13</source>
          , pages
          <fpage>447</fpage>
          -
          <lpage>454</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Magnus</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Neudecker. Matrix</surname>
          </string-name>
          <article-title>Differential Calculus with Applications in Statistics and Econometrics</article-title>
          . John Wiley and Sons, Chichester,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>A.G.G.</given-names>
            <surname>Matthews</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rowland</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.E.</given-names>
            <surname>Turner</surname>
          </string-name>
          .
          <article-title>Gaussian process behaviour in wide deep neural networks</article-title>
          .
          <source>In ICLR</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>R.H.</given-names>
            <surname>Myers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.C.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          , and
          <string-name>
            <surname>C.M. Anderson-Cook</surname>
          </string-name>
          .
          <article-title>Response Surface Methodology: Proces and Product Optimization Using Designed Experiments</article-title>
          . John Wiley and Sons, Hoboken,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>R.</given-names>
            <surname>Novak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bahri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yang</surname>
          </string-name>
          , et al.
          <article-title>Bayesian deep convolutional networks with many channels are Gaussian processes</article-title>
          .
          <source>In ICLR</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>Y.S.</given-names>
            <surname>Ong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.B.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Keane</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.W.</given-names>
            <surname>Wong</surname>
          </string-name>
          .
          <article-title>Surrogate-assisted evolutionary optimization frameworks for high-fidelity engineering design problems</article-title>
          . In Y. Jin, editor,
          <source>Knowledge Incorporation in Evolutionary Computation</source>
          , pages
          <fpage>307</fpage>
          -
          <lpage>331</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hanuš</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Koza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tumpach</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa</surname>
          </string-name>
          .
          <article-title>Interaction between model and its evolution control in surrogate-assisted CMA evolution strategy</article-title>
          .
          <source>In GECCO'21, page paper no. 358</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Repický</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Holenˇa</surname>
          </string-name>
          .
          <article-title>Boosted regression forest for the doubly trained surrogate covariance matrix adaptation evolution strategy</article-title>
          .
          <source>In ITAT 2018</source>
          , pages
          <fpage>72</fpage>
          -
          <lpage>79</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>K.</given-names>
            <surname>Rasheed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Vattam</surname>
          </string-name>
          .
          <article-title>Methods for using surrogate modesl to speed up genetic algorithm oprimization: Informed operators and genetic engineering</article-title>
          . In Y. Jin, editor,
          <source>Knowledge Incorporation in Evolutionary Computation</source>
          , pages
          <fpage>103</fpage>
          -
          <lpage>123</lpage>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>C.E</given-names>
            <surname>Rasmussen</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Hansen</surname>
          </string-name>
          .
          <source>GPML 4</source>
          .0. matlab toolbox. http://www.gaussianprocess.org/gpml/code/ matlab/doc/.
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>E.</given-names>
            <surname>Rasmussen</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Williams</surname>
          </string-name>
          .
          <article-title>Gaussian Processes for Machine Learning</article-title>
          . MIT Press, Cambridge,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ratle</surname>
          </string-name>
          .
          <article-title>Kriging as a surrogate fitness landscape in evolutionary optimization</article-title>
          .
          <source>Artificial Intelligence for Engineering Design, Analysis and Manufacturing</source>
          ,
          <volume>15</volume>
          :
          <fpage>37</fpage>
          -
          <lpage>49</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ulmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Streichert</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zell</surname>
          </string-name>
          .
          <article-title>Evolution strategies assisted by Gaussian processes with improved pre-selection criterion</article-title>
          .
          <source>In IEEE CEC</source>
          , pages
          <fpage>692</fpage>
          -
          <lpage>699</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [50] University of California in Irvine.
          <article-title>Repository of machine learning databases</article-title>
          . http://www.ics.uci.edu/ mlearn,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>V.</given-names>
            <surname>Volz</surname>
          </string-name>
          , G. Rudolph, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Naujoks</surname>
          </string-name>
          .
          <article-title>Investigating uncertainty propagation in surrogate-assisted evolutionary algorithms</article-title>
          .
          <source>In GECCO'17</source>
          , pages
          <fpage>881</fpage>
          -
          <lpage>888</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>A.G.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.P.</given-names>
            <surname>Xing</surname>
          </string-name>
          .
          <article-title>Deep kernel learning</article-title>
          .
          <source>In ICAIS</source>
          , pages
          <fpage>370</fpage>
          -
          <lpage>378</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>Z.Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.S.</given-names>
            <surname>Ong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.B.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Keane</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.Y.</given-names>
            <surname>Lum</surname>
          </string-name>
          .
          <article-title>Combining global and local surrogate models to accellerate evolutionary optimization</article-title>
          .
          <source>IEEE Transactions on Systems, Man and Cybernetics</source>
          . Part C:
          <article-title>Applications</article-title>
          and Reviews,
          <volume>37</volume>
          :
          <fpage>66</fpage>
          -
          <lpage>76</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>