<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Conformal sets in neural network regression?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Radim Demut</string-name>
          <email>demut@seznam.cz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Holena</string-name>
          <email>martin@cs.cas.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computer Science Academy of Sciences of the</institution>
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>and call Z the example space. Thus the in nite data sequence (??) is an element of the measurable space Z</institution>
        </aff>
      </contrib-group>
      <fpage>17</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>This paper is concerned with predictive regions these predictors are not suitable for neural network in regression models, especially neural networks. We use regression, therefore, we also introduce inductive conthe concept of conformal prediction (CP) to construct re- formal predictors where the prediction rule is updated gions which satisfy given con dence level. Conformal pre- only after a given number of new examples has arrived diction outputs regions, which are automatically valid, but and a calibration set is used. their width and therefore usefulness depends on the used In order to de ne a conformal predictor we need tneolnlcuosnfhoormwitdyi mereeansutrae. gAivennonecxoanmfoprlmeiitsy mwietahsurreespsehcotutldo a suitable nonconformity measure. A nonconformity other examples. We de ne nonconformity measures based measure should tell us how di erent a given example on some reliability estimates such as variance of a bagged is with respect to other examples. In chapter 3, we inmodel or local modeling of prediction error. We also present troduce two reliability estimates: variance of a bagged results of testing CP based on di erent nonconformity mea- model and local modeling of prediction error. We use sures showing their usefulness and comparing them to tra- these reliability estimates in chapter 4 to de ne norditional con dence intervals. malized nonconformity measures. Some other reliability estimates could be used, e.g. sensitivity analysis or density based reliability estimate. 1 Introduction In chapter 5, we use CP, based on nonconformity measures de ned in chapter 4, on testing data to compare our conformal regions with traditional con dence intervals and with conformal intervals where these traditional con dence intervals are used to construct the nonconformity measure.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        This paper is concerned with predictive regions for
regression models, especially neural networks. We often
want to know not only the label y of a new object, but
also how accurate the prediction is. Could the real
label be very far from our prediction or is our prediction
very accurate? It is possible to use traditional con - 2 Conformal prediction
dence intervals to answer this question but they do not
work very well with highly nonlinear regression models We assume that we have an in nite sequence of pairs
such as neural networks. We use conformal prediction
to solve this problem and construct some accurate and (x1; y1); (x2; y2); : : : ; (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
useful prediction regions.
      </p>
      <p>
        We introduce conformal prediction (CP) in chap- called examples. Each example (xi; yi) consists of an
ter 2. Conformal prediction does not output single la- object xi and its label yi. The objects are elements of
bel but a set of labels ". The size of the prediction a measurable space X called the object space and the
set depends on a signi cance level " which we want labels are elements of a measurable space Y called the
to achieve. Signi cance level is under some conditions label space. Moreover, we assume that X is non-empty
the probability that our prediction lies outside the set. and that the -algebra on Y is di erent from f;; Yg.
The set is smaller for larger ". If we have some predic- We denote zi := (xi; yi) and we set
tion rule, we will call it simple predictor and we can
use it to construct conformal predictor. We introduce Z := X Y (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
transductive conformal predictors where the
prediction rule is updated after a new example arrives. But
Z1. Usually we need only slightly weaker assumption Formally, a con dence predictor is a measurable
that the in nite data sequence (??) is drawn from a function
distribution P on Z1 that is exchangeable, that means : Z X (0; 1) ! 2Y (7)
that every n 2 IN, every permutation of f1; : : : ; ng,
and every measurable set E Z1 ful ll
that satis es (??) for all n 2 IN, all incomplete data
sequences x1; y1; : : : ; xn 1; yn 1; xn and all signi cance
P f(z1; z2; : : :) 2 Z1 : (z1; : : : ; zn) 2 Eg = levels "1 "2.
      </p>
      <p>Whether makes an error on the nth trial of the</p>
      <p>
        P f(z1; z2; : : :) 2 Z1 : (z (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ); : : : ; z (n)) 2 Eg data sequence ! = (x1; y1; x2; y2; : : :) at signi cance
We denote Z the set of all nite sequences of ele- level " can be represented by a number that is one in
ments of Z, Zn the set of all sequences of elements of Z case of an error and zero in case of no error
sohfoluenldgtnhotn.mTahkee aonrdyerdiinerwehnicceh. Ionldoredxearmtpolefsorampapleizaer &lt;&gt;&gt;8 1 if yn 2= "(x1; y1; : : : ;
this point we need the concept of a bag. A bag of size err"n( ; !) := xn 1; yn 1; xn) ; (8)
n 2 IN is a collection of n elements some of which may &gt;&gt;: 0 otherwise ;
be identical. To identify a bag we must say what
elements it contains and how many times each of these and the number of errors during the rst n trials is
elements is repeated. We write nz1; : : : ; zn= for the bag n
consisting of elements z1; : : : ; zn, some of which may Err"n( ; !) := X erri"( ; !) : (9)
be identical with each other. We write Z(n) for the i=1
set of all bags of size n of elements of a measurable
space Z. We write Z( ) for the set of all bags of
elements of Z.
(10)
(11)
(12)
      </p>
      <p>If ! is drawn from an exchangeable probability
distribution P , the number err"n( ; !) is the realized
value of a random variable, which we may designate
err"n( ; P ). We say that con dence predictor is
conservatively valid if for any exchangeable probability
distribution P on Z1 there exist two families
( n(") : " 2 (0; 1); n = 1; 2; : : :)
( n(") : " 2 (0; 1); n = 1; 2; : : :)
of f0; 1g-valued variables such that
{ for a xed "; 1("); 2("); : : : is a sequence of
independent Bernoulli random variables with parameter
";
{ for all n and ", n(") n(");
{ the joint distribution of err"n( ; P ), " 2 (0; 1), n =
1; 2; : : :, coincides with the joint distribution of
n("), " 2 (0; 1), n = 1; 2; : : :.
2.2</p>
      <p>Transductive conformal predictors
A nonconformity measure is a measurable mapping</p>
      <p>A : Z( )</p>
      <p>Z ! IR :
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Con dence predictors</title>
      <p>We assume that at the nth trial we have rstly only
the object xn and only later we get the label yn. If we
want to predict yn, we need a simple predictor</p>
      <p>D : Z</p>
      <p>X ! Y :</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) and
For any sequence of old examples x1; y1; : : : ; xn 1,
yn 1 2 Z and any new object xn, it gives
D(x1; y1; : : : ; xn 1; yn 1; xn) 2 Y as its prediction for
the new label yn.
      </p>
      <p>Instead of merely choosing a single element of Y
as our prediction for yn, we want to give subsets of Y
large enough that we can be con dent that yn will fall
in them, while also giving smaller subsets in which we
are less con dent. An algorithm that predicts in this
sense requires additional input " 2 (0; 1), which we
call signi cance level, the complementary value 1 "
is called con dence level. Given all these inputs
x1; y1; : : : ; xn 1; yn 1; xn; "</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
an algorithm
      </p>
      <p>
        that interests us outputs a subset
of Y. We require this subset to shrink as " is increased
that means it holds
"1 (x1; y1; : : : ; xn 1; yn 1; xn)
"2 (x1; y1; : : : ; xn 1; yn 1; xn)
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) To each possible bag of old examples and each possible
new example, A assigns a numerical score indicating
how di erent the new example is from the old ones.
      </p>
      <p>It is sometimes convenient to consider separately how
a nonconformity measure deals with bags of di erent
sizes. If A is a nonconformity measure, for each n =
(6) 1; 2; : : : we de ne a function
whenever "1
"2.</p>
      <p>An : Z(n 1)
for each example zi in the bag. Because a nonconfor- i = An(n(x1; y1); : : : ; (xi 1; yi 1);
mity measure (An) may be scaled however we like, the (xi+1; yi+1); : : : ; (xn; yn)=; (xi; yi)) (22)
numerical value of i does not, by itself, tell us how
unusual (An) nds zi to be. For that we de ne p-value are de ned according to (??) and (??) by the formula
for zi as
yi := Dnz1;:::;zn=(xi)
b
yb(i) := Dnz1;:::;zi 1;zi+1;:::;zn=(xi) :
(19)
n(y)gj &gt; " ;
p := jfj = 1; : : : ; n : j
n
igj :
(15)
and the formula
i :=
(yi; Dnz1;:::;zn=(xi))
from the true label yi. We can also use the deleted is equal to the set of all labels y 2 Y such that
prediction de ned as</p>
      <p>We de ne transductive conformal predictor (TCP)
by a nonconformity measure (An) as a con dence
predictor obtained by setting
respectively. It can be easily checked that in both
cases (An) form a nonconformity measure.
equal to the set of all labels y 2 Y such that
2.3 Inductive conformal predictors
n(y)gj &gt; " ;
jfi = 1; : : : ; n : i(y)
n
(17) In TCP, we need to compute the p-value (??) for all
labels y 2 Y to determine the set ". In the case of
where regression, we have Y = IR and it is not possible to
try each y 2 Y. Sometimes it is possible to generally
i(y) := An(n(x1; y1); : : : ; (xi 1; yi 1); solve equations i(y) n(y) with respect to y, and
therefore determine the set ". But if we use neural
(xi+1; yi+1); : : : ; (xn 1; yn 1); (xn; y)=; networks as simple predictor, we do not know the
gen(xi; yi)) ; 8i = 1; : : : ; n 1 ; eral form of the simple predictor, i.e. we do not know
n(y) := An(n(x1; y1); : : : ; (xn 1; yn 1)=; (xn; y)) : a functional relationship between the training set and
the trained network, because random in uences
en</p>
      <p>We now remind an important property of TCP. ter the training algorithm. Hence, we cannot solve the
The proof of the following theorem can be found in [?]. equations i(y) n(y), and it is not possible to use
tTivheelyorveamlid.1. All conformal predictors are conserva- vTeCryP.coEmvepnutiafttiohneaellqyuianteioncsiecnatn. be solved, it can be
To avoid this problem we can use inductive
confor</p>
      <p>If we are given a simple predictor (??) whose out- mal predictor (ICP). To de ne ICP from a
nonconforput does not depend on the order in which the old mity measure (An) we x a nite or in nite increasing
examples are presented, than the simple predictor D sequence of positive integers m1; m2; : : : (called update
de nes a prediction rule Dnz1;:::;zn= : X ! Y by the trials). If the sequence is nite we add one more
memformula ber equal to in nity at the end of the sequence. We
need more than m1 training examples. Then we nd
Dnz1;:::;zn=(x) := D(z1; : : : ; zn; x) : (18) k such that mk &lt; n mk+1. The ICP is determined
by (An) and the sequence m1; m2; : : : of update trials
A natural measure of nonconformity of zi is the devi- is de ned to be the con dence predictor such that
ation of the predicted label the prediction set
(21)
(23)
(24)
(25)
where the nonconformity scores are de ned by
3.2</p>
      <p>Local modeling of prediction error
j := Amk+1(n(x1; y1); : : : ; (xmk ; ymk )=; (xj ; yj )) ; We nd k nearest neighbors of an unlabeled
example x in the training set, therefore, we have a set
for j = mk + 1; : : : ; n 1 (27) N = f(x1; y1); : : : ; (xk; yk)g of nearest neighbors. We
n := Amk+1(n(x1; y1); : : : ; (xmk ; ymk )=; de ne the estimate denoted CNK for an unlabeled
ex(xn; y)) : (28) ample x as the di erence between the average label of
the nearest neighbors and the example's prediction y</p>
      <p>The proof of the following theorem can be found (using the model that was generated on all learning
in [?]. examples)</p>
      <p>Pk
CNK(x) := i=1 yi y : (33)</p>
      <p>k
The dependence on x on the right hand side of the
previous equation is implicit, but both the prediction y
and the selection of nearest neighbors depends on x.
4</p>
      <sec id="sec-2-1">
        <title>Normalized nonconformity measures</title>
        <p>We will follow a similar approach as is used in the
article [?], but we will incorporate the reliability
estimates from previous chapter and use it for neural
network regression.</p>
        <p>We will use ICP with only one update trial. Let us
have training set of size l, where l &gt; m1. We will split
it into two sets, the proper training set T of size m1
(we will further write m) and the calibration set C of
size q = l m. We will use the proper training set
for creating the simple predictor Dn(x1;y1);:::;(xm;ym)=.
The calibration set is used for calculating the p-value
of new test examples. It is good to rst normalize the
data (i.e. subtract the mean and divide data by sample
variance).</p>
        <p>We will denote ri any of the previously de ned
reliability estimates in the point xi with given simple
predictor D. We compute ri for all points in the
calibration set and de ne Ri for any given point xi as
Ri := ri : (34)</p>
        <p>medianfrj : rj 2 Cg
We de ne a discrepancy measure (??) as
Theorem 2. All ICPs are conservatively valid.</p>
        <p>For ICP combining (??) with (??) and (??) we get
Al+1( n(x1; y1); : : : ; (xl; yl)=; (x; y))
=</p>
        <p>(y; Dn(x1;y1);:::;(xl;yl);(x;y)=(x)) (29)
and</p>
        <p>Al+1( n(x1; y1); : : : ; (xl; yl)=; (x; y))
=
(y; Dn(x1;y1);:::;(xl;yl)=(x)) ;
(30)
respectively. When we de ne A by (??), we can see
that the ICP requires recomputing the prediction rule
only at the update trials m1; m2; : : :. We will use the
simplest case, where there is only one update trial m1,
therefore, we compute the prediction rule only once.
3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Reliability estimates</title>
        <p>In this chapter we are interested in di erent
approaches to estimate the reliability of individual
predictions in regression.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Variance of a bagged model</title>
      <p>We are given a learning set L = f(x1; y1); : : : ; (xn; yn)g
and take repeated bootstrap samples L(i), i = 1; : : : ; m
of size d from the learning set, i.e. for i = 1; : : : ; m
we randomly choose d points from the original
learning set L with the return and put them in L(i). The
number of points d can be chosen arbitrary. We
induce a new model on each of these bootstrap
samples L(i). Each of the models yields a prediction Ki(x),
i = 1; : : : ; m for a considered input x. The label of the
example x is predicted by averaging the individual
predictions</p>
      <p>Pm
K(x) := i=1 Ki(x) : (31)
m
We call this procedure bootstrap aggregating or
bagging. The reliability estimate of a bagged model is
dened as the prediction variance
1 m</p>
      <p>X(Ki(x)
m
BAGV(x) :=
K(x))2 :
(32) and denote
where parameter 0 controls the sensitivity to
changes of Ri. Then, we get the nonconformity score</p>
      <p>We sort nonconformity scores of the calibration
examples in descending order
(y1; y2) :=
y1</p>
      <p>y2
+ Ri</p>
      <p>;
i(y) =
y</p>
      <p>y^i
+ Ri</p>
      <p>:
(m+1)
: : :</p>
      <p>(m+q) ;
s = b"(q + 1)c :
(35)
(36)
(37)
Proposition 1. The prediction set " of the new test
example xl+g (where xl+g is from the in nite
sequence (??)) given the nonconformity score (??) is
equal to the interval
The value of this function # in the point (x1; x2; x3,
x4; x5) can be expressed as
#(x1; x2; x3; x4; x5) =</p>
      <p>A(x1; x2)</p>
      <p>We repeated the following procedure ve times for
region with signi cance level 0:1 and ve times for
region with signi cance level 0:05.
hy^l+g
where
Proof. To compute the prediction set " of the new
test example xl+g we need to nd all y 2 Y such that
for the p-value it holds
p(y) =
jfi = m + 1; : : : ; m + q; l + g : i</p>
      <p>q + 1
&gt; " :
l+g(y)gj</p>
      <p>(40)
l+g(y)gj &gt;
b"(q + 1)c (41)
We multiply the inequality by q + 1 and then it is
equivalent to</p>
      <p>jfi = m + 1; : : : ; m + q; l + g : i
and this inequality holds if and only if
(m+s)</p>
      <p>l+g(y) =
From (??) follows the assertion of the proposition.
5</p>
      <sec id="sec-3-1">
        <title>Simulation</title>
        <p>We carried out a simulation to test the normalized
nonconformity measures based on di erent reliability
estimates. We used neural networks with radial
basis functions (RBF networks) as our regression models
with Gaussian used as the basis function. Therefore,
the output of the RBF network f : IRn ! IR has the
form</p>
        <p>N
f (x) = X i exp
i=1
ijjx
cijj2
;</p>
        <p>(43)
where N is the number of neurons in the hidden layer,
ci is the center vector for neuron i, i determines the
width of the ith neuron and i are the weights of the
linear output neuron. RBF networks are universal
approximators on a compact subset of IRn. This means
that a RBF network with enough hidden neurons can
approximate any continuous function with arbitrary
precision.</p>
        <p>We used a benchmark function similar to some
empirical functions encountered in chemistry to carry out
our experiment. This function was introduced in [?].</p>
        <p>A(x1; x2) = 0:6g(x1</p>
        <p>0:35; x2</p>
        <p>B(x2; x3) = 0:4g(x2
C(x3; x4; x5) = 5 + 25[1
+0:75g(x1
Moreover, the input vectors must satisfy following
conditions</p>
        <p>5
X xi = 1
i=1
and</p>
        <p>xi 2 [0; 1]; for i = 1; : : : ; 5 : (45)
{ Randomly generate 600 points satisfying the
con</p>
        <p>ditions (??).
{ Compute the function values of function # in these</p>
        <p>points.
{ Normalize data (i.e. subtract the mean and divide</p>
        <p>data by sample variance)
{ Split this set of points into a training set of</p>
        <p>500 points and a testing set of 100 points.
{ Split the training set into a proper training set of
401 points and a calibration set of 99 points (then,
we divide the p-value in (??) by 100).
{ Split the proper training set on training set for
tting the RBF network and the validation set. Fit
the RBF network with 1; 2; 3; 4 and 5 hidden
neurons ten times using the Matlab function
lsqcurve</p>
        <p>t.
{ Choose the RBF network with the smallest error
on the validation set for each number of hidden
neurons.
{ Compute the prediction sets for each of the
100 testing points for each number of hidden
neurons.
{ Transform data and predictive regions back to the
original size (i.e. multiply by the original sample
variance and add the original mean)
{ Determine if the original point lies in our
prediction sets.
The initial values of parameters i were set as mean of predictive regions is always slightly higher than the
the response vector, initial values of i were set as the con dence level.
mean of the standard deviation of the components of
training data points. The centers ci were set randomly.</p>
        <p>Results for predictive regions based on the local
modeling of prediction errors depend a little bit on the</p>
        <p>We also computed con dence intervals using Mat- count of nearest neighbors. These intervals are valid
lab function nlpredci (denoted Conf Int). The Jaco- for all numbers of neighbors, but the tightest
interbian can be computed exactly, because the form of the vals were achieved for two neighbors. The di erence
RBF network is known and di erentiable. Therefore, between using ve or ten neighbors is not too big but
we supply the function nlpredci with this Jacobian. We lower number of neighbors works better in our model.
also use the width of this interval as another reliability This is probably caused by our data and it seems that
estimate for our normalized nonconformity measure. only a few neighbors are relevant to our prediction.</p>
        <p>We compare normalized nonconformity measures These regions are also the easiest and fastest to
combased on the following reliability estimates: the local pute.
modeling of prediction errors using nearest neighbors The best results among all predictive regions are
(CNK), the variance of a bagged model (BAGV) and achieved by those based on a variance of a bagged
the width of con dence intervals (CONF). model. These regions are the tightest of all tested and</p>
        <p>The variance of a bagged model was computed for they do not vary as much as those based on con dence
number of di erent models m = 10 and the bootstrap intervals. These regions also maintain the validity. The
samples were as big as the original sample. drawback of these regions is that we need to t a lot of</p>
        <p>The CNK estimates were computed for number of additional models which takes a lot of time in the case
neighbors k = 2; 5; 10. of neural network regression. But if time and
compu</p>
        <p>We present the results of testing CP based on dif- tational e ciency is not a problem then this method
ferent nonconformity measures in Figures ??, ?? and produces best regions.
??. There is a boxplot of all labels in Figure ?? to
compare the range of all labels with the width of di erent
predictive regions. Figures ?? and ?? show boxplots of 450
the width of prediction regions for signi cance levels
" = 0:1 and " = 0:05, respectively. It is not only in- 400
teresting whether the intervals are small enough, but 350
they should also be valid. The percentage of labels in- 300
side the predictive regions are in Tables ?? and ?? for 250
signi cance levels " = 0:1 and " = 0:05, respectively. 200</p>
        <p>The results for traditional con dence inter- 150
vals computed by Matlab function nlpredci are not 100
shown in the gures, because these results are very
di erent from the others. The median width for these 50
intervals lies between 1010 and 1014 for all counts of 1
neurons. This is probably because of the highly
nonlinear character of neural nets, while nlpredci is based Fig. 1. Boxplot of all labels.
on linearization. Moreover, during the computation of
these intervals a Jacobian matrix must be inverted but
this matrix was very often ill conditioned, therefore,
the results for con dence intervals are not too reliable.</p>
        <p>Despite what was said in the previous paragraph, Neurons CNK2 CNK5 CNK10 BAGV CONF
the predictive regions based on the width of con dence 2 91.0 91.2 90.0 91.4 92.4
intervals produce sensible results. But these prediction 3 92.6 92.4 92.6 94.0 93.6
regions show highest inconsistency between di erent 54 9924..26 9920..64 9900..00 9900..42 9900..28
neuron counts and have highest number of very large 6 92.8 89.8 91.8 91.6 91.8
intervals. These regions produce sometimes very good
results, but they are probably very dependent on the Table 1. Percentage of labels inside predictive regions for
actual t of the neural network and their results are " = 0:1.
not as consistent as the results of the other methods.</p>
        <p>However, we can see in Tables ?? and ?? that these
intervals are valid as the percentage of labels inside</p>
        <p>CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5 CNK10 BAGV
Fig. 2. Interval widths for " = 0:1.</p>
        <p>CONF
Neurons: 2</p>
        <p>CNK10
Neurons: 3</p>
        <p>CNK10
Neurons: 4</p>
        <p>CNK10
Neurons: 5
Neurons: 2</p>
        <p>CNK10
Neurons: 3</p>
        <p>CNK10
Neurons: 4</p>
        <p>CNK10</p>
        <p>Neurons: 5
100</p>
        <p>0
CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5</p>
        <p>BAGV</p>
        <p>CONF
CNK2</p>
        <p>CNK5 CNK10 BAGV
Fig. 3. Interval widths for " = 0:05.</p>
        <p>CONF</p>
        <sec id="sec-3-1-1">
          <title>Neurons</title>
          <p>2
3
4
5
6
CNK10
96.8
96.8
96.6
96.4
96.4</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>BAGV</title>
          <p>96.2
97.8
97.4
96.4
97.4
6</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusion</title>
        <p>We presented several methods for computing
predictive regions in neural network regressions. These
methods are based on the inductive conformal prediction
with novel nonconformity measures proposed in this
paper. Those measures use reliability estimates to
determine how di erent a given example is with respect
to other examples. We compared our new predictive
regions with traditional con dence intervals on
testing data. The con dence intervals did not perform
very well, the intervals were too large, it was probably
caused by the high nonlinearity of radial basis
neural networks. Predictive regions which used the width
of con dence intervals as the nonconformity measure
gave much better results. But those results were not as
consistent as the results of the other methods.
Predictive regions based on the local modeling of prediction
errors gave us good results and the computation of the
regions was very fast. A smaller number of neighbors
gave better results for these regions. The best results
were achieved by the regions based on the variance of
a bagged model. The only drawback of this method is
that a lot of models must be tted and it is, therefore,
computationally very ine cient.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bosnic</surname>
          </string-name>
          , Kononenko, I.:
          <article-title>Comparison of approaches for estimating reliability of individual regression predictions</article-title>
          .
          <source>Data &amp; Knowledge Engineering</source>
          ,
          <year>2008</year>
          ,
          <volume>504</volume>
          {
          <fpage>516</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A.</given-names>
            <surname>Gammerman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Shafer</surname>
          </string-name>
          ,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>Vovk: Algorithmic learning in a random world</article-title>
          .
          <source>Springer Science+Business Media</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. E. Uusipaikka:
          <article-title>Con dence intervals in generalized regression models</article-title>
          .
          <source>Chapman &amp; Hall</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>H.</given-names>
            <surname>Papadopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vovk</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Gammerman: Regression conformal prediction with nearest neighbours</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          <volume>40</volume>
          ,
          <year>2011</year>
          ,
          <volume>815</volume>
          {
          <fpage>840</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Valero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Argente</surname>
          </string-name>
          , et al.:
          <article-title>DoE framework for catalyst development based on soft computing techniques</article-title>
          .
          <source>Computers and Chemical Engineering</source>
          <volume>33</volume>
          (
          <issue>1</issue>
          ),
          <year>2009</year>
          ,
          <volume>225</volume>
          {
          <fpage>238</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>