<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Nonparametric Bootstrap Estimation for Implicitly Weighted Robust Regression</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Kalina</string-name>
          <email>kalina@cs.cas.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barbora Peštová</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computer Science CAS</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>1885</volume>
      <fpage>78</fpage>
      <lpage>85</lpage>
      <abstract>
        <p>Implicitly weighted robust regression estimators for linear and nonlinear regression models include linear and nonlinear versions of the least trimmed squares and least weighted squares. After recalling known facts about these estimators, a nonparametric bootstrap procedure is proposed in this paper for estimates of their variances. These bootstrap estimates are elaborated for both the linear and nonlinear model. Practical contributions include several examples investigating the performance of the nonlinear least weighted squares estimator and comparing it with the classical least squares also by means of the variance estimates. Another theoretical novelty is a proposal of a twostage version of the nonlinear least weighted squares estimator with adaptive (data-dependent) weights.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Regression methodology has the aim to model (describe,
estimate) a continuous variable (response) depending on
one or more independent variables (regressors), which
may be continuous and/or discrete. Such modelling finds
applications in an enormously wide spectrum of
applications and allows to predict values of an important variable,
which is considered to play the role of the response, for
particular values of regressors.</p>
      <p>Standard estimation methods in various linear and
nonlinear regression models are known to be too
vulnerable (sensitive) to the presence of outlying measurements
(outliers), which occur in real data for various reasons,
e.g. measurement errors, different conditions or violations
of assumptions of the model under consideration.</p>
      <p>Therefore, numerous robust regression methods have
been proposed since the development of robust
statistical estimation in 1960s as diagnostic tools for classical
methods. Some of them can be understood as reliable
self-standing procedures tailor-made to suppress the
effect of data contamination by various kinds of outliers
[20, 6, 2, 13]. In the course of time, the breakdown
point has become one of crucial measures of robustness,
which can be interpreted as a high resistance
(insensitivity) against outlying measurements in the data and one
of crucial measures of robustness of statistical estimators.
The finite-sample breakdown point is defined as the
minimal fraction of data that can drive an estimator beyond all
bounds when set to arbitrary values [13].</p>
      <p>While M-estimators represent the most widely used
robust statistical methods, they have been criticized for their
low breakdown point in linear regression. Thus, other
methods with a high value of the breakdown point
(asymptotically up to 1/2) are desirable, which are commonly
denoted as highly robust.</p>
      <p>This paper is devoted to the question of estimating the
variance of highly robust implicitly weighted estimators
in linear and nonlinear models, which has not been
investigated in literature. After recalling the least weighted
squares estimator in Section 2, a bootstrap estimate of its
variance is described and illustrated in Section 3. The
methodology is further generalized to nonlinear regression
in Section 4, where also a two-stage version of the
nonlinear least weighted squares estimator is proposed and
investigated. Three examples comparing the performance
of standard and robust methods in the nonlinear model are
presented in a separate Section 5 and conclusions are
summarized in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Least Weighted Squares</title>
      <p>This section recalls the least weighted squares estimator,
which is one of promising tools estimating parameters of
the standard linear regression model</p>
      <p>
        Yi = β0 + β1Xi1 + · · · + βpXip + ei, i = 1, . . . , n.
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
Here, Y = (Y1, . . . ,Yn)T denotes a continuous response,
(X1 j, . . . , Xn j)T is the j-th regressor for j = 1, . . . , p and
e1, . . . , en are random errors of the model. The least
squares estimator bLS of β = (β0, . . . , βp)T is well known
to be highly vulnerable to the presence of outlying
measurements in the data [13]. Therefore, numerous robust
regression methods have been proposed as alternatives to
the least squares.
      </p>
      <p>
        The least weighted squares (LWS) estimator of β in the
model (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) represents one of available robust estimators,
which was proposed in [25]. It is based on implicit
weighting of individual observations, down-weighting less
reliable observations, which might be potential outliers with a
high influence on the results. Thus, if suitable weights are
used, it may reach a high breakdown point [12].
      </p>
      <p>
        The weights are assigned to individual observations
after an (implicitly given) permutation, which is determined
only during the computation of the estimator. For a given
estimate b = (b0, b1, . . . , bp)T ∈ IRp of β , let a residual be
defined as
ui(b) = Yi − b0 − b1Xi1 − · · · − bpXip, i = 1, . . . , n. (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
Squared residuals will be considered in ascending order
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Nonparametric Bootstrap for the LWS</title>
      <p>
        u(
        <xref ref-type="bibr" rid="ref21">21</xref>
        )(b) ≤ u(
        <xref ref-type="bibr" rid="ref22">22</xref>
        )(b) ≤ · · · ≤ u(2n)(b).
      </p>
      <p>The LWS estimator of β denoted as bLW S is defined as
n
arg min ∑ wiu(2i)(b)</p>
      <p>
        i=1
over b ∈ IRp for specified (given) magnitudes of
nonnegative weights w1, w2, . . . , wn. We need to note that the
knowledge of these magnitudes is crucial for the
estimator to be properly defined (i.e. avoiding a circular
definition) as (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) are fixed as well allowing to assign the
weights directly to be non-increasing with respect to
residuals (cf. [25, 24, 23]). An adaptation of the approximate
algorithm of [20] may be used for computing the LWS
algorithm.
      </p>
      <p>Concerning the choice of weights, a general
recommendation can be given to require the sequence w1, . . . , wn
nonn
increasing with ∑i=1 wi = 1. Some choices (from the
simplest to the most complicated) include
• Linearly decreasing weights
wLD = 2(n − i + 1)
i n(n + 1)
, i = 1, . . . , n.</p>
      <p>
        (
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
• Linearly decreasing weights for a true level of
contamination ε · 100 %. Let us assume ε ∈ [0, 1/2) and
h = dεne, where
The weights equal
      </p>
      <p>dxe = min{n ∈ IN; n ≥ x}.
wi =
(h − i + 1)/h, i ≤ h,
0, i &gt; h.
• Data-dependent adaptive weights of [6]. With such
weights, the LWS estimator attains a 100 %
asymptotic efficiency of the least squares under Gaussian
errors.</p>
      <p>
        Weights according to (
        <xref ref-type="bibr" rid="ref7">7</xref>
        ) as well as the adaptive weights
of [6] ensure a high breakdown point of the LWS
estimator, which cannot be said about weights (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ).
      </p>
      <p>
        The least trimmed squares (LTS) estimator (e.g. [18])
denoted as bLT S represents a special case of the LWS with
weights equal either to zero or one. The LTS estimator
depends on the value of the trimming constant h, requiring
wh+1 = · · · = wn = 0 and n/2 &lt; h &lt; n.
(
        <xref ref-type="bibr" rid="ref8">8</xref>
        )
      </p>
      <p>
        The advantages of the LWS compared to the LTS
include sub-sample robustness, more delicate approach for
dealing with moderately outlying values, robustness to
heteroscedasticity [23], the possibility to derive
diagnostic tools and to define a corresponding robust correlation
coefficient or estimator ofσ 2 [16, 17].
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
      </p>
      <p>In this section, a new estimate of the variance for the least
trimmed squares and least weighted squares estimators is
proposed in Section 3.1. The proposal enables a
systematic comparison of their estimation performances. An
illustration on a real data set is presented in Section 3.2.
3.1</p>
      <sec id="sec-3-1">
        <title>Bootstrap Variance Estimation for the LWS</title>
      </sec>
      <sec id="sec-3-2">
        <title>Estimator</title>
        <p>The aim of this section is to apply resampling (bootstrap)
techniques to estimate var bLT S and var bLW S. The
resulting bootstrap estimates are conceptually simple and can
be computed for real data in a straightforward (although
rather computationally demanding) way.</p>
        <p>Let us first recall basic principles of bootstrap
estimation (bootstrapping), which has found a big popularity in
various statistical tasks. In general, bootstrap estimation
exploits resampling with replacement. Incorporating the
basic principles of bootstrapping, one may develop a great
variety of resampling techniques that provide us with new
possibilities of analyzing data. The range of bootstrap
methods is rather large, including residual bootstrap,
nonparametric bootstrap, semiparametric bootstrap, Bayesian
bootstrap etc. Also the terminology is not used in a unique
way. Unfortunately, not much can be said about properties
of bootstrap estimates on a general level. Interesting
comparisons of bootstrap procedures in a regression setup (but
for the total least squares) were presented in [19], where
some bootstrap approaches are valid but some are proven
not to be consistent and thus not suitable.</p>
        <p>In the specific task of estimating variability of
regression estimators, the bootstrapping approach is very
suitable. While the seminal work [3] described bootstrap
estimation from a philosophical perspective, practical
approaches to bootstrapping in linear regression were
proposed by subsequent papers [9] or [8]. Other theoretical
results were derived in [10, 11].</p>
        <p>In this paper, we focus our attention to nonparametric
bootstrap. However, a residual bootstrap may be a suitable
alternative as well.</p>
        <p>
          Under (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ), we recall that
var bLS = σ 2(X T X )−1.
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
An explicit formula for var bLW S could be derived as
analogy to [24]. Such result however remains impossible to be
directly computed for real data, as it depends on
• Unknown magnitudes of e1, . . . , en,
• The asymptotic value of (X T X )/n, which can be
however hardly evaluated for a fixed sample size.
Such approach is also complicated because of the
necessity to express the weights by means of a weight function
and there is also a more restrictive assumption of normally
distributed errors. Therefore, we take resort to a bootstrap
estimate of var bLW S. Its computation is described by
Algorithm 1, where the final estimate has the form of a
bootstrap covariance matrix in step 5. An analogous procedure
can be used for estimating var bLT S.
        </p>
        <p>Algorithm 1 Nonparametric bootstrap for the LWS in
linear regression.</p>
        <p>
          Input: Data rows (Xi1, . . . , Xip,Yi), i = 1, . . . , n
Output: Empirical covariance matrix computed from
individual estimates of γˆLW S
1: Compute the least weighted squares estimator βˆLW S
of β in model (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
2: for r = 1 to R do // repeat in order to obtain the
empirical distribution
3: Generate n new bootstrap data rows
((r)Xj∗1, . . . , (r)Xj∗p, (r)Yj∗),
j = 1, . . . , n,
(
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
4:
by sampling with replacement from the original set
of data rows (Xi1, . . . , Xip,Yi), i = 1, . . . , n
Consider a linear regression model in the form
(r)Yj∗ = (r)γ0 + (r)γ1(r)Xj∗1 + · · · + (r)γp(r)Xj∗p + (r)v j
(
          <xref ref-type="bibr" rid="ref11">11</xref>
          )
with j = 1, . . . , n and random errors v1, . . . , vn
5: Estimate (r)γ = ((r)γ0, (r)γ1, . . . , (r)γp)T in (
          <xref ref-type="bibr" rid="ref11">11</xref>
          ) by
the LWS
6: Store the estimate from the previous step as (r)γˆLW S
7: end for
8: Compute the empirical covariance matrix from values
(r)γˆLW S, r = 1, . . . , R
3.2
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Example: Linear Model for Investment Data</title>
        <p>The aim of the following example is to compare variance
estimates of the LWS estimator with variances of other
regression estimators. An investment data set that considers
a regression of n = 22 yearly values of real gross private
domestic investments in the USA in 109 USD against the
GDP is used. We consider a linear model</p>
        <p>
          Yi = β0 + β1Xi + ei, i = 1, . . . , n,
(
          <xref ref-type="bibr" rid="ref12">12</xref>
          )
while the same data set was previously analyzed
(considering another model) in [17].
        </p>
        <p>The computations were performed in R software for the
least squares, Huber’s M-estimator (see [13]), LTS and
LWS. Table 1 presents estimates of the intercept b0 and
slope b1 together with packages of R software, which were
used for the computation. The bootstrap procedure of
Section 3.1 was used to find the covariance matrix of various
robust regression estimators and the results are also
presented in Table 1. There, the standard deviation of all
estimates is denoted as s0 for the intercept and s1 for the slope.</p>
        <p>
          For the least squares, the bootstrap estimates are very
close to the exact result (
          <xref ref-type="bibr" rid="ref9">9</xref>
          ). The number of bootstrap
repetitions within Algorithm 1 is chosen as 10 000, which is
sufficient for the asymptotics. Actually, we compared
results obtained for 100 bootstrap samples and the results
were very close. This is in accordance with our experience
and a small number of bootstrap samples indeed seems to
be sufficient if a single constant is estimated rather than
the whole empirical distribution.
        </p>
        <p>The smallest variance is obtained with the least squares
estimator. Huber’s M-estimator attains only a slightly
higher variance. The loss of the LTS is remarkable.
However, we point out at the closeness of the LWS result
(computed with weights [6]) to the least squares or Huber’s
Mestimator compared to the very crude LTS result. Both the
LTS and LWS are highly robust, but their performance for
this data set without severe outliers reveals a great
difference between them in terms of efficiency. The LWS
cannot be the winner and must stay behind the least squares,
but its retardation is only mild and its superiority against
the LTS shows that the LWS (perceived as a
generalization of the LTS) eliminates the main disadvantage of the
LTS, namely its low efficiency for non-contaminated
samples [7].</p>
        <p>To the best of our knowledge, the superiority of the
LWS compared to the LTS in terms of efficiency has never
been presented in the literature. Our result indicates a
possibly strong argument in favor of the efficiency of the
LWS, at least for a single data set, while no theoretical
result on the relative efficiency of LWS compared to LWS
is available. Our result is empirical, which is obtained for
a rather simplistic data set with only a single regressor, was
obtained as an application of the nonparametric bootstrap
estimation of the robust regression estimates proposed in
Section 3.1.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Nonparametric Bootstrap in Robust</title>
    </sec>
    <sec id="sec-5">
      <title>Nonlinear Regression</title>
      <p>In this section, the standard nonlinear regression model is
recalled in Section 4.1 and the nonlinear least weighted
squares estimator in Section 4.2. A two-stage version of
the nonlinear least weighted squares is proposed in
Section 4.3 as an extension of the methodology of Section 3.1.
In addition, a two-stage version of the nonlinear least
weighted squares estimator is proposed and theoretically
investigated.
4.1</p>
      <sec id="sec-5-1">
        <title>Nonlinear Regression Model</title>
        <p>
          Let us consider the nonlinear regression model
Yi = f (β1Xi1 + · · · + βpXip) + ei, i = 1, . . . , n,
(
          <xref ref-type="bibr" rid="ref13">13</xref>
          )
where f is a given continuous nonlinear function, Y =
(Y1, . . . ,Yn)T is a continuous response and (X1 j, . . . , Xn j)T
is the j-th regressor for j = 1, . . . , p. By e1, . . . , en we
denote the model’s random errors. The classical estimator,
which is the nonlinear least squares (NLS) estimator of β ,
is vulnerable to the presence of outliers in the data.
        </p>
        <p>
          The nonlinear least trimmed squares (NLTS) estimator
represents a natural extension of the LTS estimator to the
nonlinear model (
          <xref ref-type="bibr" rid="ref13">13</xref>
          ). The breakdown point of the NLTS
was derived already in [22], other properties were later
investigated in [5]. The estimator may achieve a high
robustness, if of course a suitable value of h is used reflecting the
true contamination level in the data. An approximate
algorithm may be obtained as an extension of the algorithm
of [20]; however, it requires a tedious implementation and
we are not aware of any implementation of NLTS in
statistical software.
4.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Nonlinear Least Weighted Squares</title>
        <p>This section recalls the definition of the nonlinear least
weighted squares (NLWS) estimator, which was proposed
as our extension of the LWS estimator from the linear
regression to the nonlinear model [14] and at the same time
a weighted analogy of the NLTS estimator. So far,
theoretical properties of the NLWS have not been derived and
there is also no evidence in examples showing the
robustness and efficiency of the estimator.</p>
        <p>
          In the model (
          <xref ref-type="bibr" rid="ref13">13</xref>
          ), let
ui(b) = Yi − f (b1Xi1 − · · · − bpXip), i = 1, . . . , n, (
          <xref ref-type="bibr" rid="ref14">14</xref>
          )
denote a residual corresponding to the i-th observation for
a given estimator b = (b1, . . . , bp)T ∈ IRp of regression
parameters β = (β1, . . . , βp)T . Let us now assume the
magnitudes w1, w2, . . . , wn of nonnegative weights to be given.
The NLWS estimator of the parameters in the model (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) is
defined as
        </p>
        <p>
          n
arg min ∑ wiu(2i)(b),
i=1
(
          <xref ref-type="bibr" rid="ref15">15</xref>
          )
where the argument of the minimum is computed over all
possible values of b = (b1, . . . , bp)T and the residuals are
arranged as in (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ).
        </p>
        <p>
          The choice of weights has a determining influence on
properties of the NLWS estimator. If it is allowed to have
zero weights for the most outlying observations, then the
estimator can be conjectured to be highly robust, which
follows directly from the assignment of implicit weights
to the observations in (
          <xref ref-type="bibr" rid="ref15">15</xref>
          ). Again, weights (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) (but not
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )) ensure a high breakdown point. The NLWS estimator
with such weights is highly robust from the same reasons
as the LWS estimator in the linear regression. The main
reason for the robustness of the NLWS estimator is the
construction of the estimator itself, just like for the LWS
estimator in the linear regression.
        </p>
        <p>
          An approximate algorithm for the optimization task (
          <xref ref-type="bibr" rid="ref23">23</xref>
          )
can be obtained as a straightforward adaptation of the LTS
algorithm for the linear regression (cf. [20, 14]);
nevertheless, its properties in this context have not been
investigated. Empirical investigations will be performed on real
data sets in Section 5.
4.3
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Two-stage NLWS</title>
        <p>In this section, a version of the NLWS estimator is
proposed, which constructs data-dependent adaptive weights.
The estimator has a two-stage structure and is inspired by
a two-stage LWS estimator of [6].</p>
        <p>Cˇ ížek [6] proved his two-stage estimator with
quantilebased adaptive weights in the linear model to possess
a high breakdown point and at the same time a 100 %
asymptotic efficiency of the least squares under Gaussian
errors. Further, he evaluated its relative efficiency to be
high (over 85 %) compared to maximum likelihood
estimators in a numerical study under various distributional
models for samples of several tens of observations.</p>
        <p>
          We propose a two-stage estimator denoted as 2S-NLWS
which can be described as an improved version of the
NLWS estimator which contains a construction of
datadependent adaptive weights. The model (
          <xref ref-type="bibr" rid="ref13">13</xref>
          ) is
considered. The computation of the 2S-NLWS starts with an
initial highly robust estimator βˆ0 of β and proceeds to
proposing values of the weights based on comparing the
empirical distribution function of squared residuals with its
theoretical counterpart assuming normality.
        </p>
        <p>
          In the first stage, it is crucial to choose a suitable initial
estimator, because it influences the properties of the
resulting 2S-NLWS estimator. Therefore, it is
recommendable to use a consistent estimator which is highly robust,
i.e. NLTS with h between (say) n/2 and 3n/4 or NLWS
with weights (
          <xref ref-type="bibr" rid="ref7">7</xref>
          ). Residuals of the initial fit will be
denoted as u01, . . . , un0. We will need the notation G0n −1 for
the empirical quantile function computed from these
residuals, Fχ−1 for the quantile function of χ12 distribution and
bn = min n mn ; u(2m) &gt; 0o .
(
          <xref ref-type="bibr" rid="ref16">16</xref>
          )
(
          <xref ref-type="bibr" rid="ref17">17</xref>
          )
(
          <xref ref-type="bibr" rid="ref19">19</xref>
          )
(
          <xref ref-type="bibr" rid="ref20">20</xref>
          )
        </p>
        <p>In the second stage, the weights for the 2S-NLWS
estimator are constructed. They are defined by means of
a weight function w˜(t) for t ∈ [0, 1], where
w˜(t) =</p>
        <p>Fχ−1 (max{t, bn})
(G0n)−1 (max{t, bn})
.</p>
        <p>In other words, weights for a fixed number of
observations n are given as
w˜(t) =</p>
        <p>Fχ−1(t)
(G0n)−1 (t)
for t ∈
1 3</p>
        <p>,
2n 2n
, . . . ,
2n − 1
2n</p>
        <p>
          .
(
          <xref ref-type="bibr" rid="ref18">18</xref>
          )
        </p>
        <p>The computation of the 2S-NLWS estimator is
straightforward. In a non-contaminated model, the 2S-NLWS
estimator can be easily proven to have a full efficiency of the
least squares, just like in the linear case [6].</p>
        <p>Theorem 1. Random vectors X1, . . . , Xn are assumed to be
independent identically distributed. Let e1, . . . , en be
independent identically distributed, independent on X1, . . . , Xn
and fulfilling e i ∼ N(0, σ 2) for each i. Let the initial
estimator βˆ0 be consistent with a corresponding consistent
estimator of σ 2. Then it holds</p>
        <p>
          P
w˜(t) −→ σ 2 for each t ∈ (
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ),
        </p>
        <p>P
where −→ denotes the convergence in probability.</p>
        <sec id="sec-5-3-1">
          <title>Proof. Analogy of [6].</title>
          <p>Corollary 1. Under the assumptions of Theorem 1, the
2S-NLWS estimator βˆ2S−NLW S fulfils
ˆ P
β2S−NLW S −→ β .</p>
          <p>Concerning other properties of the 2S-NLWS estimator,
the breakdown point seems to require much more effort
to be derived. Nevertheless, it remains clear that
robustness properties of the 2S-NLWS are strongly influenced
by those of the initial estimator.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Examples on Robust Estimation in</title>
    </sec>
    <sec id="sec-7">
      <title>Nonlinear Regression</title>
      <p>This section presents three examples investigating the
performance of various estimators in the nonlinear
regression, especially focused on the soundness of the
proposed methodology of Section 4. An example
illustrating the performance of the NLWS estimator in both
noncontaminated and contaminated data is presented in
Section 5.1. The next example in Section 5.2 investigates
the tightness of an approximate algorithm for computing
the NLWS estimator. The final example presented in
Section 5.3 exploits the nonparametric bootstrap estimation of
the variance of various estimators in the nonlinear
regression model.</p>
      <sec id="sec-7-1">
        <title>Estimator b1</title>
        <p>Original data set</p>
        <p>NLS 190.8</p>
        <p>NLTS (h = 18) 191.6
NLWS (weights [6]) 191.4</p>
        <p>Contaminated data set</p>
        <p>NLS 193.9</p>
        <p>NLTS (h = 18) 191.8
NLWS (weights [6]) 191.6
b2
This example has the aim to compare the performance
of classical and robust estimators in the nonlinear model.
This will be performed on a real data set without
apparent outliers. To show the sensitivity of the NLS and on
the other hand the robustness (resistance) of the NLTS and
NLWS, the computations are repeated on a modified
version of this data set, which contains one outlier.</p>
        <p>A standard data set called Puromycin with n = 23
observations, which is available in the package datasets of
R software, is considered. The reaction velocity (rate) Y is
explained as a response of the substrate concentration X in
the nonlinear regression model</p>
        <p>Yi =</p>
        <p>
          β1Xi
β2 + Xi
+ ei, i = 1, . . . , n,
(
          <xref ref-type="bibr" rid="ref21">21</xref>
          )
where the aim is to estimate regression parameters β1
and β2.
        </p>
        <p>The results of the least squares and NLWS estimators
are shown in Table 2. The NLWS estimator turns out to
perform reliably on a data set contaminated by outlying
measurements as well as on data without such
contamination. In addition, we verified the constant R = 10 000
in Algorithm 1 to be more than sufficient in the nonlinear
regression model and a moderate sample size.</p>
        <p>Further, we also consider a contaminated data set,
obtained by modifying the value of the observation in the
Puromycin data set. Particularly, the substrate
concentration (i.e. the regressor) of the first observation was
modified from 0.02 to 0.05 to become the only outlier in the data
set. This reveals the influence of a (local) change of one
observation on the results and reveals the true advantage
of the robust estimators. Robust estimates namely remain
almost unchanged, while the contamination is revealed on
the NLS estimator. In other words, the NLS starts to differ
from the robust estimators, while all estimates were much
more similar for the original data set.
NLWS fit with the linearly decreasing weights. The NLS
fit has the tendency to fit well also influential data points.
The robust fit better explains a subset of data points, while
it considers data points corresponding to larger values of
the regressor to be outliers.</p>
        <p>Table 3 gives values of loss functions
The performance of the NLWS estimator will be now
illustrated and investigated on a numerical example with
simulated data. The data set consisting of 8 data points is shown
in Figure 1. The nonlinear regression model is used in the
form</p>
        <p>
          Yi = β0 + β1(Xi − β2)2 + ei, i = 1, . . . , n,
(
          <xref ref-type="bibr" rid="ref22">22</xref>
          )
where Y1, . . . ,Yn are values of the response, X1, . . . , Xn
values of the only regressor, β0, β1 and β2 are regression
parameters and e1, . . . , en are random errors. Figure 2 shows
fitted values corresponding to the NLS fit and also the
        </p>
      </sec>
      <sec id="sec-7-2">
        <title>Estimator NLS</title>
        <p>
          NLTS (h = 5)
NLWS (weights (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ))
        </p>
        <p>
          Loss function
(
          <xref ref-type="bibr" rid="ref23">23</xref>
          ) (
          <xref ref-type="bibr" rid="ref24">24</xref>
          ) (
          <xref ref-type="bibr" rid="ref25">25</xref>
          )
23.44 2.47 1.21
46.06 1.93 1.11
70.94 6.82 0.67
corresponding to the NLS, NLTS and NLWS, respectively.
These are evaluated for all of the three estimates.
        </p>
        <p>
          The NLS estimator minimizes (
          <xref ref-type="bibr" rid="ref23">23</xref>
          ) as expected and thus
can be expected to yield also a rather small value of (
          <xref ref-type="bibr" rid="ref25">25</xref>
          ).
The NLWS estimator has a much larger value of (
          <xref ref-type="bibr" rid="ref23">23</xref>
          )
compared to the NLS fit. However, the algorithm used for
computing the NLWS has found even a much smaller value
of (
          <xref ref-type="bibr" rid="ref25">25</xref>
          ) than the NLS. On the whole, the results of Table 3
thus give a clear evidence in favor of the reliability of the
algorithm for computing the NLWS estimator.
The bootstrap based variance estimation procedure
described in Section 3.1 can be also utilized in the context of
nonlinear regression models. The following example
incorporates such an approach in order to compare the
variances of the NLS, NLTS, and NLWS estimators.
        </p>
        <p>The validity of the nonparametric bootstrap in linear
regression model (i.e., Algorithm 3.1) is, however, going to
be verified only via a simulation study. Theoretical
justification of the nonparametric bootstrap procedure needs
to be provided by a formal proof with properly stated
assumptions. In general, bootstrapping should be used with
caution, because the nonparametric bootstrap algorithm
does not always provide a consistent estimate.</p>
        <p>The same data set is used as in Section 3.2. This time,
a nonlinear regression model</p>
        <p>Yi = β1(Xi − X¯)2 + β2Xi + β3 + ei, i = 1, . . . , n, (26)
is considered, where the centering of the regressor using
its mean X¯is done for the sake of numerical stability.</p>
        <p>Three robust estimators are computed together with
bootstrap estimates of their variances. The results are
shown in Table 4. The conclusions are analogous to those
of the example in Section 3.2, namely the NLTS
estimator loses its efficiency very much compared to the NLS.
The NLS remains to be the most efficient, i.e. retains the
smallest variance for the data set which does not contain
severe outliers. Still, the NLWS loses relatively little
compared to the NLS while it is able to outperform the NLTS
strongly. Thus, we can say that the NLWS estimator is
able to combine the robustness with efficiency reasonably
well, in comparison to the non-efficient (but much more
renowned) NLTS estimator.
6</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>This paper investigates robust estimator for the linear and
nonlinear regression methods and nonparametric bootstrap
approaches to estimating its variance. Implicitly weighted
estimators are considered, which include the least trimmed
squares and least weighted squares (in linear and nonlinear
versions).</p>
      <p>After recalling the state of the art on implicitly weighted
robust estimation in Section 2, a bootstrap method for
estimating the variance of the LWS estimator is proposed in
Section 3. A numerical example shows the LWS to have
a much smaller variance compared to the more popular
LTS estimator, which reveals another strong argument in
favor of the LWS estimator and questions whether the LTS
estimator deserves to be the most common highly robust
regression estimator.</p>
      <p>Considerations for the linear regression are further
generalized to nonlinear regression. Thus, the main
contribution can be found in Section 4 on the NLWS estimation,
which has not been much investigated in the references
so far. Our work is devoted to a bootstrap estimate
suitable for estimating the variance of the NLWS estimator.
In addition, a new version of the estimator is proposed,
which computes data-dependent adaptive weights. Their
construction allows to define the 2S-NLWS, which is
improved compared to the basic NLWS in terms of efficiency.</p>
      <p>Several examples reveal the suitability of the idea of the
NLWS for modelling a nonlinear trend in the data. It also
follows from the set of examples that an approximate
algorithm, which is available for the NLWS, turns out to be
reliable. The examples also give a warning that the NLWS
estimator behaves in a rather intricate way and the
estimator is much more complex compared to linear regression.</p>
      <p>We also computed nonparametric bootstrap estimate of
the variance of nonlinear estimators. These empirical
results allow us to conclude that there seems a major
advantage of the NLWS compared to the NLTS, although the
NLS remains to be recommendable for data without
severe outliers. The NLWS seems to be close to a
reasonable combination of the efficiency (for normal regression
errors) with high robustness (for models with
contamination), which represents a dream of robust statisticians since
the dawn of robust statistical inference.</p>
      <p>The examples investigated in this paper lead us also to
formulating the following disadvantages of robust
estimators in nonlinear regression:
• They require various tuning constants with a difficult
interpretation;
• Various robust methods yield rather different results;
• Computational intensity;
• Robustness only with respect to outliers but not to
a misspecification of the model.</p>
      <p>Important limitations of the robust nonlinear
estimation include a non-robustness to small modifications of the
nonlinear function f in the model (see [2]). Robust
nonlinear estimators also require rather tedious proofs of their
properties.</p>
      <p>Some of the properties of robust estimators valid in the
linear regression are not valid in the nonlinear model at
all. As an example let us mention diagnostic tools, which
can be derived for robust estimators in linear regression,
but would be rather controversial in the nonlinear model
[15]. The results on simulated data in Section 5.2 reveal
that the residuals are far from homoscedasticity, even if
the assumption of homoscedastic disturbances in the
regression model is fulfilled. Thus, we find residuals to be
unsuitable for making conclusions about the disturbances
(random errors). While tests from the linear regression are
no longer valid for the NLWS estimator, we do not
recommend to use residuals even for a subjective diagnostics
concerning the disturbances.</p>
      <p>We intend to apply the robust regression methods of this
paper within a future research in the area of metalearning,
which aims at comparing the suitability of various
machine learning methods for different data. Robustifying
metalearning for regression methods is however not only
a matter of using robust regression methods, but the
process of metalearning itself suffers from instability [21] and
a robust analogy of the whole process of metalearning is
highly desirable to be performed in a complex and
systematic way.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The project was supported by the project 17-01251S
“Metalearning for extracting rules with numerical
consequents” of the Czech Science Foundation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Agresti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Analysis of ordinal categorical data. 2nd edn</article-title>
          ., Wiley, New York,
          <year>2010</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baldauf</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>J.M.C.S.</given-names>
          </string-name>
          :
          <article-title>On the use of robust regression in econometrics</article-title>
          .
          <source>Economic Letters</source>
          <volume>114</volume>
          (
          <year>2012</year>
          )
          <fpage>124</fpage>
          -
          <lpage>127</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bickel</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Freedman</surname>
            ,
            <given-names>D.A.:</given-names>
          </string-name>
          <article-title>Some asymptotic theory for the bootstrap</article-title>
          .
          <source>Annals of Statistics</source>
          <volume>9</volume>
          (
          <year>1981</year>
          )
          <fpage>1196</fpage>
          -
          <lpage>1217</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Brazdil</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giraud-Carrier</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilalta</surname>
          </string-name>
          , E.: Metalearning:
          <article-title>Applications to data mining</article-title>
          . Springer, Berlin,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cˇížek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Least trimmed squares in nonlinear regression under dependence</article-title>
          .
          <source>Journal of Statistical Planning and Inference</source>
          <volume>136</volume>
          (
          <year>2006</year>
          )
          <fpage>3967</fpage>
          --
          <lpage>3988</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Cˇížek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Semiparametrically weighted robust estimation of regression models</article-title>
          .
          <source>Computational Statistics &amp; Data Analysis</source>
          ,
          <volume>55</volume>
          (
          <year>2011</year>
          )
          <fpage>774</fpage>
          -
          <lpage>788</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Cˇížek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Reweighted least trimmed squares: An alternative to one-step estimators</article-title>
          .
          <source>Test</source>
          <volume>22</volume>
          (
          <year>2013</year>
          )
          <fpage>514</fpage>
          -
          <lpage>533</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Efron</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tibshirani</surname>
            ,
            <given-names>R.J.:</given-names>
          </string-name>
          <article-title>An introduction to the bootstrap</article-title>
          . Chapman &amp; Hall/CRC, Boca Raton,
          <year>1994</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Freedman</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          :
          <article-title>Bootstrapping a regression equation: Some empirical results</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>79</volume>
          (
          <year>1984</year>
          )
          <fpage>97</fpage>
          -
          <lpage>106</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Godfrey</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Bootstrap tests for regression models</article-title>
          .
          <source>Palgrave Macmillan</source>
          , London,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , DiCiccio, T.J.,
          <string-name>
            <surname>Romano</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>On smoothing and the bootstrap</article-title>
          .
          <source>Annals of Statistics</source>
          <volume>17</volume>
          (
          <year>1989</year>
          )
          <fpage>692</fpage>
          -
          <lpage>704</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Jurczyk</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Ridge least weighted squares</article-title>
          .
          <source>Acta Universitatis Carolinae Mathematica et Physica</source>
          <volume>52</volume>
          (
          <year>2011</year>
          )
          <fpage>15</fpage>
          -
          <lpage>26</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Jurecˇková</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sen</surname>
            ,
            <given-names>P.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Picek</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Methodology in robust and nonparametric statistics</article-title>
          . CRC Press, Boca Raton,
          <year>2012</year>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kalina</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Some robust estimation tools for multivariate models</article-title>
          .
          <source>Proceedings International Days of Statistics and Economics MSED</source>
          <year>2015</year>
          , Melandrium, Slaný,
          <year>2015</year>
          ,
          <fpage>713</fpage>
          -
          <lpage>722</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Kalina</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>On robust information extraction from highdimensional data</article-title>
          .
          <source>Serbian Journal of Management</source>
          <volume>9</volume>
          (
          <year>2014</year>
          )
          <fpage>131</fpage>
          -
          <lpage>144</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Kalina</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Implicitly weighted methods in robust image analysis</article-title>
          .
          <source>Journal of Mathematical Imaging and Vision</source>
          <volume>44</volume>
          (
          <year>2012</year>
          )
          <fpage>449</fpage>
          -
          <lpage>462</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Kalina</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Least weighted squares in econometric applications</article-title>
          .
          <source>Journal of Applied Mathematics, Statistics and Informatics</source>
          <volume>5</volume>
          (
          <year>2009</year>
          )
          <fpage>115</fpage>
          -
          <lpage>125</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Mount</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Netanyahu</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piatko</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silverman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
          </string-name>
          , A.Y.:
          <article-title>On the least trimmed squares estimator</article-title>
          .
          <source>Algorithmica</source>
          <volume>69</volume>
          (
          <year>2014</year>
          )
          <fpage>148</fpage>
          -
          <lpage>183</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Pešta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Total least squares and bootstrapping with application in calibration</article-title>
          .
          <source>Statistics: A Journal of Theoretical and Applied Statistics</source>
          <volume>47</volume>
          (
          <year>2013</year>
          )
          <fpage>966</fpage>
          -
          <lpage>991</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Rousseeuw</surname>
            , P.J., van Driessen,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Computing LTS regression for large data sets</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          <volume>12</volume>
          (
          <year>2006</year>
          )
          <fpage>29</fpage>
          -
          <lpage>45</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Smith-Miles</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baatar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wreford</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Towards objective measures of algorithm performance across instance space</article-title>
          .
          <source>Computers and Operations Research</source>
          <volume>45</volume>
          (
          <year>2014</year>
          )
          <fpage>12</fpage>
          -
          <lpage>24</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Stromberg</surname>
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruppert</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Breakdown in nonlinear regression</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          <volume>87</volume>
          (
          <year>1992</year>
          )
          <fpage>991</fpage>
          -
          <lpage>997</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Víšek</surname>
            <given-names>J.Á.</given-names>
          </string-name>
          :
          <article-title>Consistency of the least weighted squares under heteroscedasticity</article-title>
          .
          <source>Kybernetika</source>
          <volume>47</volume>
          (
          <year>2011</year>
          )
          <fpage>179</fpage>
          -
          <lpage>206</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Víšek</surname>
            <given-names>J.Á.</given-names>
          </string-name>
          :
          <article-title>The least trimmed squares</article-title>
          ,
          <string-name>
            <surname>Part</surname>
            <given-names>III</given-names>
          </string-name>
          :
          <article-title>Asymptotic normality</article-title>
          .
          <source>Kybernetika</source>
          <volume>42</volume>
          (
          <year>2006</year>
          )
          <fpage>203</fpage>
          -
          <lpage>224</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Víšek</surname>
            <given-names>J.Á.</given-names>
          </string-name>
          :
          <article-title>Regression with high breakdown point</article-title>
          . In Antoch J.,
          <string-name>
            <surname>Dohnal</surname>
            <given-names>G</given-names>
          </string-name>
          . (Eds.):
          <source>Proceedings of ROBUST</source>
          <year>2000</year>
          , Summer School of J CˇMF,
          <source>JCˇ MF and Czech Statistical Society</source>
          , Prague,
          <year>2001</year>
          ,
          <fpage>324</fpage>
          -
          <lpage>356</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>