<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Statistical and Geometrical Approaches to Homogeneity Testing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dmitriy Klyushin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kateryna Golubeva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Taras Shevchenko National University of Kyiv, Ukraine</institution>
          ,
          <addr-line>03680, Kyiv, Akademika Glushkova Avenue 4D</addr-line>
        </aff>
      </contrib-group>
      <fpage>254</fpage>
      <lpage>262</lpage>
      <abstract>
        <p>Generalizations based on finite training sets are based on the fundamental compactness hypothesis, which states that objects belonging to the same class in the feature space should be close to each other and far from objects of other classes. This concept became generally accepted and was carefully studied. Meanwhile, the hypothesis of compactness is purely geometric in nature and essentially uses the concept of a metric in the feature space. In particular, it generates in a completely natural way the well-known nearest neighbor method, which uses a classifier based on geometric distance. However, this hypothesis ignores the probabilistic nature of the features. For unimodal distributions it works well, but in the general case, this hypothesis may not hold and the generalization becomes incorrect. We propose an alternative approach based on the homogeneity hypothesis. We call homogeneous objects the objects whose features have the same distribution. From a statistical point of view, this means that they belong to the same general population. The use of the universal measure of homogeneity (Petunin's p-statistic) makes it possible to effectively use the apparatus for testing statistical hypotheses about the homogeneity of features for both nonintersecting and largely overlapping samples that do not satisfy the compactness hypothesis, and also to build new variants of statistical featureless discriminant analysis. Instead of metrics in traditional featureless recognition, we propose to use the Petuninʼs heterogeneity measure. This approach is rigorously substantiated mathematically and has demonstrated high efficiency in practical applications, specifically, in breast cancer screening.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Discriminant analysis</kwd>
        <kwd>relational analysis</kwd>
        <kwd>Kolmogorov-Smirnov test</kwd>
        <kwd>Wilcoxon test</kwd>
        <kwd>compactness hypothesis</kwd>
        <kwd>homogeneity hypothesis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The complexity of the pattern recognition is closely related to the compactness hypothesis which
allows generalizations to be made based on finite training sets. The meaning of this hypothesis is
intuitively clear: similar objects should be close to each other in the feature space, and dissimilar ones
are far away. This definition has an obvious geometric character since the concepts of nearness and
farness depend on the metric used. In a completely natural way, it generates the simplest classifier that
displays a typical pattern of thinking by precedents. This simplest classifier is called the nearest
neighbor method and recognizes the tested objects by their proximity to the training objects.</p>
      <p>The hypothesis described above has a significant drawback, since it does not take into account the
probabilistic nature of the training data. Typically, the data forms samples extracted from some
sample space. For correct identification of objects represented by these samples, it is necessary to take
into account their random nature. The distance between random samples in the sample space is also
random. Therefore, a mechanism is needed that would make it possible to correctly assess the
proximity between random samples. The geometric hypothesis of compactness does not work in this
case. To solve the problem, we introduce into consideration the concept of homogeneity of objects,
which means that these objects belong to the same general population. Since objects are identified
with samples of features, their homogeneity should be assessed using criteria for testing statistical
hypotheses about homogeneity. The purpose of the chapter is to describe the hypothesis alternative to
the compactness hypothesis and propose the modification of the featureless (relational) discriminant
analysis using a homogeneity measure instead of a distance. The paper is organized in the following
way. Section 1 contains an introduction and a general description of the problem. Section 2 provides a
survey of non-parametric tests for homogeneity and describes a homogeneity measure (p-statistics)
that is the main subject of our research. Section 3 describes connection between the p-statistics and
featureless (relational) discriminant analysis. Section 4 describes results of numerical comparisons of
popular homogeneity measures (the Kolmogorov–Smirnov statistics and the Wilcoxon statistics) with
the p-statistics. Here we show how the p-statistics allow reduction of dimensions and arranging
features with respect to their significance for recognition. Section 5 contains conclusions and
describes possible directions of the work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Two-sample homogeneity measure</title>
      <p>
        Consider samples x   x1, x2 ,..., xn  G1 and y   y1, y2 ,..., yn  G2 from populations G1 and G2
and F2 that are absolutely continuous. Let the null hypothesis be
obeying distribution functions F1
F1  F2 and the alternative hypothesis be F1  F . The samples drawn from the same populations are
2
called homogeneous. There are many tests for testing the hypotheses on samples` homogeneous:
purely nonparametric (Smirnov [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], Dickson [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Wald and Wolfowitz [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Mathisen [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Wilcoxon
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Mann–Whitney [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Wilks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] etc.) and conditionally nonparametric (Pitman [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Lehmann [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
Rosenblatt [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Dwass [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Fisz [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Barnard [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], Birnbaum [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], Jockel [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], Allen [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], Efron
and Tibshirani [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], Dufour and Farhat [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] etc.). Let us consider the Klyushin–Petunin test that is
purely non-parametric and does not use any requirements to distribution functions excepting being
absolutely continuous [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. We propose to put this test in the ground of the featureless discriminant
analysis. The Hillʼs assumption A(n) [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] states that if random values x1, x2 ,..., xn  G are
exchangeable and belong to absolutely continuous distribution then
      </p>
      <p>
        P  x  xi , x j  
j  i
where x is a sample value from a population G following an absolutely continuous distribution
function F , and xi and x j are the i-th and j-th order statistics. This assumption was proved for
independent identically distributed random values [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and for exchangeable identically distributed
random values [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. It is a basis of a nonparametric test for samples homogeneity [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Thus,
computing the relative frequency hij
      </p>
      <p>of the event ym  xi , x j  for the elements of y, we can
estimate a deviation hij from</p>
      <p>. To do this we use a confidence interval for binomial proportion.</p>
      <p>Let for definiteness construct the Wilson confidence interval Iijn   pij1 , pij2  where
hij n  g 2 2  g hij (1  hij )n  g 2 4
hij n  g 2 2  g hij (1  hij )n  g 2 4
n  g 2
n  g 2
,
.
h 
2 #  pij 
j  i
n 1
n 1 n</p>
      <p> Iijn  .</p>
      <p>
        The significance level of this interval depends on the parameter g. When g = 3 the significance
level of Iijn does not exceed 0.05 [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. P-statistics, which is a homogeneity measure of samples x and
y, is defined by the equation
      </p>
      <p>
        As far as the p-statistics is the relative frequency of the event  pij  j  i  Iijn  , similar to the
n  1 
above mention case we may construct the Wilson interval I for the p-statistics an use it as a basis of
the test: if the upper bound of I is greater than 0.95, the null hypothesis is accepted, else the null
hypothesis is rejected.
(1)
(2)
(3)
[
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. If
the null hypothesis holds, lim
n n 1
0,1 , and lim
      </p>
      <p>n n 1

If the null hypothesis is true, the events  pij </p>
      <p>
         Iijn  form a generalized Bernoulli scheme
n 1
[
        <xref ref-type="bibr" rid="ref24 ref25">24, 25</xref>
        ]. If the alternative hypothesis is true, the these events form a modified Bernoulli scheme. If
the null hypothesis can be either true or false, this trial scheme is called Matveichuk–Petunin scheme
j  i i
0,1 , then the asymptotic
significance level  of a sequence of confidence intervals Iijn is less than 0.05 [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>Real samples usually contain rounded numbers and often repeated elements (ties) occur. Thus, we
must distinguish a hypothetical sample that is drawn from hypothetical population G containing
absolutely precise numbers and an empirical sample drawn from an empirical population G
containing rounded measurement. Therefore, we shall have a sample x   x1, x2 , ..., xn  approximating
a hypothetical sample
x   x1, x2 , ..., xn </p>
      <p>Let x(1)  x(2)  ...  x(n)
and
x(1)  x(2)  ...  x(m) be
variational series of hypothetical and empirical samples.</p>
      <p>If a number x is drawn from G independently from x the Hill assumption (1) holds, hence:
where k  0,1, ..., n , x(0)  , and x(n1)   .Then
p x  x(k) , x(k1)  </p>
      <p>1</p>
      <p>,
p x  x(i) , x( j)    i  i1  ...  j1 
j  i
,
where tl  t  x(l)   is the multiplicity of x(l) . If x does not contain ties then  i  0.</p>
      <p>Consider the null hypothesis that hypothetical absolutely continuous distribution functions F1 и
F2 of hypothetical populations G1 and G2 are identical. Suppose, we have empirical samples
x   x1,..., xn  G1 and y   y1,..., yn  G2 , where G1
and G2
are corresponding empirical
populations. Construct the Wilson confidence interval Iij   pi(j1) , pi(j2)  for the probability of the event
yk  x(i) , x( j)  using its observed relative frequency. Let us denote N  # Iij 
n  n 1
2
and
compute the empirical p-statistics h 
1  j  i</p>
      <p># 
N  n 1</p>
      <p>
 Iij . Then, construct a confidence interval</p>
      <p>
  j  i 
I   p(1) , p(2)  for probability p    Iij  using h . If the upper bound of I is greater than
  n  1 
0.95, the null hypothesis is accepted, else the null hypothesis is rejected.</p>
    </sec>
    <sec id="sec-3">
      <title>3. P-statistics and relational discriminant analysis</title>
      <p>
        The relational discriminant analysis is developed in the papers of Petunin et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], Duin and
Pekalska [
        <xref ref-type="bibr" rid="ref28 ref29 ref30 ref31 ref32 ref33 ref34 ref35">28–35</xref>
        ] etc. The main idea of relational discriminant analysis consists in the replacing a
presentation of objects in feature space via vectors of feature (hypothesis on vector space) by a
proximity (similarity) to some training set using a distance in a metric space. This idea is very
productive and valuable, but implicitly it used a concept of geometric proximity space with a metrics.
This approach is invalid for random samples. Let us imagine data on a pool of cells (e.g. their areas)
measured in microscopic research. A researcher obtains a sample of real values but not an ordered
vector. So, to use a distance to measure proximity to training samples is impossible. That is why the
concept of homogeneity measure is very useful for such cases. Despite the large number of statistical
tests for samples homogeneity only Kolmogorov-Smirnov statistics, U-statistics (Wilcoxon test) and
p-statistics have the properties allowing numerical estimating the samples homogeneity (similarity in
the sense of belonging to the same population).
      </p>
      <p>For example, we can use the p-statistics to solve the problem of dimension reduction and feature
selection. Compute the proximity measure between samples from G1 and G2 with respect to two
(4)
(5)
features: ith and jth. Consider the matrixes of features of kth object from G1 and lth object from G2 ( n
is the number of features and m is the number of measured values of every feature) :</p>
      <p>Consider the ith columns corresponding to ith feature from uk and vl: Xik   x1ik , x2ki ,..., xmki T and
Yil   y1il , y2li ,..., ymli T . Then, compute p-statistics for samples (not vectors!) Xik and Yil and
construct the vector of p-statistics for uk and vl with respect to every feature:
k1l   X1k ,Y1l , k2l   X 2k ,Y2l , ..., knl   X Nk ,YNl .</p>
      <p>Then, compute the average p-statistics
k1 
1 N</p>
      <p>k1t , 2 
N t1 k
1 N</p>
      <p> k2t , ..., n 
N t1 k
1 Nn</p>
      <p>kt</p>
      <p>N t1
for uk and an object from G2 with respect to ith feature. This scheme may be applied for
comparing uk with other object from G1:
k1 </p>
      <p>N11 s1N,s1k1s , k2 </p>
      <p>N11 s1N,s1k2s , ..., kn </p>
      <p>1 N kns.</p>
      <p>N 1 s1,s1</p>
      <p>Pairing p-statistics we form a proximity vector space corresponding to ith and jth features:
ti , t j  and si , s j  , i, j  1, 2,..., m; t, s  1, 2,..., n. Now, in the proximity vector space we have
two sets of points consisting of average interclass homogeneity measure and average intraclass
homogeneity measure. Thus, we may use in the proximity space any classificator developed for
metric spaces. The average intraclass homogeneity measure allows estimating intrinsic diversity of
objects in the population, and the average intraclass homogeneity measure allows estimating the
feature significance.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and results</title>
      <p>
        To estimate the sensitivity and specificity of the tests we have carried out numerical experiments
using samples from normal distribution with various parameters describing the degree of overlapping.
We used samples containing 40 random numbers with the same mean and different standard
deviations (location shift) and with different means and the same standard deviation (scale shift). We
computed the p-statistics with its lower and upper confidence bounds, the Kolmogorov–Smirnov
statistics and its p-value and the Wilcoxon statistics and its p-value. The sensitivity of the Klyushin–
Petunin test [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] was estimated as the relative frequency of the event when the upper confidence
bound for the p-statistics is less that 0.95 when distributions are different. The sensitivity of the
Kolmogorov–Smirnov test and the Wilcoxon signed-rank test was estimated as the relative frequency
of the event when corresponding p-value  0.05 when distributions were different. The sensitivity of
the Klyushin–Petunin test is considered as the relative frequency of the event when the upper
confidence bound for the p-statistics is greater that 0.95 for identical distributions. The specificity of
the Kolmogorov–Smirnov and the Wilcoxon signed-rank test is considered as the relative frequency
of the event when p-value  0.05 when distributions are identical. In this way we have tested two
statistical hypotheses: on location shift and on a scale. The null hypothesis on location shift states that
the locations of both distributions are the same. The null hypothesis on scale shift states that the
variances of both distributions are the same. The alternative hypotheses, in opposite, state that the
distribution functions are different. These cases are illustrated at Fig. 1 and Fig. 2. The results are
provided in Tables 1–8. Note a remarkable property of the p-statistics. It not only correctly recognizes
heterogeneous samples but also demonstrate monotonic decreasing as the location shift increases.
      </p>
      <p>As it was expected, in the case of location shift both the Kolmogorov–Smirnov test and the
Wilcoxon signed-rank test work perfectly. The effectiveness of the Kolmogorov–Smirnov test is
explained by the increasing discrepancy between the empirical distributions as the location shift
increases. The Wilcoxon signed rank test was developed namely for this case and effectively
recognizes inversions. However, when we test the scale shift hypothesis the situation changes. Now,
the distribution functions are largely overlapped and the discrepancy between them is not very
significant. Moreover, the Wilcoxon signed-rank test poorly recognizes the inversions between
largely overlapped samples. These statements are justified by the following results (Table 4–6).</p>
      <p>As we see, the Kolmogorov–Smirnov test has failed in the case of largely overlapped samples in
more than almost a half of the cases, and the Wilcoxon signed-rank test has failed in all the cases. The
Klyushin–Petunin test fails in almost a third of cases of very overlapped samples following the
distributions N(0,3), N(0,4) and N(0,5).</p>
      <p>In practice, data are often rounded and ties occur in samples. To simulate this effect we use the
same samples as in previous experiments but have rounded them up to two decimal digits. After
rounding, samples in average contained four ties. The results are provided in Tables 7–12. As
expected from the theoretical point of view, the results of the Kolmogorov–Smirnov test and the
Wilcoxon signed rank test have not changed comparing with the case without ties because the ties do
not affect on the Kolmogorov–Smirnov test statistics and the Wilcoxon signed rank test statistics.
Thus, we do not provide the results for these cases because they are the same as in Tables 5 and 6.</p>
      <p>It is easy to see, that ties have affected on the p-statistics and have slightly changed the monotonic
decreasing of the p-statistics with respect to the increasing of the location shift. The Klyushin-Petunin
test, as the Kolmogorov–Smirnov test, fails comparing distributions N(0,3), N(0,4) and N(0,5),
analogously to the case without ties. But it was effective in the cases when the Kolmogorov–Smirnov
test failed and the Wilcoxon signed rank test failed.</p>
      <p>Thus, we have demonstrated the prevalence of the p-statistics over the Kolmogorov–Smirnov and
the Wilcoxon signed rank tests. The Klyushin–Petunin test on homogeneity based on the p-statistics
has high sensitivity and specificity both in cases location and scale shifts and in the cases when
sample arbitrary overlapped. It is a universal test for test homogeneity of two samples and may be
successfully use in application of featureless discriminant analysis as a substitution of a metrics. The
p-statistics effectively estimates the homogeneity of the sample and takes values from the interval
(0,1). This greatly facilitates its use in the classification of objects defined by samples, in comparison
with the Kolmogorov-Smirnov test and Wilcoxon signed rank test.</p>
      <p>Since relational discriminant analysis is based on describing the proximity of data, the concept of
statistical homogeneity fits perfectly with its concept. Instead of comparing feature vectors
representing objects, we can compare samples containing their features. This makes it possible to
replace the hypothesis of compactness with the hypothesis of homogeneity and to reduce recognition
to assessing the homogeneity of samples. Thus, the p-statistic and the Klyushin-Petunin test can be a
valuable tool for pattern recognition in the paradigm of relational discriminant analysis.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and future work</title>
      <p>For correct generalization based on finite training sets, it is necessary to correctly state the
fundamental postulates. Relational discriminant analysis is based on the compactness hypothesis,
which states that objects belonging to the same class in the feature space must be close, and objects
from different classes must be distant. This geometric hypothesis is not correct in the case of
assessing the proximity between objects that are characterized not by vectors (ordered sets of
features), but by random samples (unordered sets of features), since in such cases it makes no sense to
talk about a geometric distance (metric). We propose to base the relational analysis on the hypothesis
of homogeneity, which states that objects from the same class (homogeneous) belong to the same
general population, that is, the samples of features that characterize them have the same distribution,
and objects from different classes belong to different general populations (heterogeneous), that is, the
samples of features that characterize them have different distributions. For a numerical assessment of
the homogeneity of the samples, we propose to use the Petunin p-statistics, which showed high
sensitivity and specificity in experiments both in testing the hypothesis of a mean shift and in testing
the hypothesis of the scale shift, in contrast to the statistics of Kolmogorov-Smirnov and Wilcoxon
signed rank tests. The proposed approach is strictly mathematically justified and has demonstrated
high efficiency in practical applications.</p>
      <p>The future direction of the work is to estimate theoretical power of the proposed test and develop
its multivariate version. In particular, Petunin's ellipses and ellipsoids, which are constructed on the
basis of Hill's assumption, are of great interest. With their help, one can unambiguously order random
points in a multidimensional space according to their statistical depth, similar to the Mahalanobis
distance, detect outliers, and change points in multidimensional time series. Such tasks often arise in
control systems for the timely detection of deviations from the normal operation. In the
onedimensional case, p-statistics, in contrast to the Kolmogorov–Smirnov statistics, are resistant to
random noise and universal, in contrast to the Mann–Whitney–Wilcoxon statistics.</p>
      <p>Promising methods are the transformation of multidimensional data into one-dimensional samples
for subsequent classification. For this, Fisher's linear discriminant analysis transformations can be
applied. Similarly, in the multivariate case, one can consider the average p-statistic calculated from
one-dimensional samples, followed by an estimate of its average value. Corresponding experiments
show the high efficiency of p-statistics in comparison with traditional methods.</p>
      <p>The nonparametric approach allows one to effectively solve the problems of classifying
onedimensional and multidimensional data, identify the most and least probable elements of the sample,
and rank them using statistical peeling. This is of great importance in medical applications, because
due to the unambiguous ranking of multivariate data, it becomes possible to assess the individual risk
of a particular patient, and not just the probability of his belonging to a certain group. Such
applications have become the focus of a new area of research in artificial intelligence, which is called
explainable artificial intelligence. The approach described in the paper is fully consistent with the
concept of explained artificial intelligence, since machine learning problems solved using p-statistics
allow for an accurate probabilistic interpretation.</p>
      <p>The universal and robust nature of p-statistics (its robustness to outliers and independence from the
type of hypothesis on mean or variance shift) makes it an indispensable useful tool in statistical
studies of data of any size, both small and large. With a small sample size of the strategy, p-statistics
work well in combination with the bootstrap and jackknife methods. In such cases, similarly to the
multidimensional case, the averaged p-statistics is used, which surpasses traditional analogues in its
properties. In the case of large data, the p-statistic can be computationally difficult, but fragmenting
large samples and averaging the p-statistic solves this problem as well.</p>
      <p>Thus, the presented work proves the high accuracy, sensitivity and specificity of p-statistics, its
robustness and superiority over the traditional Kolmogorov–Smirnov and Mann–Whitney–Wilcoxon
statistics. Excellent characteristics and a wide field of applications allow us to hope that p-statistics
will become a very useful tool for solving many problems of data analysis and machine learning.
6. References</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Smirnov</surname>
          </string-name>
          ,
          <article-title>Estimate of difference between empirical distribution curves in two independent samples</article-title>
          .
          <source>Bulletine of Moscow State University</source>
          <volume>2</volume>
          (
          <year>1939</year>
          )
          <fpage>3</fpage>
          -
          <lpage>14</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Smirnov</surname>
          </string-name>
          ,
          <article-title>On the deviations of an empirical distribution curve</article-title>
          .
          <source>Matematicheskii Sbornik</source>
          <volume>6</volume>
          (
          <year>1939</year>
          )
          <fpage>3</fpage>
          -
          <lpage>26</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>W. G.</given-names>
            <surname>Dixon</surname>
          </string-name>
          ,
          <article-title>A criterion for testing the hypothesis that two samples are from the same population</article-title>
          .
          <source>Annals of Mathematical Statistics</source>
          <volume>11</volume>
          (
          <year>1940</year>
          )
          <fpage>199</fpage>
          -
          <lpage>204</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177731914
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wolfowitz</surname>
          </string-name>
          ,
          <article-title>On a test whether two samples ate from the same population</article-title>
          .
          <source>Annals of Mathematical Statistics</source>
          <volume>11</volume>
          (
          <year>1940</year>
          )
          <fpage>147</fpage>
          -
          <lpage>162</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177731909
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H. C.</given-names>
            <surname>Mathisen</surname>
          </string-name>
          ,
          <article-title>A method of testing the hypothesis that two samples are from the same population</article-title>
          .
          <source>Annals of Mathematical Statistics</source>
          <volume>14</volume>
          (
          <year>1943</year>
          )
          <fpage>188</fpage>
          -
          <lpage>194</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177731460
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wilcoxon</surname>
          </string-name>
          ,
          <article-title>Individual comparisons by ranking methods,. Biometrika 1 (</article-title>
          <year>1945</year>
          )
          <fpage>80</fpage>
          -
          <lpage>83</lpage>
          . doi:
          <volume>10</volume>
          .2307/3001968
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H. B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Whitney</surname>
          </string-name>
          ,
          <article-title>On a test of whether one of the random variables is stochastically larger than other</article-title>
          .
          <source>Annals of Mathematical Statistics</source>
          <volume>18</volume>
          (
          <year>1947</year>
          )
          <fpage>50</fpage>
          -
          <lpage>60</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177730491
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Wilks</surname>
          </string-name>
          ,
          <article-title>A combinatorial test for the problem of two samples from continuous distributions</article-title>
          .
          <source>In: Proceeding of Fourth Berkeley Symposium on Mathematical Statistics and Probability</source>
          , volume
          <volume>1</volume>
          ,
          <year>1961</year>
          , pp
          <fpage>707</fpage>
          -
          <lpage>717</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>E. J. G</surname>
          </string-name>
          . Pitman,
          <article-title>Significance tests which may be applied to samples from any populations</article-title>
          ,
          <source>Journal of Royal Statistical Society Series A. 4</source>
          (
          <year>1937</year>
          )
          <fpage>119</fpage>
          -
          <lpage>130</lpage>
          . doi:
          <volume>10</volume>
          .2307/2984124
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <article-title>Consistency and unbiasedness of certain nonparametric tests</article-title>
          ,
          <source>Annals of Mathematical Statistics</source>
          <volume>22</volume>
          (
          <year>1947</year>
          )
          <fpage>165</fpage>
          -
          <lpage>179</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177729639
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Rosenblatt</surname>
          </string-name>
          ,
          <article-title>Limit theorems associated with variants of the von Mises statistic</article-title>
          .
          <source>Annals of Mathematical Statistics</source>
          <volume>23</volume>
          (
          <year>1952</year>
          )
          <fpage>617</fpage>
          -
          <lpage>623</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177729341
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dwass</surname>
          </string-name>
          ,
          <article-title>Modified randomization tests for nonparametric hypotheses</article-title>
          ,.
          <source>Annals of Mathematical Statistics</source>
          <volume>28</volume>
          (
          <year>1957</year>
          )
          <fpage>181</fpage>
          -
          <lpage>187</lpage>
          , doi:10.1214/aoms/1177729038
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fisz</surname>
          </string-name>
          ,
          <article-title>On a result be M. Rosenblatt concerning the Mises-Smirnov test</article-title>
          ,
          <source>Annals of Mathematical Statistics</source>
          <volume>31</volume>
          (
          <year>1960</year>
          )
          <fpage>427</fpage>
          -
          <lpage>429</lpage>
          . doi:
          <volume>10</volume>
          .1214/aoms/1177705905
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Barnard</surname>
          </string-name>
          , Comment on “
          <article-title>The spectral analysis of point processes” by M.S. Bartlett</article-title>
          .
          <source>Journal of Royal Statistical Society Series B</source>
          ,
          <volume>25</volume>
          (
          <year>1963</year>
          )
          <fpage>294</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Z. W.</given-names>
            <surname>Birnbaum</surname>
          </string-name>
          ,
          <article-title>Computers and unconventional test-statistics</article-title>
          . In:
          <string-name>
            <given-names>F.</given-names>
            <surname>Prochan</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. J</given-names>
            .
            <surname>Serfling</surname>
          </string-name>
          (Eeds).
          <source>Reliability and Biometry. SIAM</source>
          , Philadelphia, PA,
          <year>1974</year>
          , pp.
          <fpage>441</fpage>
          -
          <lpage>458</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>K.-H. Jockel</surname>
          </string-name>
          ,
          <article-title>Finite sample properties and asymptotic efficiency of Monte Carlo tests</article-title>
          .
          <source>Annals of Statistics</source>
          <volume>14</volume>
          (
          <year>1986</year>
          )
          <fpage>336</fpage>
          -
          <lpage>347</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Allen</surname>
          </string-name>
          ,
          <article-title>Hypothesis testing using L1-distance bootstrap</article-title>
          .
          <source>American Statistician</source>
          <volume>51</volume>
          (
          <year>1997</year>
          )
          <fpage>145</fpage>
          -
          <lpage>150</lpage>
          . doi:
          <volume>10</volume>
          .1080/00031305.
          <year>1997</year>
          .10473949
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>B.</given-names>
            <surname>Efron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          ,
          <article-title>An Introduction to the Bootstrap</article-title>
          . Vol.
          <volume>57</volume>
          of Monographs on Statistics and Applied Probability. New York: Chapman-Hall,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>J.-M. Dufour</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Farhat</surname>
          </string-name>
          ,
          <article-title>Exact nonparametric two-sample homogeneity tests for possibly discrete distributions. Center for Interuniversity research in Quantitative Economics (CIREQ)</article-title>
          .
          <source>Preprint</source>
          <year>2001</year>
          -
          <volume>23</volume>
          . California Press,
          <year>2001</year>
          , pp.
          <fpage>707</fpage>
          -
          <lpage>717</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Klyushin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Yu. I.</given-names>
            <surname>Petunin</surname>
          </string-name>
          ,
          <article-title>A Nonparametric Test for the Equivalence of Populations Based on a Measure of Proximity of Samples,</article-title>
          .
          <source>Ukrainian Mathematical Journal</source>
          <volume>55</volume>
          (
          <year>2003</year>
          )
          <fpage>181</fpage>
          -
          <lpage>198</lpage>
          . doi:
          <volume>10</volume>
          .1023/A:1025495727612
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>B.</given-names>
            <surname>Hill</surname>
          </string-name>
          ,
          <article-title>Posterior distribution of percentiles: Bayes' theorem for sampling from a population</article-title>
          ,
          <source>Journal of the American Statistician Association</source>
          ,
          <volume>63</volume>
          1968
          <fpage>677</fpage>
          -
          <lpage>691</lpage>
          . doi:
          <volume>10</volume>
          .1080/01621459.
          <year>1968</year>
          .
          <volume>11009286</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>I. Madreimov</surname>
          </string-name>
          ,
          <string-name>
            <surname>Yu. I. Petunin</surname>
          </string-name>
          ,
          <article-title>Characterization of a uniform distribution using order statistics</article-title>
          .
          <source>Teoriya Veroyatnostey i Matematicheskaya Statistika</source>
          ,
          <volume>27</volume>
          (
          <year>1982</year>
          )
          <fpage>96</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R. I.</given-names>
            <surname>Andrushkiw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Klyushin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Lysyuk</surname>
            <given-names>V. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
          </string-name>
          . I.
          <article-title>Petunin, Construction of the bulk of general population in the case of exchangeable sample values</article-title>
          .
          <source>In: Proceedings of the International Conference of Mathematics and Engineering Techniques in Medicine and Biological Science (METMBS'03)</source>
          , Las Vegas, Nevada, USA,
          <year>2003</year>
          , pp
          <fpage>486</fpage>
          -
          <lpage>489</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Matveichuk</surname>
          </string-name>
          , Yu. Petunin,
          <article-title>A generalization of the Bernoulli model occurring in order statistics</article-title>
          . I.,
          <source>Ukrainian Mathematical Journal</source>
          <volume>42</volume>
          1990
          <fpage>459</fpage>
          -
          <lpage>466</lpage>
          . doi:
          <volume>10</volume>
          .1007/BF01071335.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>S.</given-names>
            <surname>Matveichuk</surname>
          </string-name>
          , Yu. Petunin,
          <article-title>A generalization of the Bernoulli model occurring in order statistics</article-title>
          . II.,
          <source>Ukrainian Mathematical Journal</source>
          <volume>43</volume>
          (
          <year>1991</year>
          )
          <fpage>728</fpage>
          -
          <lpage>734</lpage>
          . doi:
          <volume>10</volume>
          .1007/BF01058940.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>N.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , S. Kotz,
          <article-title>Some generalizations of Bernoulli and Polya-Eggenberger contagion models</article-title>
          ,
          <source>Statistical Papers</source>
          <volume>32</volume>
          (
          <year>1991</year>
          )
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1007/BF02925473.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Yu. I. Petunin.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Klyushin</surname>
          </string-name>
          ,
          <string-name>
            <surname>R. I. Andrushkiw</surname>
          </string-name>
          ,
          <article-title>Nonlinear algorithms of mattern recognition for computer-aided diagnosis of breast cancer</article-title>
          .
          <source>Nonlinear analysis</source>
          ,
          <volume>30</volume>
          (
          <year>1997</year>
          )
          <fpage>5431</fpage>
          -
          <lpage>5336</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          , D. de Ridder,
          <string-name>
            <surname>D. M. J. Tax</surname>
          </string-name>
          ,
          <article-title>Experiments with a featureless approach to pattern recognition</article-title>
          .
          <source>Pattern Recognition Letters</source>
          <volume>18</volume>
          (
          <year>1997</year>
          )
          <fpage>1159</fpage>
          -
          <lpage>1166</lpage>
          . doi:
          <volume>10</volume>
          .1016/S0167-
          <volume>8655</volume>
          (
          <issue>97</issue>
          )
          <fpage>00138</fpage>
          -
          <lpage>4</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          , E. Pȩkalska D. de Ridder,
          <article-title>Relational discriminant analysis</article-title>
          ,
          <source>Pattern Recognition Letters</source>
          <volume>20</volume>
          (
          <year>1999</year>
          )
          <fpage>1175</fpage>
          -
          <lpage>1181</lpage>
          . doi:
          <volume>10</volume>
          .1016/S0167-
          <volume>8655</volume>
          (
          <issue>99</issue>
          )
          <fpage>00085</fpage>
          -
          <lpage>9</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Loog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pekalska</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. M. J. Tax</surname>
          </string-name>
          ,
          <article-title>Feature-based dissimilarityspace classification</article-title>
          . In: D.
          <string-name>
            <surname>Unay</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Cataltepe</surname>
            , S. Aksoy (Eds.), Recognizing Patterns in Signals, Speech, Images, and Videos,
            <given-names>ICPR</given-names>
          </string-name>
          <year>2010</year>
          , vol.
          <volume>6388</volume>
          , Springer,
          <year>2010</year>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pekalska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          ,
          <article-title>On combining dissimilarity representations</article-title>
          . In: J.
          <string-name>
            <surname>Kittler</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Roli</surname>
          </string-name>
          (Eds.),
          <source>Multiple Classifier Systems</source>
          , volume.
          <source>2096 of Lecture Notes in Computer Science</source>
          . Springer-Verlag,
          <year>2001</year>
          , pp.
          <fpage>359</fpage>
          -
          <lpage>368</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pekalska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          .
          <article-title>The Dissimilarity Representation for Pattern Recognition. Foundations and Applications</article-title>
          . World Scientific, Singapore,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pe</surname>
          </string-name>
          <article-title>̨ kalska</article-title>
          ,
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          ,
          <article-title>Dissimilarity-based classification for vectorial representations</article-title>
          .
          <source>In: Proceedings of 18th International conference on pattern recognition</source>
          ,
          <source>Hong Kong, China</source>
          ,
          <year>2006</year>
          , volume III, pp.:
          <fpage>137</fpage>
          -
          <lpage>140</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pe</surname>
          </string-name>
          <article-title>̨ kalska</article-title>
          ,
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          ,
          <article-title>Beyond traditional kernels: Classification in two dissimilarity-based representation spaces</article-title>
          .
          <source>IEEE Transactions on Systems, Man, and. Cybernetics</source>
          , Part C:
          <article-title>Applications</article-title>
          and Reviews 38 (
          <year>2008</year>
          )
          <fpage>729</fpage>
          -
          <lpage>744</lpage>
          . doi:
          <volume>10</volume>
          .1109/TSMCC.
          <year>2008</year>
          .2001687
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>E.</given-names>
            <surname>Pe</surname>
          </string-name>
          <article-title>̨ kalska</article-title>
          ,
          <string-name>
            <given-names>R. P. W.</given-names>
            <surname>Duin</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>Paclı́k, Prototype selection for dissimilarity-based classifiers</article-title>
          .
          <source>Pattern Recognition</source>
          .
          <volume>39</volume>
          (
          <year>2006</year>
          )
          <fpage>189</fpage>
          -
          <lpage>208</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.patcog.
          <year>2005</year>
          .
          <volume>06</volume>
          .012
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>