<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>D. Tasche)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Minimising quantifier variance under prior probability shift</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dirk Tasche</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Independent Researcher</institution>
          ,
          <addr-line>Zurich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>For the binary prevalence quantification problem under prior probability shift, we determine the asymptotic variance of the maximum likelihood estimator. We find that it is a function of the Brier score for the regression of the class label on the features under the test data set distribution. This observation suggests that optimising the accuracy of a base classifier, as measured by the Brier score, on the training data set helps to reduce the variance of the related quantifier on the test data set. Therefore, we also point out training criteria for the base classifier that imply optimisation of both of the Brier scores on the training and the test data sets.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Prior probability shift</kwd>
        <kwd>quantifier</kwd>
        <kwd>class distribution estimation</kwd>
        <kwd>Cramér-Rao bound</kwd>
        <kwd>maximum likelihood estimator</kwd>
        <kwd>Brier score</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        or should we rather strive to develop an optimal classifier,
possibly based on an comprehensive feature selection
The survey paper [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] described the problem to estimate process?
prior class probabilities (also called prevalences) on a test Some researchers indeed suggest that the accuracy of
set with a diferent distribution than the training set (the the base classifiers is less important for quantification
quantification problem) as “Given a labelled training set, than for classification. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] made the following statements:
induce a quantifier that takes an unlabelled test set as in- • From the abstract of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: “These strengths can
put and returns its best estimate of the class distribution.” make quantification practical for business use, even
As becomes clear from [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and also more recent work on where classification accuracy is poor.”
the problem, it has been widely investigated in the past • P. 165 of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: “The efort to develop special
purtwenty years. pose features or classifiers could increase the cost
      </p>
      <p>
        A lot of diferent approaches to quantification of prior significantly, with no guarantee of an accurate
clasclass probabilities has been proposed and analysed (see, sifier. Thus, an imperfect classifier is often all that
e.g. [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]), but it appears that the following question is available.”
has not yet received very much attention: • P. 166 of [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: “It is suficient but not necessary to
      </p>
      <p>Is it worth the efort to try to train a good (accurate) hard have a perfect classifier in order to estimate the
(or soft or probabilistic) classifier as the ‘base classifier’ for class distribution well. If the number of false
posthe task of quantification if the class labels of individual itives balances against false negatives, then the
instances are unimportant and only the aggregate prior overall count of predicted positives is nonetheless
class probabilities are of interest? correct. Intuitively, the estimation task is easier</p>
      <p>
        In principle, there is a clear answer to this question. for not having to deliver accurate predictions on
The accuracy of the classifier matters at least in the ex- individual cases.”
treme cases:
• If a classifier is least accurate because its
predictions and the true class labels are stochastically
independent, then quantification is not feasible.
• If a classifier is most accurate in the sense of making
perfect predictions then perfect quantification is
easy by applying Classify &amp; Count [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>But if no perfect classifier is around, can we be happy to
deploy a moderately accurate classifier for quantification</p>
      <p>
        The point on the mutual cancellation of false positives
and false negatives is mentioned also by a number of
other researchers like for instance [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. On p. 74, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] wrote:
“Equation 1 [with the definition of the  -measure] shows
that 1 deteriorates with (  +   ) and not with
|  −   |, as would instead be required of a function
that truly optimizes quantification.”
      </p>
      <p>
        There are also researchers that hold the contrary
position, at least as quantification under an assumption
of prior probability shift (see (1) below for the formal
definition) is concerned:
• [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] noted for the class of ‘ratio estimators’ they
introducted that it was both desirable and
feasible to construct estimators with small asymptotic
variances.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        • [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] demonstrated by a simulation study that
estimating the prior class probabilities by means of
a more accurate base classifier may entail much
shorter confidence intervals for the estimates.
      </p>
      <p>
        Prior probability shift is a special type of data set shift,
see [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for background information and a taxonomy of
• [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] pointed out the eficiency of the maximum data set shift. In the literature, also other terms are used
likelihood estimator (MLE) for the quantification for prior probability shift, for instance ‘global drift’ [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
task and made cases for its application. or ‘label shift’ [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        The problem of estimating the test set prior class
probIn the following, we revisit the question of the usefulness abilities can also be interpreted as a problem to estimate
of accurate classifiers for quantification: the parameters of a ‘mixture model’ [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] where the
component distributions are learnt on a training set. See
      </p>
      <p>
        After giving an overview of related research in Sec- [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] for an early work on the properties of the maximum
tion 2 and specifying the setting and assumptions for likelihood (ML) estimator in this case. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] revived the
the binary quantification problem in Section 3, we recall interest in the ML estimator for the unknown prior class
the technical details of the definition of the MLE for the probabilities in the test set by specifying the associated
positive class prevalence in Section 4. In particular, we ‘expectation maximisation’ (EM) algorithm.
show that the MLE is well-defined under the mild condi- [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] proposed to take recourse to the notion of Fisher
tion that the test sample consists of at least two diferent consistency as a criterion to identify completely
unsuitpoints, see (6) below. able approaches to the quantification problem that do not
      </p>
      <p>
        In Section 5, we describe, based on its representation have this property. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] then proved Fisher consistency
in terms of the Fisher information, the Cramér-Rao lower of the ML estimator under prior probability shift.
bound for the variances of unbiased estimators of the The ML approach has been criticised for its sometimes
positive class prevalence, see (9) below. This lower bound, moderate performance and the efort and amount of
trainat the same time, is the large sample variance of the MLE ing data needed to implement it. However, recently some
defined in Section 4. Thus, the Cramér-Rao lower bound researchers [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] began to vindicate the ML approach.
is achievable in theory by MLEs – which might explain They focussed on the need to properly calibrate the
posto some extent the superiority of MLEs as observed in terior class probability estimate in order to improve the
[
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. eficiency of the MLE for the positive class prevalence on
      </p>
      <p>
        In Section 6, we show for the test distribution of the the test set. Complementing the work of [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], in this
feature vector that its associated Fisher information with paper we study the role that constructing more powerful
respect to the positive class prevalence is – up to a factor (or accurate) classifiers by selection of more appropriate
depending only on the true prevalence – just the vari- features may play to reduce the variances of the related
ance of the posterior positive class probability (Proposi- quantifiers.
tion 6.1 below). The variance of the posterior probability In the following, we revisit the well-known asymptotic
is closely related to the Brier score for the regression of eficiency property of ML estimators in the special case
the class label on the feature vector. While the Brier score of the MLE for prior class probabilities in the binary
decreases when the information content of the feature setting and investigate how it is impacted by the power
vector increases, the variance of the posterior probabil- (accuracy) of the classifier the MLE is based on.
ity decreases with shrinking information content of the
features. In any case, these observations imply that the
large sample variance of the MLE (or the Cramér-Rao 3. Setting
lower bound for the variances of unbiased estimators)
may be reduced when a feature vector with larger infor- We consider the binary prevalence quantification
probmation content is selected. Section 6 concludes with two lem in the following setting:
suggestions of how this can be achieved in practice (Brier
curves, ROC analysis).
      </p>
      <p>
        In Section 7, we illustrate the observations of
Sections 5 and 6 with a numerical example. Table 1 below
demonstrates variance reduction through more powerful
features both for the ML quantifier and a non-ML
quantiifer. Figure 1 below suggests that diferences in eficiency
between the ML quantifier and non-ML quantifiers may
depend both on the information content (power) of the
feature vector and on the true value of the positive class
prevalence.
• There is a training (or source) data set
(1, 1), . . . , (, ) ∈ X × { 0, 1}. It is
assumed to be an i.i.d. sample of a random vector
(,  ) with values in X × { 0, 1}. The vector
(,  ) is defined on a probability space (Ω ,  ),
the training (or source) domain. The elements 
of Ω are the instances (or objects). Each instance
 belongs to one of the classes 0 and 1, and its
class label is  () ∈ {0, 1}. In addition, each
instance  has features () ∈ X. Often, X
is the -dimensional Euclidian space such that light on the importance of both maximum likelihood
estiaccordingly  is a real-valued random vector. See mation and accurate classifiers for the eficient estimation
Appendix B.1 of [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] for more detailed comments of the unknown positive class prevalence  in the test
of how this setting avoids the logical problems data set.
that arise when feature vectors and instances are
considered to be the same thing.
• Under the training distribution  , both the features 4. The ML estimator for the
() and the class labels  () of the instances are positive class prevalence
observed in a series of  independent experiments
resulting in the sample (1, 1), . . . , (, ). Assume that the conditional distributions in (1) have
The sample can be used to infer the joint distribu- positive densities ,  = 0, 1. Then the unconditional
tion of  and  under  , and hence, in particular, density of the features vector  under  is
also the distribution of  (the class distribution)
under  .  ()() = (1 − ) 0() +  1(),  ∈ X. (3)
• There is a test (or target) data set 1, . . . ,  ∈ X.
      </p>
      <p>It is assumed to be an i.i.d. sample of the random Hence the likelihood function
vector  with values in X, under a probability mea- () = (; 1, . . . , )
sure  on Ω that may be diferent to the training
distribution  . for the sample 1, . . . ,  is given by
• Under the test distribution , only the features 
() of the instances are observed in a series of  () = ∏︁(︀  (1() − 0()) + 0()︀) . (4)
independent experiments resulting in the sample =1
1, . . . , . The sample can be used to infer the
distribution of  under . This implies for the first two derivatives of the
log• The goal of quantification is to infer the dis- likelihood with respect to :
torfibfuetaitounreosf 1, .u.n. d,er ge,nbearsaetdedounntdheer samapnlde  log () = (5a)
(on1,th1e)j,o.i.n.t, (samp,leo)f gfeenateurraetesdanudndcelrass .laIbteilss ∑︁ 1() − 0() ,
not possible to design a method for this inference =1  (1() − 0()) + 0()
without any assumption on the relation of  and
. 2 log2 () = (5b)
• Icbnoyntphdriiistoirpopanrpaoelbrfa,ebwaitleuitaryesssdhuiimsftte,riibtnhutathtieonsesanansrdeetthhaeatrsteharmeeclealatuesnsd-- − ∑=︁1 (︂  (1(1)(−  )0−(0)()+)0() ︂) 2 ≤ 0.
der  and , i.e. it holds that</p>
      <p>We assume that there is at least one  ∈ {1, . . . , } such
 [ ∈  |  = ] = [ ∈  |  = ] (1) that
for  ∈ {0, 1} and all measurable subsets  of X.</p>
      <p>Denoting  [ = 1] =  and [ = 1] = , (1)
implies that the distribution of the features  under 
and  respectively can be represented as
 [ ∈  ] =</p>
      <p>(2a)
  [ ∈  |  = 1] + (1 − )  [ ∈  |  = 0],</p>
      <p>(2b)
[ ∈  ] =
  [ ∈  |  = 1] + (1 − )  [ ∈  |  = 0],</p>
      <p>
        Under (6),  ↦→ log () is strictly concave in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ].
      </p>
      <p>
        Hence (see Example 4.3.1 of [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]) the equation
1( ) ̸= 0( ).
 log  () = 0
 
(6)
(7)
has a solution 0 &lt;  &lt; 1 if and only if
and
&gt; 1.
      </p>
      <p>
        (8a)
for  ⊂ X. In the following, we assume that the com- This solution is then the unique point in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] where
ponents ,  [ ∈  |  = 1] and  [ ∈  |  = 0] () takes its absolute maximum value. By strict
concan be perfectly estimated from the training sample cavity of log , if (8a) is not true then either
(1, 1), . . . , (, ).
is iBnafesaicsaiblllye,.thTishemaesasnusmleptttiinogn hel=ps∞,howwheicvhero,btvoiosuhseldy 1 ∑︁ 1() ≤ 1 and 1 ∑︁ 0() &gt; 1, (8b)
=1 0()
      </p>
      <p>
        =1 1()
applies or
and
[
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] lies at  = 0 while under (8c), the unique maximum
of log  in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] is taken at  = 1.
      </p>
      <p>
        In summary, under the natural assumption (6), the
likelihood function  of (4) has an absolute maximum
in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] at a unique point * ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]. As a consequence,
the maximum likelihood (ML) estimate ˆ of the test
set positive class prevalence  is well-defined by setting
ˆ = * .
      </p>
    </sec>
    <sec id="sec-3">
      <title>5. The Cramér-Rao bound for unbiased estimators</title>
      <p>function  : X
In the setting of Section 3, let ̃︀ be any unbiased
estimator of the positive class prevalence  under the test
distribution , i.e. ̃︀ = (1, . . . , ) for some</p>
      <p>→ R such that
[̃︀] = [ = 1] = ,
in the sense that √ (ˆ −</p>
      <p>Actually, since ˆ from Section 4 is the ML
estimator of , (2b) presents not only a lower
bound for the</p>
      <p>variances of unbiased estimators
of  but also the large sample variance of ˆ
︃(</p>
      <p>) converges in
dis︂[ (︁ 1()− 0() )︁ 2]︂ − 1)︃</p>
      <p>()()
tribution toward</p>
      <p>0, 

︂[ (︁ 1()− 0() )︁ 2]︂ − 1</p>
      <p>()()
the normal distribution with mean 0 and variance</p>
      <p>
        , the asymptotic variance of ˆ
conditional distributions of the feature vector , such
that the density  () of  under  is given by (3).
tion of the feature vector under the test distribution ,
then the variance of ̃︀ is bounded from below by the
inverse of the product of the Fisher information of the test
distribution with respect to  and the size of the sample
(Cramér-Rao bound, see, e.g., Corollary 7.3.10 of [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]):
var[̃︀] ≥
=
 
 
1
1
︂[ (︁  log () ()
      </p>
      <p>︁) 2]︂
︂[ (︁ 1()− 0() )︁ 2]︂</p>
      <p>()()
that there are positive densities 0 and 1 of the class- in terms of the variance of  .
ture probability measure  of (2b). Assume additionally
where  denotes the expected value under the mix- result on the representation of the Fisher information
mentioned above in the context of the Cramér-Rao bound
If (1, . . . , ) is an i.i.d. sample from the distribu- a density as specified by
(3) then the Fisher information
.</p>
      <p>(9) ˆ:
,</p>
      <p>):
bound of (9) may not hold for ˆ because it need not be</p>
      <p>
        In summary, if the densities of the class-conditional
feature distributions are known, the ML estimator of 
has asymptotically the smallest variance of all unbiased
estimators of  on i.i.d. samples from the test distribution
of the features. This is demonstrated in the two upper
panels of Table 4 of [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which shows on simulated data
that confidence intervals for  – which are primarily
driven by the standard deviations of the estimators –
based on the ML estimator are the shortest if the training
sample is infinite and the test sample is large.
      </p>
    </sec>
    <sec id="sec-4">
      <title>6. The asymptotic variance of the</title>
    </sec>
    <sec id="sec-5">
      <title>ML estimator</title>
      <p>Denote by  () the posterior positive class probability
given  =  under . Assume that the feature vector
 under  has a density that is given by (3). Then it
holds that
 () =
 1()
 ()()
.</p>
      <p>(10)
This representation immediately implies the following
Proposition 6.1. If the feature vector  under  has

︂[ (︁  log ()</p>
      <p>︁) 2]︂
with respect to  can be represented as follows:
()</p>
      <p>of the distribution of  under 

[︃
︂(  log  ()</p>
      <p>︂) 2]︃
()
=
var[ ()]
2 (1 − )2 .</p>
      <p>
        From Proposition 6.1 we obtain the following
representation of the asymptotic variance of the ML estimator
︂[ (︁ 1() − 0() )︁ 2]︂ − 1
In (12), the optimal Brier Score () is also called
refinement loss , while var[ ] and var[ ()] are
known as uncertainty and resolution respectively [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>By (11), the asymptotic variance of the ML estimator
ˆ is reduced if a feature vector  with greater variance
of  () is found (or if the Brier Score with respect to
 decreases). Consider, for instance, a feature vector ′
that is a function of the feature vector , i.e. it holds that
′ =  () for some function  : X → X′. Since 
potentially maps diferent values  ∈ X onto the same
value ′ ∈ X′ the amount of information carried by ′ is
reduced compared to the amount of information carried
by . Therefore, the approximation of the class label 
by regression on ′ is less close than the approximation
of  by regression on .</p>
      <p>From this observation, it follows that (′) ≥
(). This in turn implies by (12) and (11) for the
asymptotic variances of the ML estimators ˆ(′) and
ˆ() that
var[ˆ(′)] ≤
var[ˆ()].</p>
      <p>(13)
Observe that ′ =  () also implies  (′) ≥
 (), i.e. also under the training distribution  , the
posterior positive class probability   () based on  is
a better predictor of  than   (′) which is based on ′.
By the assumption underlying this paper,  (′) and
 () are observable while (′) and ()
are not, because the class label  is not observed in the
test data set.
(′) ≥</p>
      <p>But does  (′) ≥</p>
      <p>() always imply
() and therefore also (13), thus
gen ().
eralising the implication “′ =  () ⇒ (13)”?</p>
      <p>
        This paper has no fully general answer to this
question. Instead we can only point to alternative conditions
on ′ and  that imply both  (′) ≥
and (′) ≥ (), but are weaker than ′ =
Brier curves: Recall the notion of Brier curve from [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
(with the slightly modified definition of [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]). If the
Brier curve for   (′) dominates the Brier curve
for   (), then by item 6) of Proposition 5.2 and
Proposition 4.1 of [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], it follows that  (′) ≥
 () and (′) ≥ () hold.
      </p>
      <p>ROC analysis: Recall the notion of Receiver Operating
(′) ≥ () hold.</p>
      <p>
        Characteristic (ROC) as defined, for instance, in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>
        If the ROC for the density ratio associated with 
dominates the ROC for the density ratio associated
with ′, then by Remark 5.4 and Proposition 4.1
of [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], it follows that  (′) ≥  () and
1–8
(16)
(17)
following explicit formula:
 () in Section 3 above, the estimator ˆ, is given by the
      </p>
      <p>
        By Eq. (2) of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], in the setting of this paper as specified
      </p>
    </sec>
    <sec id="sec-6">
      <title>7. Example: Binormal model</title>
      <p>
        In this section, we numerically compare the variance of
the Sample Mean Matching (SMM) estimator ˆ, of
class prevalences [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and the Cramér-Rao bound of (9)
(which is also the large sample variance of the ML
estimator ˆ, as specified in Section 4 above). In order to be
able to do this, we take recourse to the univariate
binormal model with equal variances of the class-conditional
distributions: The two normal class-conditional
distributions of the feature variable  are given by
 |  =  ∼ 
( ,  2),
 = 0, 1,
      </p>
      <p>(14a)
for conditional means  0 &lt;  1 and some  &gt;</p>
      <p>0. For the
sake of simplicity, we choose
 0 = 0,  1 &gt; 0,  = 1.</p>
      <p>(14b)
Greater values of  1 imply less overlap of the
classconditional feature distributions, corresponding to more
powerful (or accurate) models. Or in other words, for
greater values of  1, the feature variable  carries more
information on the class label  .</p>
      <p>
        As stated in Section 3, we assume we are dealing with
an infinitely large training sample and a test sample of
size . By Section 6, then for large  the variance of
ˆ, is approximately
1


[︃
︂( 1() − 0() )︂ 2]︃− 1
refer to Theorem 3 of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], case  = ∞ in the notation
of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Observe that in the case of SMM and  = ∞, it
holds that the representation of the variance is exact, not
only approximate. Hence we obtain
var[ˆ, ] =
=
 2 +  (1 − ) ( 02 +  12)
      </p>
      <p>( 1 −  0)2
1 (︂</p>
      <p>1
  21 +  (1 − ) .</p>
      <p>
        ︂)
In the model specified by (14a) and (14b), the
classification power (or accuracy) is driven by the diference of
the conditional means, i.e. by the mean conditional on
the positive class  1. If we measure the power by the
Area under the Curve (AUC, see for instance Section 6.1
of [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]) in order to obtain a measure which is
independent of the class prevalences, AUC is a simple function
of  1 =  1 −  0:
 ( 1) = Φ
      </p>
      <p>(18)
︂(  1 ︂)
√2</p>
      <p>,
and
√︁
√︁
  =</p>
      <p>var[ˆ, ]
  =</p>
      <p>var[ˆ,].
with Φ denoting the standard normal distribution
function. 8. Conclusions</p>
      <p>From (18) und (17), it is clear that for fixed test sample
size  the variance of ˆ, decreases when the power In this paper, we have revisited the binary quantification
of the model increases. This is less obvious from (15) for problem, i.e. the problem of estimating a binary prior
the asymptotic variance of ˆ, but if follows from (11) class distribution on the test data set when training and
in that case. test distributions are diferent.</p>
      <p>Table 1 above illustrates these observations. In the
table, we use the notation
• Specifically, under the assumption of prior
probability shift we have looked at the asymptotic variance
of the maximum likelihood estimator (MLE).
• We have found that this asymptotic variance is
closely related to the Brier score for the regression
of the class label variable  against the features
vector  under the test set distribution. In
particular, the asymptotic variance can be reduced by
selection of a more powerful feature vector.
• At the end of Section 6, we have pointed out
sufifcient conditions and associated training criteria
(Brier curves and ROC analysis) for minimising
both the Brier score on the training data set and
the Brier score on the test data set.
• These findings suggest methods to reduce the
variance of the ML estimator of the prior class
probabilities (or prevalences) on the test data set. Due to
the statistical consistency of ML estimators, by
reducing the variance of the estimator also its mean
squared error is minimised.
• The large sample variance of the MLE associated
with its asymptotic variance is identical to the
Cramér-Rao lower bound for the variances of
unbiased estimators of the prior positive class
probability. Therefore, it seems likely that also other
estimators benefit from improved performance of
the underlying classifiers or feature vectors.</p>
      <p>
        Indeed, results of a simulation study in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and theoretical
ifndings in [
        <xref ref-type="bibr" rid="ref26 ref7">7, 26</xref>
        ] suggest that improving the accuracy
of the base classifiers used for quantification helps to
reduce not only the variances of ML estimators but also
of other estimators. The example of the Sample Mean
Match (SMM) estimator we have discussed in Section 7
supports this conclusion.
      </p>
      <p>
        However, these findings must be qualified in so far
as the observations made in this paper apply only to the
case where both training data set and test data set are
large. This is a severe restriction indeed as [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] pointed
out the importance of quantification methods specifically
in the case of small training data sets, for cost eficiency
reasons.
      </p>
      <p>
        Further research on developing eficient quantifiers
for small or moderate training and test data set sizes
therefore is highly desirable. A promising step in this
direction has already been done by [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], with a proposal for
the selection of the most suitable quantifiers for problems
on data sets with widely varying sizes.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>The author thanks four reviewers for their useful and
supportive comments.
1–8</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Castaño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chawla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Coz</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          <article-title>Review on Quantification Learning</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>50</volume>
          (
          <year>2017</year>
          )
          <volume>74</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>74</lpage>
          :
          <fpage>40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maletzke</surname>
          </string-name>
          , G. Batista,
          <article-title>Accurately Quantifying a Billion Instances per Second</article-title>
          ,
          <source>in: 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . doi:
          <volume>10</volume>
          .1109/DSAA49011.
          <year>2020</year>
          .
          <volume>00012</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <article-title>QuaPy: A PythonBased Framework for Quantification</article-title>
          ,
          <source>arXiv preprint arXiv:2106.11057</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Forman</surname>
          </string-name>
          , Counting Positives Accurately Despite Inaccurate Classification,
          <source>in: European Conference on Machine Learning (ECML</source>
          <year>2005</year>
          ), Springer,
          <year>2005</year>
          , pp.
          <fpage>564</fpage>
          -
          <lpage>575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Forman</surname>
          </string-name>
          ,
          <article-title>Quantifying counts and costs via classification</article-title>
          ,
          <source>Data Mining and Knowledge Discovery</source>
          <volume>17</volume>
          (
          <year>2008</year>
          )
          <fpage>164</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Esuli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abbasi</surname>
          </string-name>
          ,
          <article-title>Sentiment quantification</article-title>
          ,
          <source>IEEE intelligent systems 25</source>
          (
          <year>2010</year>
          )
          <fpage>72</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Izbicki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stern</surname>
          </string-name>
          ,
          <article-title>Prior Shift Using the Ratio Estimator</article-title>
          , in: A.
          <string-name>
            <surname>Polpo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Stern</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Louzada</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Izbicki</surname>
          </string-name>
          , H. Takada (Eds.),
          <source>International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Engineering</source>
          , Springer,
          <year>2017</year>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tasche</surname>
          </string-name>
          ,
          <article-title>Confidence intervals for class prevalences under prior probability shift</article-title>
          ,
          <source>Machine Learning and Knowledge Extraction</source>
          <volume>1</volume>
          (
          <year>2019</year>
          )
          <fpage>805</fpage>
          -
          <lpage>831</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Alexandari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kundaje</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shrikumar</surname>
          </string-name>
          ,
          <article-title>Maximum Likelihood with Bias-Corrected Calibration is HardTo-Beat at Label Shift Adaptation</article-title>
          , in: International Conference on Machine Learning, PMLR,
          <year>2020</year>
          , pp.
          <fpage>222</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Garg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lipton</surname>
          </string-name>
          ,
          <article-title>A Unified View of Label Shift Estimation</article-title>
          , in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>M. F.</given-names>
          </string-name>
          <string-name>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>3290</fpage>
          -
          <lpage>3300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Moreno-Torres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Raeder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alaiz-Rodriguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chawla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <article-title>A unifying view on dataset shift in classification</article-title>
          ,
          <source>Pattern Recognition</source>
          <volume>45</volume>
          (
          <year>2012</year>
          )
          <fpage>521</fpage>
          -
          <lpage>530</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hofer</surname>
          </string-name>
          , G. Krempl,
          <article-title>Drift mining in data: A framework for addressing drift in classification</article-title>
          ,
          <source>Computational Statistics &amp; Data Analysis</source>
          <volume>57</volume>
          (
          <year>2013</year>
          )
          <fpage>377</fpage>
          -
          <lpage>391</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lipton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smola</surname>
          </string-name>
          ,
          <article-title>Detecting and Correcting for Label Shift with Black Box Predictors</article-title>
          , in: J.
          <string-name>
            <surname>Dy</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Krause (Eds.),
          <source>Proceedings of the 35th International Conference on Machine Learning</source>
          , volume
          <volume>80</volume>
          <source>of Proceedings of Machine Learning Research, PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>3122</fpage>
          -
          <lpage>3130</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Frühwirth-Schnatter</surname>
          </string-name>
          ,
          <article-title>Finite Mixture and Markov Switching Models: Modeling and</article-title>
          Applications to Random Processes, Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Coberly</surname>
          </string-name>
          ,
          <article-title>The numerical evaluation of the maximum-likelihood estimate of mixture proportions</article-title>
          ,
          <source>Communications in Statistics - Theory and Methods</source>
          <volume>5</volume>
          (
          <year>1976</year>
          )
          <fpage>1127</fpage>
          -
          <lpage>1135</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Saerens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Latinne</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Decaestecker, Adjusting the Outputs of a Classifier to New a Priori Probabilities: A Simple Procedure</article-title>
          ,
          <source>Neural Computation</source>
          <volume>14</volume>
          (
          <year>2001</year>
          )
          <fpage>21</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tasche</surname>
          </string-name>
          ,
          <article-title>Fisher Consistency for Prior Probability Shift</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>18</volume>
          (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          . URL: http://jmlr.org/papers/v18/
          <fpage>17</fpage>
          -
          <lpage>048</lpage>
          . html.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>M.-J. Zhao</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Edakunni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Pocock</surname>
          </string-name>
          , G. Brown,
          <article-title>Beyond Fano's Inequality: Bounds on the Optimal F-Score, BER, and Cost-Sensitive Risk</article-title>
          and
          <string-name>
            <given-names>Their</given-names>
            <surname>Implications</surname>
          </string-name>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>14</volume>
          (
          <year>2013</year>
          )
          <fpage>1033</fpage>
          -
          <lpage>1090</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Titterington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Makov</surname>
          </string-name>
          ,
          <article-title>Statistical analysis of finite mixture distributions</article-title>
          , Wiley New York,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Casella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Berger</surname>
          </string-name>
          , Statistical Inference, second ed., Duxbury Press,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hand</surname>
          </string-name>
          ,
          <article-title>Construction and Assessment of Classification Rules</article-title>
          , John Wiley &amp; Sons, Chichester,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hernández-Orallo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Flach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ferri</surname>
          </string-name>
          , Brier Curves:
          <article-title>A New Cost-Based Visualisation of Classifier Performance</article-title>
          ,
          <source>in: Proceedings of the 28th International Conference on Machine Learning (ICML</source>
          <year>2011</year>
          ),
          <source>International Machine Learning Society</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>585</fpage>
          -
          <lpage>592</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tasche</surname>
          </string-name>
          , Calibrating suficiently,
          <source>arXiv preprint arXiv:2105.07283</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Reid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Williamson</surname>
          </string-name>
          , Information, Divergence and
          <article-title>Risk for Binary Experiments</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <year>2011</year>
          )
          <fpage>731</fpage>
          -
          <lpage>817</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R Core</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <string-name>
            <surname>R:</surname>
          </string-name>
          <article-title>A Language and Environment for Statistical Computing, R Foundation for Statistical Computing</article-title>
          , Vienna, Austria,
          <year>2019</year>
          . URL: https:// www.R-project.
          <source>org/.</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Izbicki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stern</surname>
          </string-name>
          ,
          <article-title>Quantification Under Prior Probability Shift: the Ratio Estimator and its Extensions</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>20</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          . URL: http://jmlr.org/papers/ v20/
          <fpage>18</fpage>
          -
          <lpage>456</lpage>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maletzke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hassan</surname>
          </string-name>
          , D. dos Reis, G. Batista,
          <article-title>The Importance of the Test Set Size in Quantification Assessment</article-title>
          ,
          <source>in: Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2640</fpage>
          -
          <lpage>2646</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>