<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Pitfalls in Quantification Assessment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Waqar Hassan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>André Maletzke</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gustavo Batista</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quantifier Q</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Engineering and Exact Sciences, Western Paraná State University</institution>
          ,
          <addr-line>Foz do Iguaçu, PR</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Mathematics and Computer Science, University of São Paulo</institution>
          ,
          <addr-line>São Carlos, SP</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Science and Engineering, University of New South Wales</institution>
          ,
          <addr-line>Sydney, NSW</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Testing sets with natural</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>5</volume>
      <issue>2021</issue>
      <abstract>
        <p>Quantification is a research area that develops methods that estimate the class attribute prevalence in an independent sample. Like the other fields in Machine Learning, quantification researchers often use experimental assessment to evaluate and compare the performance of their proposals. Therefore, the design of the experimental protocols is critical to the research in quantification. Currently, two protocols have dominated the assessment of quantifiers: the artificial-prevalence protocol (APP) and the natural-prevalence protocol (NPP). APP is the most employed since it allows the use of classification datasets, plenty available. However, APP has a shortcoming: the synthetic class prevalence of the test sets. This paper discusses the practical consequences of this shortcoming and shows simple examples of quantifiers that exploit these limitations to improve their performance artificially. We propose a baseline quantifier, lazy, and radar charts as tools to identify situations where the proposed quantifiers are performing poorly.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Machine learning</kwd>
        <kwd>Quantification</kwd>
        <kwd>Experimental Protocol</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Quantification is the research area that develops methods</title>
        <p>to estimate the class prevalence in a data sample. In
the last decade, we have witnessed a rapid evolution of
the field supported by a thriving research community
and significant developments in theoretical and practical
aspects of quantification.</p>
        <p>Similarly to other areas of Machine Learning,
quantification researchers often avail of experimental
assessments to demonstrate the eficacy of their proposals.</p>
        <p>Therefore, experimental setups are a critical part of
quantification methods development. They support a fair
comparison of algorithms allowing the researchers to assess
research progress and guide future investigation eforts.</p>
        <p>
          Several researchers have investigated how the
experimental design decisions may influence the performance of
quantifiers. Recent examples are the study of the impact
of classifier hyper-parameters [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and test set size [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
in quantification accuracy and the investigation of the
properties of performance measures for quantification
assessment [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          Quantification research has employed two main
experimental setups in assessing their proposals:
artificialprevalence protocol (APP) and natural-prevalence protocol
(NPP) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. APP is the most common one since it allows
the assessment of quantification methods using
classifi
        </p>
        <p>Error</p>
        <p>Q</p>
        <p>Error
Testing sets with different
class distribution</p>
        <p>Testing sets with different
class distribution
Samples
randomly
extracted from
the test set
varying class
distribution
...</p>
        <p>Training
set
Quantifier</p>
        <p>Q
Evaluation
Evaluation
0.5
P(⊕)
1
0.5
P(⊕)</p>
        <p>1
cation datasets including benchmark data that are plenty
available online.</p>
        <p>In summary, APP consists of splitting a classification
dataset into training and test sets. The test set class
distribution is artificially manipulated through sub-sampling,
creating multiple test set samples. A typical design
decision is to generate test samples with class prevalence
spanning across the whole spectrum of class distributions.</p>
        <p>The same number of test samples is often generated for
each class distribution, creating a uniform distribution
of class prevalence across all test samples. Therefore, the synthetic characteristics to inflate their performance.
quantification methods are assessed for a whole spec- It is not the objective of this paper to reject APP or
trum of class distributions, being each class distribution to propose a diferent experimental protocol. It is our
equally probable. Fig. 1-left illustrates the APP setup. opinion that APP is currently the best protocol to assess</p>
        <p>
          In contrast, NPP requires specialised quantification quantifiers due to the lack of NPP datasets. However,
datasets. Those datasets come with multiple naturally as APP employs partially natural and partially synthetic
occurring test samples. For example, in an insect surveil- data distributions, the research community must know
lance problem [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], we may have several insect traps in- the APP limitations.
stalled in the field capturing diferent amounts of insects. We conclude with a few recommendations when
imThe final objective is to count the number of captured plementing quantifiers, particularly of the distribution
insects by species using quantification methods. Each matching category. We also recommend the adoption of
test sample is likely to have a diferent class distribution. a new baseline quantifier for APP named lazy. Lazy is a
Fig. 1-right illustrates the NPP setup. quantifier that constantly predicts the expected positive
        </p>
        <p>APP has an artificially controlled component: the class class prevalence across all test sets. Lazy has a similar
distribution of the test sets. For simplicity, virtually ev- semantic as the majority-class classifier, often used as a
ery paper in quantification that uses APP has opted for baseline in classification assessment.
a uniform distribution of class prevalence values across This paper is organised as follows: Section 2 discusses
the test sets. Thus, we can expect experimental setups to the artificial-prevalence protocol (APP) and Section 3
have the same amount of test sets for each class preva- reviews the natural-prevalence protocol NPP. Section 4
lence. However, for real problems, such uniform distri- presents the related work, including a brief review of the
bution rarely holds. In the insect example, we can expect quantifiers employed in this paper. Section 5 discusses
some species are less prevalent than others, and the class two limitations of APP and examples of simple
quantidistribution across all test sets is unlikely to be uniform. fiers that inappropriately exploit them to improve their</p>
        <p>
          Such artificial characteristic introduces a considerable performance. Section 6 brings a discussion about the
risk of biasing the experimental results. For instance, we APP limitations and recommend the use of lazy as a
baseoften summarise the quantifiers performance across all line quantifier and radar charts as a tool to inspect the
test sets with a statistic like the mean error to facilitate a performance of quantifiers for all possible class
distribuperformance comparison. Such numbers are then com- tions. Finally, Section 7 concludes our work and presents
pared directly or through a hypothesis test. Therefore, directions for future work.
the best performing method in an APP experiment will
have the best average performance considering all class
distributions equally likely, and thus evenly important. 2. Artificial-prevalence Protocol
However, such average performance may not be a
relevant criterion for a particular problem. The practice The artificial-prevalence protocol (APP), proposed by
may demand approaches that best perform non-uniform Forman [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], is the most used experimental setup to assess
distributions, such as skewed class distributions. and compare quantification methods.
        </p>
        <p>The widespread use of APP may lead the community For a binary classification dataset with a set of classes
to focus on methods that perform well on average across  = {⊕ , ⊖} , APP creates multiple test sets through the
all class distributions. Through “evolutionary” research application of sub-sampling. It means that APP randomly
pressure, researchers are likely to propose variations and removes examples from either ⊕ or ⊖ to generate test
improvements of the best-performing methods. One may sets with predetermined class distributions. Typically,
say that researchers can adapt APP to generate diferent the experimental design uses class distributions across
class distributions across all test sets besides the uniform the entire spectrum of possibilities, such as  =  (⊕ ) ∈
one. However, such a choice of distribution is arbitrary {0, .01, .02, . . . , .99, 1}.
since we do not expect to have a dominant class distribu- APP involves a stochastic decision and therefore is
pastion over a variety of real-world datasets. sive of variability due to chance. Thus, most researchers</p>
        <p>
          This paper examines the two main experimental de- prefer to repeat this experiment to decrease variance.
Resigns used in quantification, APP and NPP. We discuss searchers often generate between 10 and 100 samples for
some potential limitations of APP related to the creation each class distribution. This procedure leads to an
assessof multiple test sets with artificial class distributions. We ment over a large number of test sets. For instance, 
address two possible weaknesses of the APP protocol: with increments of .01 in the range [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] and 10 samples
the uniform class distribution across all test sets and the per value of  leads to 1010 (101 × 10) test samples.
discrete nature of the class distributions. To make our We can now perceive the two main limitations of APP.
contributions concrete, we show two simple scenarios The first is the discrete nature of the  values.
Experiin which poorly designed quantifiers use these two APP mental designs generate a class distribution with fixed
increments and for all values in the range. The second is expected value under certain circumstances. A simple
the uniform distribution across all test samples. Looking way to think about this is to realise that .5 is the middle
at  as a random variable,  () has a uniform distri- value in the range [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ], and it is the constant prediction
bution. APP evaluations use classification datasets, and that minimises the error under a uniform distribution of
there is no natural distribution for the test sets. Therefore, classes in the test sets.
most researchers decide to use a uniform distribution and Such observation is also the motivation for
incorporatgenerate the same number of test sets for each value of ing a simple baseline quantifier that we call lazy. The lazy
. quantifier predicts E[] for every test set independently
        </p>
        <p>
          Quantification methods inappropriately make use of of its actual class distribution.
both limitations, giving them a significant advantage Section 5.2 shows a simple example in which we
proover other methods. A simple example occur with quan- pose a quantifier that applies Laplace smoothing to the
tification methods that explicitly search over the space output of the Classify &amp; Count (CC) method (cf.
Secof possible class distributions, such as HDy [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and the tion 4.2). Such a new quantifier appears to outperform
methods of the DyS family [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. We show in Section 5.1 CC; however, its performance improvement is merely an
that by leaking the actual values of  and searching di- artifact of the experimental design.
rectly over these values, HDy can artificially boost its
performance.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. The Natural-prevalence</title>
    </sec>
    <sec id="sec-3">
      <title>Protocol</title>
      <sec id="sec-3-1">
        <title>The natural-prevalence protocol (NPP) requires datasets</title>
        <p>Evaluation Measure Definition in which multiple test samples occur naturally. These
Absolute Error (AE) |1| ∑︀∈ | ^() − () | test samples are often a result of data collection
occurNormalized Absolute Error (NAE) ∑2︀(1∈−m|i^n()− (()))| trhinegi ninsevcatrsiouurvseliollcaanticoenasp,ppleirciaotdiosnorrebqoutihr.esFothreindsetapnlocye-,
Relative Absolute Error (RAE) |1| ∑︀∈ |^()−()()| ment of insect traps in diferent locations to estimate
(NNoRrmAEa)lized Relative Absolute Error |∑|−︀1∈+1−|m^m(ini)n−()(()|()) itnhteersepsatt.iotemporal distribution of the insect species of
An essential characteristic of NPP datasets is the
presSquared Error (SE) |1| ∑︀∈ (^() − ())2 ence of a class distribution drift in the test samples.
OtherDiscordance Ratio (DR) |1| ∑︀∈ m|a^x(()(−),(^()|)) wise, we can trivially solve the problem by estimating the
Kullback-Leibler Divergence (KLD) ∑︀∈ () log ^(()) empirical class probability in the training set. Returning
to the insect surveillance example, we can expect that the
gNeonrcmea(lNizKedLD)Kullback-Leibler Diver- 2 ((,^,)^+)1 − 1 insect distribution will vary both spatially (for instance,
Pearson Divergence (PD) |1| ∑︀∈ (()^−(^)())2 adsuewtaotetrheanpdrefsoeondc)eaonfdfatevmouproarballelylo(fcoarl ceoxanmdiptiloen,dsuseuctho
seasonal factors).</p>
        <p>
          The second limitation is exploitable in a more subtle NPP datasets are rare because they require a
tremenway. Frequently, quantification papers report the average dous efort to collect and label the data, including the
performance across all test sets. Table 1 summarises the multiple test samples. One example of a dataset with
most used error measures in quantification research, and such characteristics is the Woods Hole Oceanographic
we point the interested readers to [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] for an insightful Institution (WHOI) Plankton dataset [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The dataset
analysis of these measures properties. All these measures consists of nine years of data collected by an automated
are pointwise, i.e., they take into consideration only two system that samples 5 ml of seawater every 20 minutes
values: the actual  and the predicted ˆ prevalences of resulting in nearly 1 billion images. Recently, González
the positive class in a single test sample. Therefore, re- et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] used the NPP setup to assess several
quantisearchers often average those numbers across all test ifers to a subset of this dataset consisting of 3.4 million
samples to provide a single number that summarises the annotated images organized in 964 samples.
quantifier performance for an entire dataset. NPP has none of the mentioned limitations of APP. The
        </p>
        <p>The main issue with this experimental setup is that class prevalence of the test set occur naturally and vary
 has a fixed and known expected value across all test from dataset to dataset. Therefore, NPP has no fixed
exsets. For instance, E[] = .5 if we generate the same pected value for , and designing a baseline method such
number of test sets for all possible class distributions as lazy would require a diferent value for each dataset.
with a constant increment. Therefore, a quantifier can We can safely conclude that the risks of overfitting the
appear more accurate if it predicts values closer to this test sets with NPP are much lower than with APP. In
particular, methods that explicitly search for ˆ cannot have
access to a fixed set of artificial  values. Also, avoiding
extreme ˆ predictions, such as 0 and 1, is less likely to
make the method appear more accurate.</p>
        <p>
          We could conclude that NPP is superior to APP, but a
scan of the literature shows that APP is the experimental
setup of choice of most of the quantification papers [
          <xref ref-type="bibr" rid="ref11 ref4 ref7 ref8">11, 4,
12, 8, 13, 7, 14, 15</xref>
          ]. The reason is the lack of NPP datasets
and the dificulties in creating them. A possible solution
is using large classification datasets that can be split into
various test sets. One example is the study in [16] that
assesses quantifiers in 5148 binary test sets created with
the RCV1-v2 dataset [17]. The authors created those test
sets by splitting the one-year worth of data in RCV1-v2
in 52 weeks. They also consider each of the 99 classes in
the data, in turn, leading to the 5148 test sets (52 × 99).
        </p>
        <p>However, even in this case, these datasets may not pose
a challenging problem for quantification due to the lack
of class distribution variability among test sets [18].</p>
        <p>
          The creation of NPP datasets may involve a
significant amount of work in gathering data from
quantification applications. Sentiment analysis is an example
of an application domain that has used quantification
frequently and a candidate to provide NPP benchmark
datasets. However, a relevant study performed in this
area using NPP setup [19] have been re-assessed with
APP recently [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The main reason is the reduced
number of test sets used in the experiments and the potential
risks of generalising those experimental results based on
limited evidence [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Therefore, we see a legitimate need to use APP datasets
due to the lack of NPP datasets. However, as we discuss
and show empirical evidence in this paper, researchers
must be mindful of the limitations of APP setup to avoid
proposing methods that take advance of the experimental
design to provide competitive performance.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Related Work</title>
      <sec id="sec-4-1">
        <title>This section provides a summary of the quantification algorithms included in our experiments. We start by providing some initial background and notation that will be useful to explain the quantification approaches.</title>
        <sec id="sec-4-1-1">
          <title>4.1. Background</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Quantification is a supervised machine learning task that</title>
        <p>shares similarities with classification. The main one is
the attribute-value representation for observation and
their relation to a nominal class attribute.</p>
        <p>Formally, in the case of binary classification, let  =
{(x1, 1), . . . , (x, )} be the labeled set, in which
each example x ∈  is a vector in the -dimensional
feature space  , and  ∈  = {⊕ , ⊖} is its
respective class label. In such a setting, we can define a binary
classifier as a model ℎ induced from  such that:
ℎ :  →− {⊕
, ⊖}</p>
      </sec>
      <sec id="sec-4-3">
        <title>The objective of ℎ is to predict the class label of pre</title>
        <p>viously unseen examples accurately. In this scenario, a
scorer ℎ is defined such that:
ℎ :  →−</p>
        <p>R
which aims to predict, for each example, a numerical
value that correlates to P( = ⊕| x), that is, the
posterior probability of an example belonging to the positive
class. To simplify the remaining of the text, we refer to
scores of positive examples as positive scores and scores
of negative examples as negative scores. We can
conveniently convert a scorer into a classifier by setting a
threshold: only examples scored above the threshold are
classified as positive.</p>
        <p>
          In binary quantification, we are not interested in
individual class labels. Instead, we predict the proportion of
the positive examples in a test sample. A quantification
model ℎ is, therefore, defined as follows:
ℎ : S →−
[
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]
where S denotes the universe of possible samples from
 and ℎ estimates the prevalence of the positive class
in a given sample  ∈ S .
        </p>
        <p>
          In the last decade, we have witnessed the proposal of
several quantification algorithms. Although they share
the same objectives, their introduction by diferent
communities led to diferent names for the quantification
task, such as prevalence estimation [20], class prior
estimation [21], and class distribution estimation [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We
recommend the survey of Gonzalez et al. [22] as an
organised and comprehensive review of the most relevant
quantification methods proposed in the literature. In
what follows, we provide a summary of the
quantification algorithm included in our experiments.
        </p>
        <sec id="sec-4-3-1">
          <title>4.2. Quantification Methods</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>The most straightforward quantification approach is</title>
        <p>Classify &amp; Count (CC). It is a naive adaptation of
classiifers to quantification problems. Forman [ 15] has
demonstrated CC has a systematic error that monotonically
increases as we move away from a distribution that CC
provides optimal counting.</p>
        <p>CC uses a classifier to label each instance in the test
sample. Afterwards, it counts the number of examples
belonging to each class. CC provides optimal quantification
results with a perfect classifier. However, classifiers with
balanced errors, such as a binary classifier that commits
an equal number of false-positive and false-negative
errors, are also optimal. Intuitively, in these situations, CC
benefits from the fact that opposite mistakes can nullify
each other.</p>
        <p>Negative distribution
Positive distribution</p>
        <p>Threshold
False positive</p>
        <p>False negative</p>
        <p>Fig. 2 illustrates the probability density functions of the
scores for a binary classification problem for a
hypothetical test set. We chose the threshold so that the number
of false positives matches the false negatives. Therefore,
the CC quantifier provides perfect quantification, albeit
the underlying classifier is not perfect.</p>
        <p>The CC outcome is the count of every observation
with a score above the threshold. In other words, it is
the sum of true positives and false positives. However,
the actual count is the sum of true positives and false
negatives. Fig. 2 helps us to understand the motivation
behind several quantification methods that correct the
counts by estimating false-positive and false-negative
errors.</p>
        <p>
          A well-known approach of this approach is Adjusted
Classify &amp; Count (ACC) [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In absolute numbers, ACC’s
correction factor adds the false negatives to CC’s output
and then subtracts the false positives. However, ACC is
more commonly expressed as frequencies, in the
following manner:
a search mechanism to find the parameters that best
match a mixture of positive and negative training set
score distributions with the unlabelled score distribution
of the test set. The computation of the parameters of this
mixture leads to the quantification estimate.
        </p>
        <p>
          The HDy algorithm [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] represents each score
distribution as a histogram. A weighted sum of these
histograms gives the mixture between the positive and
negative score distributions, where the weights sum up to 1.
The weights that minimize the Hellinger Distance (HD)
between the mixture and the unlabeled (test) score
distribution (⊙ ) are considered to be the proportion of the
corresponding classes in the unlabeled sample. The next
equation details this computation:
ˆ
 HDy(⊕ ) =
arg min {︀ HD (︀  [⊕ ] + (1 −  )[⊖ ], [⊙ ])︀}
0≤  ≤ 1
where HD represents Hellinger distance and [· ]
indicates an operation that converts a set of scores into a
histogram. Fig. 3 illustrates this process.
        </p>
        <p>HD
]α)</p>
        <p>S⊕
+ (1 - α)</p>
        <p>S⊖
)
(2)
]</p>
      </sec>
      <sec id="sec-4-5">
        <title>HDy uses histograms to represent the positive, neg</title>
        <p>
          ative and unlabelled score distributions. A histogram is
ˆ  (⊕ ) = ˆ(⊕|⊕  (⊕ )) − − ((⊕ |⊕|⊖ ⊖ )) (1) tahdeisncurmetbeerreopfrebsinens1t.aHtioDnythauatthhoarss arerceolemvmanetnpdaarpapmlyeitnegr,
the method over a range of bins from 10 to 110 with an
wabnyhdIeCfrwCe(eˆi⊕|n⊕knthee()w⊕itset)hstithessetterhtture.uepe-po-(pos⊕|⊖iostisitviitevi)vecielsaarnstahsdteepf.rfaealvslseae-le-ppnoocsesitiptivirveoevrriadateteeds, iensctTirmhemeatoeernditgpionofsai1lt0iHv.eDTdyhiesmtfinreiabthluootiduotnupssueatsciarsoltsihnseeaamlrl besedinaiasrcnvhaoltfuotefinhsd.e
in the test set, then ACC would be a perfect quantifier. the alpha that minimises the Hellinger distance. Some
However, as the test set is unlabelled, the best we can minor improvements to HDy are the use of ternary
do is to estimate these quantities in a validation set. Es- search to make HDy more eficient and the use of Laplace
timating these values in the validation set often makes smoothing [25] to compute the bin values [26].
ACC far from perfect and not as accurate as the state-of- The HDy method inspired a recently proposed
framethe-art [12]. work named DyS [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] that supports the use of diferent
        </p>
        <p>Another class of quantification methods is known as distance measures besides HD:
distribution matching. These methods include algorithms
that mixture distributions on the training set to match
the test set distribution. A practical way of matching ˆ DyS(⊕ ) =
distributions is to consider the scores obtained on an arg min {︀ DS (︀  [⊕ ] + (1 −  )[⊖ ], [⊙ ])︀}
unlabeled set follow a parametric mixture between two 0≤  ≤ 1
known distributions (one for the positive and another
for the negative class). In general, these methods use
1Bins divide the entire range of score values into a series of
intervals, so we can count how many values fall into each interval.
where DS is a dissimilarity measure to estimate the match
between the distributions of training scores and test
scores, and [· ] is an operation that converts a set of
scores into a suitable representation for DS, such as a
histogram in the case of HD.</p>
        <p>
          The DyS approach with Topsøe distance is among of
the most accurate quantifiers in the literature [
          <xref ref-type="bibr" rid="ref11 ref8">11, 8</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Exploring APP</title>
      <p>This section discusses two limitations of APP: the
discrete nature (Section 5.1) and the uniform distribution
(Section 5.2) of the  values. We show simple examples
of how quantifiers can exploit these limitations leading
to inflated performance.
[27], OpenML [28], PROMISE [29], and Reis [26]
repositories. Table 2 briefly describes the main features of the
datasets.</p>
      <sec id="sec-5-1">
        <title>We control two main parameters: the</title>
        <p>5.1. Exploring the Discrete Nature of  c{a1r0d,i2n0a,li3ty0, 4o0f, 5 0th,e100s,e1ts50,200a, n2d50,Σ 3 0.0, 35|0,4|00, 4∈50,
APP requires the explicit specification of the test class 500} and |Σ | ∈ {10, 20, 30, 40, 50, 100, 150}. We use
prevalences that will be generated to assess the quanti- two state-of-the-art quantifiers, HDy and DyS. The
ifers. The common practice is to subsample the positive results are expressed as the mean absolute error (MAE)
or negative classes, producing class prevalences across (c.f. Table 1) across all datasets2.
the entire spectrum of possibilities.</p>
        <p>The nature of the generated artificial class prevalences HHDDyy−−1200 HHDDyy−−4300 HHDDyy−−15000 HDyDSy−−T15S0
is inherently discrete. Therefore, it is a researcher’s deci- 0.03
sion how many diferent test distributions they will
generate. For example, we can create test set distributions 0.02
with .1 increments such as  = {0, .1, . . . , .9, 1} or with AE
.01 increments such as  = {0, .01, . . . , .99, 1}. Apart 0M.01
from the obvious computational overhead caused by the
smaller increments, it is unclear how such a decision will 0.00
afect the assessment of quantifiers’ performance. 0 N50umb1e0r0of 1d5is0cre2t0e0pos2i5t0ive c3l0a0ss d35is0trib4u0t0ions45|A0| 500</p>
        <p>A possible issue with APP discrete distributions occurs
when this design information is inadvertently leaked to
the quantification method. To illustrate this problem, we Figure 4: Experimental results for HDy and DyS quantifiers
use HDy, a state-of-the-art quantifier. The use of HDy is as || and |Σ| vary.
motivated by this algorithm searching over the possible
class distributions, looking for the match that provides
the smallest Hellinger distance. The HDy is formalised in
Equation 2. We restate this equation making an explicit
statement of which values of  will be tested:
a very stable performance for all variations of HDy when the uniform distribution. In the case of binary
quantifi|| ≥ 100. Therefore, we have evidence to recommend cation, the estimates provided by LCC will be closer to
|| = 100 since greater values would require additional .5.
computational processing. We assessed CC and LCC, adding to the datasets of</p>
        <p>
          Recent publications with large experimental assess- Table 2 additional datasets listed in Table 3. Our
experments have used || ≪ 100, such as || = 12 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] and iments use test sets of size 100 and smoothing factor
|| = 21 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. These smaller values of || are necessary  = 10 and Random Forests with 200 trees as the
underto avoid a combinatorial explosion since these papers lying classifier.
also vary the class distribution in the training set. In the
particular case of these two papers, there is no leakage Table 3
of test information to the quantifier. Additional Datasets Description.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], the authors use convex optimisation as a search
procedure and a ternary search when the first procedure Dataset Size Features Repository
fails to find a response. Regarding [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the implementa- BCarendkitMCaarrkdet[i3n3g] [32] 3405,,020101 1263 UUCCII
tion in the QuaPy framework [31] uses |Σ | = 100 by HTRU2 17,898 8 UCI
default. In our experiments, |Σ | = 100 and || = 20 JM1 10,880 21 PROMISE
leads to slightly optimistic MAE estimates. We recom- Letter Recognition 20,000 16 UCI
mend that QuaPy replaces the search procedure of Distri- Pollen 3,848 6 OpenML
bution Matching algorithms with ternary [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] or convex Numerai 96,320 22 OpenML
optimisation procedure or a combination of both [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>Table 4 summarises the results, showing the mean
5.2. Exploring the Uniform Distribution absolute error (MAE) of each quantifier (CC and LCC)
and the − value estimated according to the Friedman
Let us suppose we want to create a new quantifier based test with 95% confidence. The best quantification result
on the well-known Classify and Count (CC) approach. for each dataset is shown in bold. Underlined values
CC is a quantifier that estimates the prevalence of the represent a statistical diference between CC and LCC.
positive class,  =  (⊕ ), by counting the output of a
classifier. Let us suppose that we want to improve the
empirical probabilities provided by CC using a “smoothed” Table 4
estimator such as Laplace smoothing [25]. We name this CC and LCC mean absolute error per dataset.
variation of CC as Laplace CC or simply LCC.</p>
        <p>Laplace smoothing is a technique often applied when
computing empirical probabilities from data. In the case
of binary quantification, the class estimate provided by
the Classify &amp; Count (CC) is given by
ˆ =
∑︀|=|1 1[ℎ(x) = ⊕ ]</p>
        <p>||
where  is a test sample and x ∈ .</p>
        <p>Laplace smoothing provides a smoothed estimator by
adding a pseudo-count  . In the case of ˆ given by CC,
we have</p>
        <p>Dataset
s AedesSex
tse AedesQuinx
taa Anuran Calls
lD ArabicDigit
iiagn EBENGGE(yvoetSet)ate
r
O Nomao
ts Bank Marketing
sea Credit Card
ta HRTU2
laD JM1
iitodn LPeotltleenr Recognition
d Numerai
A</p>
      </sec>
      <sec id="sec-5-2">
        <title>LCC seems to outperform CC for the additional</title>
        <p>where  ≥ 0 is the smoothing parameter and  is the datasets. In the datasets that LCC is better than CC, the
number of classes, i.e., two for binary classification. The performance of LCC is consistently superior to CC for
resulting estimate is between the empirical probability ˆ all possible values of . Figure 5 illustrates this
perforand the uniform probability 1/. mance improvement using radar plots for four datasets</p>
        <p>Although Laplace smoothing is a well-established tech- from Table 33. The radar plots use the angle to represent
nique, there is no reason to expect LCC to outperform the test set positive-class distribution  and the radius to
CC in any experimental setting. Laplace smoothing is a describe the mean absolute error.
technique that pushes the prevalence estimates towards</p>
      </sec>
      <sec id="sec-5-3">
        <title>3The results for all datasets are available on the paper website.</title>
        <p>LCC</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>CC</p>
      <p>0.0
0.5 0.9 0.1 0.6 0.9 0.1 The purpose of this paper is to discuss the APP
experi0.4 0.4 mental setup and some of its limitations. These
deficienAEM000...231 00..78 00..32 AEM0.2 00..78 00..32 ccireesataeretersetlasteetds twoitthheaswynidtheertaicndgaetoafdcislatrsisbduitsiotrnisbuutsieodnsto.</p>
      <p>Our objective is to illustrate with simple examples
0.6 0.5 0.4 0.6 0.5 0.4 how quantifiers can exploit APP limitations to improve
their performance artificially. These examples are
simHRTU2 JM1 plifications of some scenarios our research group has</p>
      <p>CC LCC CC LCC experienced while developing quantification methods.
0.12 0.9 0.0 0.1 0.6 0.9 0.0 0.1 Abescqoumaenstimficaotrieonchmaleltehnogdinsgintcoreidaesnetiinfystohpehsiestiiscsauteios.n, it
E00..0048 0.8 0.2 E00..24 0.8 0.2 We have spent countless hours looking for a theoretical
AM 0.7 0.3 AM 0.7 0.3 feoxrpmlasnaaqtiuoannotiffierwhy tao qfinudaonutitfietrhat
emappiprircoaxlilmyaotuestptehre0.6 0.5 0.4 0.6 0.5 0.4 lazy method under challenging situations, such as tiny
test set sizes.</p>
      <p>This situation becomes more dificult to diagnose given
Figure 5: Radial plots for CC and LCC for four datasets. the need to assess Machine Learning proposals in
several datasets. The more datasets we use in an empirical
evaluation, the less likely it becomes to look at individual</p>
      <p>We can notice in Table 4 that the additional datasets numbers. There is a need to summarise the results into a
often show larger quantification errors than the original relatively small amount of data to compare them more
datasets. Therefore, we could be tempted to conclude easily.
that Laplace smoothing is a technique that can help to im- By presenting and analysing summarised results, we
prove quantification performance for complex problems. are less prone to identify potential issues. We increase
However, this conclusion is incorrect. the chance of having a biased model being recognised as</p>
      <p>The reason for the performance enhancement is, in legit by authors and reviewers.
fact, an artifact of the experimental design. As the prob- In this sense, using the lazy baseline quantifier helps
lem becomes more challenging and the quantification identify the conditions the models are not performing
error increases, it pays of to forecast class prevalences well enough. Informally, the community has used CC as
closer to E[]. In our experiment, if we constantly predict a baseline classifier. However, we have seen situations
ˆ = .5 for all test sets, the expected MAE is just .25. where the performance of CC (and other quantifiers)</p>
      <p>We may argue that this is not a limitation of APP. In was much worse than lazy. These situations gave us the
other experimental setups, we could also have an ap- impression that we were making progress when we were
parent performance improvement by forecasting values far from a practical performance.
closer to E[] under certain circumstances, such as dif- A simple plot to compare the performance for multiple
ifcult datasets. However, what makes APP particularly quantifiers across all class distributions is the radar chart.
dangerous is that E[] is fixed for all datasets. Therefore, Figure 6 illustrates two of these plots for the datasets
a single mistake can consistently improve the quantifier Numerai (left) and Bank Marketing (right) and quantifiers
performance for several datasets. DyS, HDy and lazy.</p>
      <p>Before we conclude this section, let us define a baseline The lazy quantifier shows a well-defined heart shape.
quantifier known as lazy. We can use lazy as a reference For the Bank Marketing dataset, the performance of DyS
quantifier and require all assessed methods to outperform and HDy is worse than the baseline for most class
dislazy in an APP setup. tributions. Meanwhile, for the Numerai dataset, DyS</p>
      <p>Lazy is a quantifier with a constant output equals to outperforms HDy for all distributions. These two
quantiE[]. Lazy is the best constant quantifier we can conceive, ifers also outperform lazy for the majority of the  values.
and in APP setup with uniform  distribution, it has a Radar charts are helpful to spot when a quantifier
perconstant expected MAE of .25 across all test sets. forms diferently according to . For instance, HDy
performs better for small  values (such as .25) than large
ones (around .75) in the Numerai dataset. Figure 5 shows
more extreme cases of such performance imbalance that
are dificult to identify when we analyse mean
performance numbers, such as in Table 4</p>
      <p>We are aware that the community has used other plots
to visualise the performance of quantifiers. A plot
often seen in the literature is a scatter plot with the -axis
representing the true and -axis the predicted class
prevalences. Thus, the diagonal line represents the optimal
quantifier. However, we note that the radar plot and this
scatter plot convey diferent information. The radar plot
informs the average MAE, while the scatter plot presents
the average ˆ for each value of .</p>
      <p>We find the scatter plot interesting to access if a
certain quantifier provides biased estimates, i.e., under or
overestimates ˆ for diferent values of . However, it may
not be a handy tool to access the accuracy of unbiased
quantifiers. In contrast, the radar plot directly shows
the quantification error for diferent values of .
Therefore, it can be a more valuable tool to assess the error
distribution according to .</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion and Future Work</title>
      <sec id="sec-7-1">
        <title>This paper reviews the experimental protocols used in</title>
        <p>quantification and discusses the shortcomings of APP.
As the name suggests, APP has an artificial component
that modifies the test set distribution to create multiple
test sets with diferent class distributions.</p>
        <p>The artificial part of the APP creates some structural
regularities that are not present in the real world. Such
regularities can be subject to exploitation by quantifiers,
leading to an improper increase in counting accuracy.</p>
        <p>The objective of this paper is to call the attention of
the community to the shortcomings of APP. Therefore,
reducing the chance of inadvertently taking advantage
of them when proposing new quantification methods.</p>
        <p>We propose lazy, a baseline quantifier that can help
us to identify when quantifiers are performing poorly.
Lazy constantly outputs E[]. Any useful quantifier must
outperform lazy.</p>
        <p>Also, we propose the use of radar charts as a visual
tool to compare quantifiers. These plots are helpful to
understand how the performance of a quantifier varies</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <sec id="sec-8-1">
        <title>We thank the anonymous reviewers for their insightful comments. This work has been partially funded by TWAS-CNPq (139467/2017-3).</title>
        <p>comparative evaluation of quanticfiation methods, 564–575.</p>
        <p>arXiv preprint arXiv:2103.03223 (2021). [24] A. Maletzke, D. dos Reis, E. Cherman, G. Batista,
[12] W. Hassan, A. Maletzke, G. Batista, Accurately On the need of class ratio insensitive drift tests
quantifying a billion instances per second, in: IEEE for data streams, in: International Workshop on
International Conference on Data Science and Ad- Learning with Imbalanced Domains: Theory and
vanced Analytics (DSAA), IEEE, 2020, pp. 1–10. Applications, volume 94 of Proceedings of Machine
[13] L. Milli, A. Monreale, G. Rossetti, F. Giannotti, D. Pe- Learning Research, PMLR, ECML-PKDD, Dublin,
dreschi, F. Sebastiani, Quantification trees, in: 2013 Ireland, 2018, pp. 110–124. URL: http://proceedings.
IEEE 13th International Conference on Data Min- mlr.press/v94/maletzke18a.html.</p>
        <p>ing, IEEE, 2013, pp. 528–536. [25] Q. Yuan, G. Cong, N. M. Thalmann, Enhancing
[14] A. Bella, C. Ferri, J. Hernández-Orallo, M. J. Ramirez- naive bayes with various smoothing methods for
Quintana, Quantification via probability estimators, short text classification, in: Proceedings of the 21st
in: IEEE International Conference on Data Mining, International Conference on World Wide Web, 2012,
IEEE, 2010, pp. 737–742. pp. 645–646.
[15] G. Forman, Quantifying counts and costs [26] D. dos Reis, A. Maletzke, D. F. Silva, G. E. A.
via classification, Data Mining and Knowl- P. A. Batista, Classifying and counting with
reedge Discovery 17 (2008) 164–206. URL: http://dx. current contexts, in: ACM SIGKDD International
doi.org/10.1007/s10618-008-0097-y. doi:10.1007/ Conference on Knowledge Discovery and Data
s10618-008-0097-y. Mining, ACM, 2018, pp. 1983–1992. doi:10.1145/
[16] A. Esuli, F. Sebastiani, Optimizing text quanti- 3219819.3220059.</p>
        <p>ifers for multivariate loss functions, ACM Transac- [27] D. Dheeru, E. Karra Taniskidou, UCI machine
learntions on Knowledge Discovery from Data (TKDD) ing repository, 2017. URL: http://archive.ics.uci.edu/
9 (2015) 1–27. ml.
[17] D. D. Lewis, Y. Yang, T. Russell-Rose, F. Li, Rcv1: A [28] J. Vanschoren, J. N. van Rijn, B. Bischl, L. Torgo,
new benchmark collection for text categorization Openml: Networked science in machine learning,
research, Journal of machine learning research 5 ACM SIGKDD Explorations Newsletter 15 (2013)
(2004) 361–397. 49–60. doi:10.1145/2641190.2641198.
[18] D. Card, N. A. Smith, The importance of calibration [29] J. Sayyad Shirabad, T. Menzies, The PROMISE
reposfor estimating proportions from annotations, in: itory of software engineering databases., School of
Proceedings of the 2018 Conference of the North Information Technology and Engineering,
UniverAmerican Chapter of the Association for Computa- sity of Ottawa, Canada, 2005. URL: http://promise.
tional Linguistics: Human Language Technologies, site.uottawa.ca/SERepository.</p>
        <p>Volume 1 (Long Papers), 2018, pp. 1636–1646. [30] L. Candillier, V. Lemaire, Design and analysis of
[19] W. Gao, F. Sebastiani, From classification to quantifi- the nomao challenge active learning in the
realcation in tweet sentiment analysis, Social Network world, in: ALRA: Active Learning in Real-world
Analysis and Mining 6 (2016) 19. Applications, Workshop ECML-PKDD, 2012.
[20] J. Barranquero, P. González, J. Díez, J. J. del Coz, [31] A. Moreo, A. Esuli, F. Sebastiani, Quapy: A
pythonOn the study of nearest neighbor algorithms for based framework for quantification, arXiv preprint
prevalence estimation in binary problems, Pattern arXiv:2106.11057 (2021).</p>
        <p>Recognition 46 (2013) 472–482. URL: http://dx.doi. [32] S. Moro, P. Cortez, P. Rita, A data-driven approach
org/10.1016/j.patcog.2012.07.022. doi:10.1016/j. to predict the success of bank telemarketing,
Decipatcog.2012.07.022. sion Support Systems 62 (2014) 22–31.
[21] Y. S. Chan, H. T. Ng, Estimating class priors in do- [33] I.-C. Yeh, C.-h. Lien, The comparisons of data
minmain adaptation for word sense disambiguation, in: ing techniques for the predictive accuracy of
prob21st International Conference on Computational ability of default of credit card clients, Expert
SysLinguistics and the 44th Annual Meeting of the tems with Applications 36 (2009) 2473–2480.
Association for Computational Linguistics,
Association for Computational Linguistics, Stroudsburg,</p>
        <p>PA, USA, 2006, pp. 89–96.
[22] P. González, A. Castaño, N. V. Chawla, J. J. D. Coz, A
review on quantification learning, ACM Computing</p>
        <p>Surveys (CSUR) 50 (2017) 74.
[23] D. M. dos Reis, A. G. Maletzke, E. Cherman, G. E.</p>
        <p>Batista, One-class quantification, in: European
Conference on Machine Learning, Dublin, 2018, pp.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <article-title>Re-assessing the “classify and count” quantification method</article-title>
          , arXiv preprint arXiv:
          <year>2011</year>
          .
          <volume>02552</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Maletzke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>M. dos</article-title>
          <string-name>
            <surname>Reis</surname>
            ,
            <given-names>G. E. Batista,</given-names>
          </string-name>
          <article-title>The importance of the test set size in quantification assessment</article-title>
          .,
          <source>in: IJCAI</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2640</fpage>
          -
          <lpage>2646</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <article-title>Evaluation measures for quantification: An axiomatic approach</article-title>
          , Inf. Retr. J.
          <volume>23</volume>
          (
          <year>2020</year>
          )
          <fpage>255</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sebastiani</surname>
          </string-name>
          ,
          <article-title>Tweet sentiment quantification: An experimental re-evaluation, arXiv preprint</article-title>
          arXiv:
          <year>2011</year>
          .
          <volume>08091</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Why</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Batista</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mafra-Neto</surname>
          </string-name>
          , E. Keogh,
          <article-title>Flying insect detection and classification with inexpensive sensors</article-title>
          ,
          <source>Journal of visualized experiments: JoVE</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Forman</surname>
          </string-name>
          ,
          <article-title>Counting positives accurately despite inaccurate classification</article-title>
          ,
          <source>in: ECML</source>
          , Springer,
          <year>2005</year>
          , pp.
          <fpage>564</fpage>
          -
          <lpage>575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>González-Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alaiz-Rodríguez</surname>
          </string-name>
          , E. Alegre,
          <article-title>Class distribution estimation based on the hellinger distance</article-title>
          ,
          <source>Information Sciences 218</source>
          (
          <year>2013</year>
          )
          <fpage>146</fpage>
          -
          <lpage>164</lpage>
          . doi:http://dx.doi.org/10.1016/j.ins.
          <year>2012</year>
          .
          <volume>05</volume>
          .028.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maletzke</surname>
          </string-name>
          , D. dos
          <string-name>
            <surname>Reis</surname>
            , E. Cherman,
            <given-names>G. E. A. P. A.</given-names>
          </string-name>
          <string-name>
            <surname>Batista</surname>
          </string-name>
          ,
          <article-title>Dys: a framework for mixture models in quantification</article-title>
          ,
          <source>in: 33th AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sosik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. E.</given-names>
            <surname>Peacock</surname>
          </string-name>
          , E. Brownlee,
          <article-title>Annotated plankton images data set for developing and evaluating classification methods</article-title>
          ,
          <year>2015</year>
          . URL: https: //hdl.handle.
          <source>net/10</source>
          .1575/
          <year>1912</year>
          /7341.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>González</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Castano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. E.</given-names>
            <surname>Peacock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Díez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Del Coz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Sosik</surname>
          </string-name>
          ,
          <article-title>Automatic plankton quantification using deep features</article-title>
          ,
          <source>Journal of Plankton Research</source>
          <volume>41</volume>
          (
          <year>2019</year>
          )
          <fpage>449</fpage>
          -
          <lpage>463</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Schumacher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Strohmaier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lemmerich</surname>
          </string-name>
          , A
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>