<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop on AI Evaluation Beyond Metrics, July</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Relevance of Non-Human Errors in Machine Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ricardo Baeza-Yates</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marina Estévez-Almenzar</string-name>
          <email>marina.estevez@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Machine Learning, Responsible AI, Evaluation, Error Analysis, Non-Human Errors</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DTIC, Pompeu Fabra University</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Experiential AI, Northeastern University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>25</volume>
      <issue>2022</issue>
      <abstract>
        <p>The current practice of focusing the evaluation of a machine learning model on the accuracy of validation has been lately questioned, and has been declared as a systematic habit that is ignoring some important aspects when developing a possible solution to a problem. This lack of diversity in evaluation procedures reinforces the diference between human and machine perception on the relevance of data features, and reinforces the lack of alignment between the fidelity of current benchmarks and human-centered tasks. Hence, we argue that there is an urgent need to start paying more attention to the search for metrics that, given a task, take into account the most humanly relevant aspects. We propose to base this search on the errors made by the machine and the consequent risks involved in moving human logic away from that of the machine. If we work on identifying these errors and organize them hierarchically according to this logic, we can use this information to provide a reliable evaluation of machine learning models, and improve the alignment between training processes and the diferent considerations humans make when solving a problem and analyzing outcomes. In this context we define the concept of non-human errors, exemplifying it with an image classification task and discussing its implications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Imagine that you enter a skyscraper and the elevator has
a sign that says: “Works 99% of the time”. Would you take
the elevator? Most people would not. However, if the
sign says “Does not work 1% of the time and when that
happens, stops”, you probably would use it, because you
perceive that you will be safe thanks to the explanation of
the error and the possibility to evaluate the consequences:
“The elevator may fail, but when it does, the fail consists
of stopping”. Today, Machine Learning (ML) models are
evaluated primarily on the basis of success rather than
failure. Worse, this evaluation does not take into account
the potential harm of its mistakes, like is done in the
pharmaceutical or the food industry.</p>
      <p>Along the same lines as this example, the current
benchmarks fidelity to human-centered tasks has
recently been called into question [1, 2, 3]. The practice
curacy has been stated as a dangerous habit [4, 5] that
is ignoring some important aspects of the human
perception when developing a solution for a problem, such
as carefully studying the risks of the solution and its
diferent points of operation. This lack of diversity in
evaluation procedures reinforces the diference between</p>
      <sec id="sec-1-1">
        <title>Therefore, after discussing the state of the art in Sec</title>
        <p>tion 2, we introduce the concept of non-human errors in</p>
      </sec>
      <sec id="sec-1-2">
        <title>Section 3. This concept allow us to build our error tax</title>
        <p>of centering the model evaluation on the validation ac- fully represent human perception but should, at least,
onomy. Then we use an image classification task to do have been proposed. For example, [5] points out that
aca proof of concept that uses our taxonomy in Section 4, curacy alone cannot distinguish between strategies. Two
ending with a discussion of its consequences in Section 5. systems – brains or algorithms – may achieve similar
acA simple notebook to illustrate this work is available in curacy with very diferent strategies. In their study, they
Github.1 conclude that the consistency between human errors and
errors made by deep learning models is not far away from
what can be expected by chance alone, indicating that
2. Related Work machines still employ very diferent perceptual
mechanisms. There are also some benchmarks that approach
The practice of centering the model evaluation on accu- decentralisation with respect to accuracy. In [7] the
auracy has been and still is being questioned. For example, thors propose to study the impact of errors and compare
[6] warns that the only-use of accuracy to measure ma- them according to their type, but both the proposed error
chine performance works as a limitation to humans when classification they ofer and the associated impact focus
analyzing machines, and states that adversarial vulner- only on pose estimation algorithms.
ability results from the susceptibility of models to data In [8], the author also proposes to pay particular
atfeatures that are potentially candidates for generalization. tention to the analysis of error from a quantitative
perThey recall that the fact that we train machines to solely spective. He proposes to focus this analysis on those
maximize accuracy is making its learning system use any errors whose correction has the greatest impact in terms
available signal to achieve this goal, even those signals of improving the accuracy of the algorithm. Even though
that look incomprehensible to humans. this analysis is fundamental, we propose to focus on</p>
        <p>This idea is also supported by [1], who states that train- improving our models in qualitative terms. Given a
coning a model in a robust way leads to a reduction of accu- text, if a type of error is suficiently serious, illogical or
racy. They argue that this trade-of between the accuracy risky, it does not matter if it is made infrequently: we
of a model and its robustness to adversarial perturbations should work on minimising this type of error in order to
is a consequence of robust classifiers learning fundamen- minimise the possible harmful consequences. Another
tally diferent feature representations than standard clas- important discussion that the author mentions is how
sifiers. These diferences, in particular, seem to result to define human-level performance in order to compare
in unexpected benefits: the representations learned by it with the performance of a machine. This highlights
robust models tend to align better with salient data char- the importance of considering the context in which the
acteristics and human perception. Other trade-ofs of machine learning model is applied. The human-level
perthis kind were recently addressed for language models formance we choose to consider will depend on the task
and their intrinsic risks in [3], where authors state that itself, or the harm risks it may pose.
researchers are extending the state of the art on a wide To solve the problem of the misalignment between
array of tasks as measured by classification scores on state-of-the-art benchmarks and human-centered tasks,
some benchmarks, following the methodology of using some of the works mentioned above propose as the main
some pre-trained models and then fine-tuning them for solution either to correct the labels in data sets or redefine
specific tasks. In this scenario, they take a step back and the way we store and represent these labels. Even when
pay attention to the possible risks associated with this these corrections are essential, using correctly labeled or
technology in terms of environmental and financial costs. redefined labels to evaluate models could still not be
suf</p>
        <p>Another research that calls into question the results ifcient to cover the diversity of human perception. While
obtained with current state-of-the-art benchmarks was working on improving the representations associated
done by [4], that highlights the fact that some data sets with the inputs in ML systems is very necessary, we need
contain errors in their labels, and they expose a subse- to do a similar efort on improving the way we interpret
quent study about the potential for these label errors to and analyze the outputs. These outputs are mainly
charafect benchmark results. Surprisingly, they find that acterized by two elements: successes and failures. Until
lower capacity models may be practically more useful now, ML evaluation metrics have been mainly based on
than higher capacity models in real-world data sets with successes, giving visibility to the accuracy of the
algohigh proportions of erroneously labeled data. They con- rithm over other possible ways of measuring the overall
clude that ML practitioners must be careful when choos- performance of the algorithm. We propose to change the
ing which model to deploy based on validation or test focus and start prioritizing the analysis of errors, as well
accuracy. as their classification according to the potential damage</p>
        <p>In an attempt to overcome the limitation that entails they may cause in the context in which they occur.
the only-use of accuracy in ML evaluation, other metrics
1https://github.com/ealmenzar/non-human-errors
crete use case is specified, we wonder whether we can
determine which errors are the most relevant in terms of
human harm risk. Although harm risk may be perceived
very diferently depending on the performer (human or
machine), it is clear that we are interested in avoiding
harmful consequences for the humans involved, directly
or indirectly, in the task at hand. It is reasonable to think
(a) (b) that the errors related to these consequences are those
that are unexpected and atypical for humans, and
thereFmiagnucrees.1O: nVtihseu alelft,rtehperreesdentrtiaatniognleorefphruesmenatns athnedMMLLmpoedrefol,r- fore those that are dificult for us to explain and control.
the blue ellipse represents the human, and the green sphere Since, as humans, we are accustomed to human errors,
represents the ground truth. They are positioned in the solu- we might expect that those errors that are furthest away
tion space of a binary prediction problem. For every figure, from the errors that a human might make could be
conpositive answers are inside, and negative answers are outside, sidered risky: we refer to these types of disparate errors
being the correct answers determined by the green sphere. On as non-human errors (see Figure 3).
the right we can see the yellow region representing the correct We can also formalise this idea in terms of
mathematanswers obtained by both the human and the ML model. ical sets. This will help us to formally define the
diferent types of errors mentioned and graphically expressed
above. We denote  as the green sphere,  as the red
trian3. Non-Human Errors gle, and  as the blue ellipse (see in Figure 1a). Following
the logic explained above, we could consider these sets
Finding new methodologies and metrics that allow us to of points in the solution space (and their complementary
exploit the valuable information in errors done by ma- sets, noted as  ,  , and  respectively) as follows:
chines is not trivial. To illustrate in a simplified way the
error exploration that we are proposing, let us consider  ≡ true positives
a problem with a binary solution space. In this space, we  ≡ positives predicted by the model
are given the ground truth, so we are capable of deter-  ≡ positives predicted by the human
mining whether a point from the space is a correct or an
incorrect answer, as well as its impact. For example, a  ≡ true negatives
typical assumption when solving a binary classification  ≡ negatives predicted by the machine
problem, is to consider that false negatives have the same  ≡ negatives predicted by the human
weight of false positives. This is not always correct,
because their harm might be quite diferent. Indeed, when Focusing on the errors shown in Figure 2, now we
depredicting an illness, a physician will prefer to see many note the false positives errors made only by the machine
more healthy patients just to avoid missing any ill one. as   (Figure 2d), and we define this set of errors as
One solution is to use a weighted accuracy but still the
operational point might be diferent because here the re-   =  ∩  ∩ 
call of ill patients is much more relevant than the overall Similarly, we denote and define the false positives
eraccuracy (weighted or not). rors made only by the human (Figure 2e), the false
pos</p>
        <p>In Figure 1 we can see this ground truth represented itive errors made by both the machine and the human
as a green sphere, positioned on the solution space, such (Figure 2f), the false negative errors made only by the
that those points that fall into the green area are the true machine (Figure 2d), the false negative errors made only
positive answers, and the rest of the points are the true by the human (Figure 2b), and the false negative errors
negative answers. We can see two more shapes in this made by both the machine and the human (Figure 2c) as
space that symbolize the perceptual agents; a red triangle follows, respectively:
representing the machine predictions, and a blue ellipse
representing the human predictions. Following the pre-  ℎ =  ∩  ∩ 
vious logic, in Figure 1b we can see the true answers   =  ∩  ∩ 
correctly predicted by both the model and a human. In
Figure 2, we focus on the errors. Here we are able to dis-   =  ∩  ∩ 
tinguish between two kinds of errors: false positives and  ℎ =  ∩  ∩ 
false negatives. And we make another distinction based
on the entity that is making the error (human and/or ML   =  ∩  ∩ 
model). Note that all these sets are disjoint because of the
excluIn this general and abstract scenario, where no con- sivity imposed when considering which agent commits
points. The distinction of non-human errors is based on
the distinction between the successes and mistakes made
by the diferent perceptual agents (human and machine).</p>
        <p>Also, this distinction only makes sense in a context in
which we can expect reasonable human performance.</p>
        <p>Thus, the category of non-human errors can be found in
those human centered tasks that can be at least partially
solved in a reasonable way by the humans and where a
ML algorithm is applied instead. However, in practice,
consideration of the specific use case will be decisive.</p>
        <p>We next apply this idea to a simple but illustrative
Figure 3: Non-human errors stressed in red: both false neg- problem: classifying images of dogs and cats according
atives and false positives done by the ML model but not by to their breed [9]. This translates into a fine-grained
imhumans (cases (a) and (d) in Figure 2). age classification problem that is mainly solved by using
deep neural networks. Based on expert sources in the
classification of these animals ( FIFe and FCI Federations),
the error. Now the sets of interest arise from the union of we have been able to construct a taxonomy that
represome of the previous sets. We note  as the non-human sents the possible errors that can be made in this task.
errors (those committed by the machine but not by the Following our definition of non-human errors, in this
human) explained above,  as those errors committed by problem we can identify them as those errors that are
the human but not by the machine, and  as those errors fundamentally diferent from the errors that a human
committed by both the human and the machine together: solving this task would commit. Therefore, we define
as non-human errors those cases in which the machine
 =   ∪   classifies a dog as a cat, or vice versa (see Figure 4).
No =  ℎ ∪  ℎ tice that there might be other non-human errors when
 =   ∪   comparing among only cats or dogs, but those are much
less important and less common than the definition that</p>
        <p>In this paper we focus on  , non-human errors, which we use for this proof of concept and provides a lower
we believe are the errors that we should address first be- bound for non-human errors.
cause of the harm risks that could be involved in making So far, we have selected one of the top-ranked
algoerrors that escape human logic. But how can we precisely rithms for solving this specific task, the Big Transfer (BiT)
determine these errors? How can we measure how far model from [10], which achieved 93% of accuracy. In
Figan answer should be from human logic in order to call it ure 6 we give the full confusion matrix of 3,312 prediction
a non-human error? We address these challenges in the pairs among 25 dog breeds (top-left) and 12 cat breeds
next section. (bottom-right), where we can see that there are 4 pairs
that are hard to classify (two breeds of Terriers and 3
4. Proof of Concept pairs of cat breeds). Notice that this confusion matrix
is in general non-symmetric, as the output of the model
may difer because the input and the prediction for each
pair is diferent.</p>
        <p>Here we found that more than 3% of the errors were
non-human errors (8 of 241 errors), which appear as light</p>
      </sec>
      <sec id="sec-1-3">
        <title>Approaching a problem by adopting the previous abstract perspective allows us to visualize it with some independence from the use case or real-world application, which is good for understanding the wide range of operating</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5. Discussion</title>
      <sec id="sec-2-1">
        <title>Why should we care if algorithms mistake dogs for cats?</title>
        <p>This is clear when similar tasks are proposed in fields
where the lives of human beings and their
fundamental rights are at risk of being left unprotected. In these
ifelds, even in the case of a low percentage of non-human
errors, the consequences could have a catastrophic and
irreversible impact. One concrete such example happened
(b) in 2018, when a Uber self-driving car was not able to
recognize a woman in a bicycle crossing a road at night in
Figure 5: Two of the non-human errors obtained when run- Tempe, Arizona.2. A human most probably would have
ning the BiT model [10] over the Oxford-IIIT Pets data set [9]. recognized the woman and hence this is a non-human
In (a) a Chihuahua is mistaken for an Abyssinian cat with a error. We do not know if the backup driver could have
confidence of 46.24%. In (b) a Bengal cat is mistaken for a Chi- reacted on time, but she was seeing a video as the car
thoutahheuaonweitohf tahceosnefcidonendcoepotifo2n0.i8n3t%h,ealipsterocfebnrteaegdesvseorrytecdlobsye was working well until then. Finally, she was charged
their probability of being selected as the tag for that image. of negligence, as Uber quickly settled with the family of
the victim to avoid being sued [11].Hence, this event at
the end impacted the lives of two women.</p>
        <p>One related issue that we do not discuss is another bad
squares in the top-right and bottom-left of Figure 6. Two habit: predicting an answer even when we have low
conof these errors are shown in Figure 5, where a Chihuahua ifdence. For example, in Figure 5 (b), any smart/honest
is classified as an Abyssinian cat (Figure 5a) and a Bengal person would say “I don’t know” with such low
conficat is classified as a Chihuahua (Figure 5b). However, dence. Even in case (a), if there is a harm risk, not giving
there is a notable diference between these two errors: an answer might be a safer output. In the Uber example is
the certainty of the answer provided by the algorithm. the same. Predicting ”I don’t know” and stopping, might
This supports the need to start providing new metrics. be safer than predicting ”there is no human in front of me
In this case, for instance, it would be interesting to focus and is safer to run over the object” (notice that the later
on the extent to which an algorithm is, under unreliable assumption might be still dangerous for the passengers).
certainty, either predicting correctly or erring, regardless
of whether the answer is right or wrong.</p>
        <p>2https://www.theguardian.com/technology/2018/mar/19/uberself-driving-car-kills-woman-arizona-tempe</p>
      </sec>
      <sec id="sec-2-2">
        <title>A. Madry, From imagenet to image classification:</title>
        <p>Contextualizing progress on benchmarks, in:
International Conference on Machine Learning, PMLR,
2020, pp. 9625–9635.
[3] E. M. Bender, T. Gebru, A. McMillan-Major,</p>
        <p>S. Shmitchell, On the dangers of stochastic
parrots: Can language models be too big?, in: FAccT
’21: 2021 ACM Conference on Fairness,
Accountability, and Transparency, Virtual Event / Toronto,</p>
        <p>Canada, March 3-10, 2021, ACM, 2021, pp. 610–623.
[4] C. G. Northcutt, A. Athalye, J. Mueller, Pervasive
Figure 7: Classification of white blood cells. As it happened label errors in test sets destabilize machine learning
with cats and dogs, after further investigation on the risk benchmarks, arXiv preprint 2103.14749 (2021).
caossuoldcibaeteudsetdoams iasttaaxkoenoonmeycteolldfeofrinaendoitshpearra[t1e2,er1r3o]r,st.hImis atgreees [5] R. Geirhos, K. Meding, F. A. Wichmann, Beyond
collected by [14]. accuracy: Quantifying trial-by-trial behaviour of
cnns and humans by measuring error consistency,
arXiv preprint arXiv:2006.16736 (2020).
[6] A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom,
In fact the self-driving car did predict a bicycle one of the B. Tran, A. Madry, Adversarial examples are
times [11]. not bugs, they are features, arXiv preprint</p>
        <p>We are currently working on a problem that is tech- arXiv:1905.02175 (2019).
nically very similar to the classification of dogs and cats [7] M. Ruggero Ronchi, P. Perona, Benchmarking and
according to breed, but is a real-life application that is error diagnosis in multi-instance pose estimation,
much more relevant to humans: the classification of in: Proceedings of the IEEE international
conferwhite blood cells. This problem is also formulated as ence on computer vision, 2017, pp. 369–378.
a fine-grained image classification problem and, even [8] A. Ng, Machine learning yearning,
when the number of diferent classes of elements is much 2018. URL: https://info.deeplearning.ai/
smaller than in the previous example (see Figure 7), their machine-learning-yearning-book.
diferentiation is very important. Indeed, [ 12] points out [9] O. M. Parkhi, A. Vedaldi, A. Zisserman, C. Jawahar,
that neutrophil levels were associated with breast can- Cats and dogs, in: 2012 IEEE conference on
comcer risk, including advanced stages of breast cancer. In puter vision and pattern recognition, IEEE, 2012,
the meta-analysis proposed by [13], it was shown that pp. 3498–3505.
breast cancer patients with a higher ratio of neutrophils [10] A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver,
to lymphocytes had a higher relapse and lower overall J. Yung, S. Gelly, N. Houlsby, Big transfer (bit):
survival. General visual representation learning, in: 16th</p>
        <p>The importance of including an in-depth study of the European Conference on Computer Vision, Part V
errors that an algorithm could make in this classification 16, Springer, 2020, pp. 491–507.
is evident, just as it is fundamental that in these complex [11] L. Smiley, ‘I’m the operator’: The aftermath of a
use cases, both the evaluations of the algorithms and their self-driving tragedy, Wired (2022).
publication are accompanied by the corresponding pa- [12] Y. Okuturlar, M. Gunaldi, E. E. Tiken, B.
Oztorameters or new metrics that make visible the diferent er- sun, Y. O. Inan, T. Ercan, S. Tuna, A. O. Kaya,
rors made, their frequency and their associated risk based O. Harmankaya, A. Kumbasar, Utility of
periphon professional knowledge. Moreover, these parameters eral blood parameters in predicting breast cancer
could provide not only transparency and explainability risk, Asian Pacific Journal of Cancer Prevention 16
to the model, but also valuable clues to researchers that (2015) 2409–2412.
would allow the algorithms to be improved in terms of [13] B. Wei, M. Yao, C. Xing, W. Wang, J. Yao, Y. Hong,
human-centered responsibility and accountability. Y. Liu, P. Fu, The neutrophil lymphocyte ratio is
associated with breast cancer prognosis: an updated
References systematic review and meta-analysis, OncoTargets
and therapy 9 (2016).
[1] D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, [14] X. Zheng, Y. Wang, G. Wang, J. Liu, Fast
A. Madry, Robustness may be at odds with accuracy, and robust segmentation of white blood cell
imarXiv preprint arXiv:1805.12152 (2018). ages by self-supervised learning, Micron 107
[2] D. Tsipras, S. Santurkar, L. Engstrom, A. Ilyas, (2018) 55–71. URL: https://www.sciencedirect.com/
science/article/pii/S0968432817303037.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>