<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Metrics for Explainable AI: Challenges and Prospects. CoRR abs/</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Position: We Can Measure XAI Explanations Beter with Templates</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jonathan Dodge</string-name>
          <email>dodgej@eecs.oregonstate.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Explainable AI, Evaluating XAI Explanations, Empirical Studies,</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Margaret Burnett</string-name>
          <email>burnett@eecs.oregonstate.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Heuristic Evaluations</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Oregon State University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1812</year>
      </pub-date>
      <volume>04608</volume>
      <issue>3</issue>
      <fpage>1328</fpage>
      <lpage>1334</lpage>
      <abstract>
        <p>This paper argues that the Explainable AI (XAI) research community needs to think harder about how to compare, measure, and describe the quality of XAI explanations. We conclude that one (or a few) explanations can be reasonably assessed with methods of the “Explanation Satisfaction” type, but that scaling up our ability to evaluate explanations requires more development of “Explanation Goodness” methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Human-centered computing → User studies.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>As AI plays an ever-increasing role in our lives, society needs a
variety of tools to inspect them. Explanations have emerged to fill
that role, but measuring their quality continues to prove challenging.
Hofman et al. [ 10] ofers two terms to describe mechanisms for
measuring the quality of an explanation, quoted at length here,
with highlighting added to assist this paper’s discussion.
“Explanation Goodness: Looking across the scholastic and
research literatures on explanation, we find assertions about
what makes for a good explanation, from the standpoint of
statements as explanations. There is a general consensus on
this; factors such as clarity and precision. Thus, one can look
at a given explanation and make an a priori (or
decontextualized) judgment as to whether or not it is “good.” ... In a proper
experiment, the researchers who complete the checklist
[evaluation of the explanation] with reference to some particular
AI-generated explanation, would not be the ones who created
the XAI system under study.”
“Explanation Satisfaction: While an explanation might be
deemed good in the manner described above, it may at the
same time not be adequate or satisfying to users-in-context.
• Contextualization: ExpS is defined relative to a task, while</p>
      <p>ExpG is not.
• Actor: ExpS is measured from the perspective of a user
performing a task, while ExpG is from the perspective of
researchers (ideally dispassionate bystanders, but often the
designers themselves).
• Timing: Because ExpS is defined relative to a task, it must
be measured after the task is completed, while ExpG can be
measured anytime.</p>
      <p>The main thesis of this paper will be that we, as a research
community, need to think harder about how to compare, measure,
and describe the ExpG of explanation templates (clarified later in
Section 5), as a potentially strong complement to ExpS.</p>
      <p>To develop the main thesis, we attempt to argue several points:
• Background: Most current research has focused on ExpS.
• Tasks: ExpS is easier to operationalize, but incurs a great deal
of experimental noise.
• Benefits: ExpS’s usefulness is hampered by participants’
limited exposure to the system.
• Scope: ExpG afords the opportunity to consider a wider range
of behaviors, making ExpG mechanisms particularly well
suited to reasoning about explanation templates.</p>
      <p>Through these points, we hope to provoke thought about how
explanation designers can better validate their design decisions via
ExpG for explanation templates. This is of particular importance
because a great many design decisions are never evaluated via
ExpS mechanisms.
2</p>
    </sec>
    <sec id="sec-3">
      <title>BACKGROUND: MOST CURRENT</title>
    </sec>
    <sec id="sec-4">
      <title>RESEARCH HAS FOCUSED ON EXPS</title>
      <p>
        How have past researchers evaluated XAI design decisions? Most
have used ExpS mechanisms, but a few have used ExpG mechanisms.
For both types, an important criterion is rigor. The dangers of
departure from rigorous processes for validating design decisions
can be potentially severe if it devolves into “I methodology” [
        <xref ref-type="bibr" rid="ref4">24</xref>
        ], i.e.,
designers relying solely on their own views and assumptions about
what their users will need and how they will use the functionalities
the designers decide to provide.
2.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Research that uses ExpS mechanisms</title>
      <p>Through extensive literature review, Hofman et al. [ 10] identified a
group of existing ExpS methods for mental model elicitation (their
Table 4). Among them, many are essentially qualitative and focus
on things people say (e.g. Think Aloud or Interview techniques).
We felt the “Retrospection Task” [18] and the “Prediction Task” [20]
looked to be the most well suited for quantitative study, and chose
to use them for Anderson et al.’s empirical studies [4]. Approaching
the problem from another angle, Dodge et al. [8] investigated several
aspects of perceptions of fairness and explanations in a decision
support setting.</p>
      <p>
        Other researchers have used ExpS to understand a wide
variety of efects in explanation. Providing explanation has been
shown to improve mental models [15, 16]. Of particular
importance to moderating the efects of explanation is the explanation’s
soundness and completeness [17]; most easily described with
the phrase “the whole truth (completeness) and nothing but the
truth (soundness)” about how the system is really working. Note
that neither soundness nor completeness are binary properties, but
a smooth continuum—with 100% soundness or completeness not
always achievable. Explanation has also been shown to increase
satisfaction (here we mean in the colloquial sense, the user’s
selfreported feeling) [
        <xref ref-type="bibr" rid="ref2">2, 12</xref>
        ], and understanding—particularly in low
expertise observers [
        <xref ref-type="bibr" rid="ref7">27</xref>
        ]. Several diferent kinds of explanation have
also been shown to improve user acceptance via setting appropriate
expectations (e.g. by showing an accuracy gauge) [14]. There are
many other researchers studying explanations using ExpS
mechanisms, and we refer the reader to Abdul et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for a recent
literature review.
2.2
      </p>
    </sec>
    <sec id="sec-6">
      <title>Research that uses ExpG mechanisms</title>
      <p>
        One XAI tool for ExpG is the checklist proposed by Hofman et
al. [10]’s Appendix A, composed of 8 yes/no questions that
researchers and designers ask themselves (e.g. “The explanation of
the [software, algorithm, tool] is suficiently detailed.” ). Other
approaches include Amershi et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]’s guidelines for interactive AI,
Kulesza et al. [15]’s design principles for explanatory debugging,
and Wang et al. [
        <xref ref-type="bibr" rid="ref5">25</xref>
        ]’s guidelines to match human reasoning
processes with XAI techniques. However, most XAI research is not
explicit about usage of these or other ExpG mechanisms.
3
      </p>
    </sec>
    <sec id="sec-7">
      <title>TASKS: EXPS IS “EASY” TO</title>
    </sec>
    <sec id="sec-8">
      <title>OPERATIONALIZE, BUT NOISY</title>
      <p>ExpS sings a siren song of sorts; it appears simple to evaluate, one
must simply define a task and criteria to measure performance at
that task. Easy enough, right? Wrong.</p>
      <p>We have been using one of the XAI tasks which we felt would
be most well suited for quantitative study, the “Prediction Task”,
proposed by Muramatsu et al. [20]. However, we have run into a
number of challenges using it in our XAI studies, some of which
are apparent in Figure 1, taken from Anderson et al. [4].</p>
      <p>First, participants’ ability to perform the task (predict an AI’s
next actions) in a domain are moderated by a number of other
things, such as their need for cognition, interest in the task, domain
experience, etc. To illustrate the efect of variability in
explanation consumers, imagine a scientist efectively describing quantum
100%
rcya 80%
ccu 60%
A%
gae 40%
rve
A 20%
0%
1
2
3
4
5
6
7
8
9
10
11
12
13
14
computing to another scientist. Then imagine the same person
giving the same explanation to a child1. The explanation itself
could be high quality, it just was not appropriate for that
audience and needed reformulation. Thus, empirically measuring the
explanation’s quality is entangled with many factors beyond the
explanation itself.</p>
      <p>Second, there is a great deal of variability in the state/action
space (as we observed in [4]). This leads some choices to be easy,
causing all treatments to have nearly 100% participant prediction
accuracy—even those without explanation (e.g. the 5th decision
point in Figure 1). In contrast, others are much harder, and all
treatments had nearly 0% prediction accuracy (e.g. the 4th decision
point in Figure 1). As a result of these floor and ceiling efects, some
of the variation between treatments is obscured.</p>
      <p>Third, it is dificult to assign “partial credit” for participants’
predictions. In Figure 1’s case, participants faced a choice of 4 options,
leading random guessing to be right 25% of the time. However, AI
is regularly used in domains with much larger action spaces, so the
probability a participant picks right can be vanishingly small. As
a result, it seems natural to think about which answers might be
considered better than others. To do so, one might consider
similarity in the action space (actions that look similar) or in the value
space (actions that produce similar consequences), but either way it
is a challenge to design rigorously.
4</p>
    </sec>
    <sec id="sec-9">
      <title>BENEFITS: EXPS’S USEFULNESS IS</title>
    </sec>
    <sec id="sec-10">
      <title>HAMPERED BY LIMITED EXPOSURE</title>
      <p>Consider that in-lab user studies are typically designed to be
executed within a 2-hour window for a variety of reasons (e.g.
reliability). As a result, the amount of participant exposure to the system is
actually quite low. As an example, in Anderson et al.’s study [4] we
showed 14 decision points to participants over the available 2 hours.
In that paper, we point to this limited exposure as a possible reason
that we did not observe any learning efect (evident in Figure 1).</p>
      <p>Other challenges also surround exposing participants to the
system when performing a ExpS evaluation. In particular, which
decision points do we show to participants? Because this agent has
1https://www.youtube.com/watch?v=OWJCfOvochA conducts a similar exercise,
though Dr. Gershon changes the explanation for 5 diferent audiences.
***Prediction: likely to reoffend
The training set contained 8
individuals matching this one.
2 of them reoffended (25%)
***Prediction: likely to reoffend
The training set contained 151
individuals matching this one.
93 of them reoffended (61%)
been training for 30,000 episodes, scrutinizing the training data in
its entirety would be a daunting task. After training is complete,
one could imagine presenting test cases, some of which could be
handcrafted. One way to select test cases to present to assessors is a
recent approach to solving the problem of which decision points to
show by Huang et al. [11]. Their approach measures the “criticality”
of each state, and chooses the ones where the agent perceived its
choice to matter the most.</p>
      <p>Note again, the large variability in state/action space, which we
commented on in Section 3. When combined with limited exposure,
participants are essentially gazing into a vast expanse of behavior
through a tiny peephole.
5</p>
    </sec>
    <sec id="sec-11">
      <title>SCOPE: EXPG CAN CONSIDER A WIDER</title>
    </sec>
    <sec id="sec-12">
      <title>RANGE OF BEHAVIORS, VIA TEMPLATES</title>
      <p>One great advantage of ExpG is, it supports explanation templates.
5.1</p>
    </sec>
    <sec id="sec-13">
      <title>What is an explanation template?</title>
      <p>The earliest evidence we could find for what we call “explanation
templates” is from Khan et al. [13], and we think studying them can
help address some of the problems discussed earlier in this paper.
Explanation templates operate on a diferent granularity than an
explanation. If an explanation describes or justifies an individual
action, the explanation template is like the factory that creates the
explanation.</p>
      <p>We have built templates inspired by Binns et al. [5], who used a
wizard of oz methodology to generate multiple types of explanation
for a decision support setting. The decision they were trying to
explain was an auto insurance quote ([5]’s Figure 2). One of their
types, which they term “Case-based explanations”, is demonstrated
in the left side of Figure 2 as we used them in Dodge et al. [8] to
explain an AI system’s judicial sentencing recommendations. Note
that the templates shown in this paper are for textual explanations,
but the idea extends naturally to other types of explanations, like
the visual explanations in Mai et al.’s Figures 1 and 2 [19].
5.2</p>
    </sec>
    <sec id="sec-14">
      <title>Why consider explanation templates?</title>
      <p>An explanation template can be combined with appropriate
software infrastructure and a test set to generate a large set of
explanation instances. We argue that examining the distribution of
thousands of explanations generated in this way can be more
illuminating than seeing individual ones (e.g., Figure 2, right).</p>
      <p>Although ExpG mechanisms can be used on any explanation
template one desires, consider using a case-based explanation
template. One way to produce these explanations is by finding training
examples “near” the input, then characterize how well the labels of
the resultant set match the label of the input (Figure 2, left). If we
run an explanation generator on the whole test set, we can create a
histogram from those matching percentages, shown on the right
side of Figure 2.</p>
      <p>However, this introduces an issue with explanation soundness.
To see why, note that many instances fall near 100%—these are
the explanations a user might find “convincing”. However, there
is also a good chance the explanation finds that around 50% of
the nearby training examples matched the input’s label—these are
the explanations that do not substantiate any claim (note that the
classifier is binary). Even worse, there are instances that fall near
0%—these are the self-refuting explanations2 along the lines of “This
instance was labelled an A because all the nearby training examples
were B’s.”.</p>
      <p>In this circumstance, the lack of soundness arises from the fact
that the explanation uses nearest neighbors while the underlying
classifier does not. Note also that the fix for these two problems is
diferent. When the explanation lacks evidence, one should go find
more evidence. But when the explanation self-refutes, one must
ifgure out why the contradiction exists.</p>
      <p>Note that if we were evaluating these explanations with ExpS, the
result would depend strongly on whether the provided explanation
was “convincing” or “self-refuting”—but the explanation template is
the same in both cases, only the input changed. On the other hand,
ExpG allows us to consider the wider scope—that the template is
capable of generating explanations which will occasionally refute
itself —and decide if that is acceptable.</p>
      <p>To continue with the example of case-based self-refutation,
suppose a member of the research team proposed an alternative
explanation template that avoids refuting itself3. To do so, we adjust
the static text and the calculation that fills in the variable parts, as
illustrated in Figure 3. Note that the new proposal only highlights
counts on things that match, which has the efect of essentially
ignoring nearby training examples that do not match the input’s
predicted label.</p>
      <p>The advantage of this proposal is that this alternative will not
refute itself. But the advantage came at a cost: according to known
taxonomies, it brings a decrease in completeness, as the explanation
is telling less of the “whole truth.”</p>
      <p>So did the overall ExpG go up or down? We think most would
argue down... but we cannot measure how much. This example
exposes a critical weakness in the ExpG approach: the vocabulary
and calculus currently available to us cannot adequately describe
and measure the implications of a single design decision.
5.3</p>
    </sec>
    <sec id="sec-15">
      <title>Scalability: Many design decisions are only</title>
      <p>validated with ExpG
So where does this weakness leave us?</p>
      <p>During a full design cycle, XAI designers face many design
decisions, of which only a few can be evaluated with ExpS. To illustrate,
consider that case-based explanation as originally proposed by
Binns et al. [5] would be implemented by showing the single
nearest neighbor. That approach could be extended by showing the k
nearest neighbors for a number of diferent k. Or, it could be
implemented by showing whatever neighbors lie within some feature
space volume—which is the approach used by Dodge et al. [8] and
illustrated in the left side of Figure 2. The space of these possible
2The original reason we generated the histogram on the right of Figure 2 was not to
see if the explanation would self-refute, but to compare the two classifiers: one trained
on raw data (raw) and another trained on the processed data (proc). “Processing” the
data refers to the use of a preprocessor by Calmon et al. [6] intended to debias the
data—perhaps inducing a classifier people consider more fair. In this efort, we looked
to see what classifications were diferent, compared confidence score histograms, etc.
3Here we use a strawman explanation template that is known to be bad, in order to
explore the extent of our ability to characterize how bad it is. Correll performed a
similar exercise in the visualization community, proposing Ross-Chernof glyphs as a
strawman, “...as a call to action that we need a better vocabulary and ontology of bad
ideas in visualization. That is, we ought to be able to better identify ideas that seem prima
facie bad for visualization, and better articulate and defend our judgments.” [7].
The training set contained 80 individuals matching this one.
20 of them reoffended (25%)
design decisions is large enough that we cannot hope to evaluate
them all with ExpS mechanisms.</p>
      <p>
        Given this, we suspect many XAI design decisions are never
assessed, unless the XAI designers use ExpG—strictly due to the
impracticality of doing ExpS at the scale needed for full coverage. In
a sense, explanation is a user interface, and user interface designers
have long used a wide variety of techniques relating to ExpG (e.g.
design guidelines [21, 22], cognitive dimensions [9], cognitive
walkthroughs [
        <xref ref-type="bibr" rid="ref6">26</xref>
        ], etc). Fortunately, there are some promising works
that demonstrate the use of these approaches (e.g., [
        <xref ref-type="bibr" rid="ref3">3, 15, 23</xref>
        ]).
      </p>
      <p>Hypothesis: Studying one (or a few) explanations is
wellsuited to ExpS oriented methods, but the template level may
require ExpG methods.</p>
      <p>This suggests that, to create XAI systems according to rigorous
science—especially XAI systems that generate explanations with
a template—we must develop improvements in the rigor and
measurability of ExpG mechanisms. These will bring outsize benefit to
designers and researchers as compared to improvements in ExpS.</p>
    </sec>
    <sec id="sec-16">
      <title>6 ACKNOWLEDGMENTS</title>
      <p>This work was supported by DARPA #N66001-17-2-4030. We would
like to acknowledge all our co-authors on the work cited here, with
specific highlights to Andrew Anderson, Alan Fern, Q. Vera Liao,
Yunfeng Zhang, Rachel Bellamy, and Casey Dugan—research work
does not happen in a vacuum, nor do ideas.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ashraf</given-names>
            <surname>Abdul</surname>
          </string-name>
          , Jo Vermeulen, Danding Wang, Brian Y Lim, and
          <string-name>
            <given-names>Mohan</given-names>
            <surname>Kankanhalli</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Trends and trajectories for explainable, accountable and intelligible systems: An hci research agenda</article-title>
          .
          <source>In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM</source>
          , New York, NY, USA,
          <volume>582</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Amershi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cakmak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Knox</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Kulesza</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Power to the people: The role of humans in interactive machine learning</article-title>
          .
          <source>AI Magazine</source>
          <volume>35</volume>
          ,
          <issue>4</issue>
          (
          <year>2014</year>
          ),
          <fpage>105</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Saleema</given-names>
            <surname>Amershi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Dan</given-names>
            <surname>Weld</surname>
          </string-name>
          , Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal,
          <string-name>
            <given-names>Paul N.</given-names>
            <surname>Bennett</surname>
          </string-name>
          , Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and
          <string-name>
            <given-names>Eric</given-names>
            <surname>Horvitz</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Guidelines for HumanAI Interaction</article-title>
          .
          <source>In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow</source>
          , Scotland Uk)
          <article-title>(CHI '19)</article-title>
          . ACM, New York, NY, USA, Article
          <volume>3</volume>
          , 13 pages. https://doi.org/10.1145/3290605.3300233
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Nelly</given-names>
            <surname>Oudshoorn</surname>
          </string-name>
          and
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Pinch</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>How Users Matter: The Co-Construction of Users and Technology (Inside Technology)</article-title>
          . The MIT Press, Cambridge, MA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Danding</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qian Yang</surname>
            ,
            <given-names>Ashraf</given-names>
          </string-name>
          <string-name>
            <surname>Abdul</surname>
          </string-name>
          , and Brian Y Lim.
          <year>2019</year>
          .
          <article-title>Designing Theory-Driven User-Centric Explainable AI</article-title>
          .
          <source>In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. CHI</source>
          , Vol.
          <volume>19</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Cathleen</surname>
            <given-names>Wharton</given-names>
          </string-name>
          , John Rieman, Clayton
          <string-name>
            <surname>Lewis</surname>
            , and
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Polson</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>The cognitive walkthrough method: A practitioner's guide</article-title>
          .
          <source>In Usability inspection methods</source>
          .
          <volume>105</volume>
          -
          <fpage>140</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Robert</surname>
            <given-names>H Wortham</given-names>
          </string-name>
          , Andreas Theodorou, and
          <string-name>
            <surname>Joanna</surname>
          </string-name>
          J Bryson.
          <year>2017</year>
          .
          <article-title>Improving robot transparency:real-time visualisation of robot AI substantially improves understanding in naive observers</article-title>
          ,
          <source>In IEEE RO-MAN</source>
          <year>2017</year>
          .
          <article-title>IEEE RO-MAN 2017</article-title>
          . http://opus.bath.ac.uk/55793/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>