<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Taxonomy for Human Subject Evaluation of Black-Box Explanations in XAI</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Chromik</string-name>
          <email>michael.chromik@ifi.lmu.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Schuessler</string-name>
          <email>schuessler@tu-berlin.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LMU Munich</institution>
          ,
          <addr-line>Munich</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Technische Universität Berlin</institution>
          ,
          <addr-line>Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>The interdisciplinary field of explainable artificial intelligence (XAI) aims to foster human understanding of black-box machine learning models through explanation methods. However, there is no consensus among the involved disciplines regarding the evaluation of their efectiveness - especially concerning the involvement of human subjects. For our community, such involvement is a prerequisite for rigorous evaluation. To better understand how researchers across the disciplines approach human subject XAI evaluation, we propose developing a taxonomy that is iterated with a systematic literature review. Approaching them from an HCI perspective, we analyze which study designs scholar chose for diferent explanation goals. Based on our preliminary analysis, we present a taxonomy that provides guidance for researchers and practitioners on the design and execution of XAI evaluations. With this position paper, we put our survey approach and preliminary results up for discussion with our fellow researchers.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Human-centered computing → HCI design and evaluation
methods.
explainable artificial intelligence; explanation; human evaluation;
taxonomy.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>We have witnessed the widespread adoption of intelligent systems
into many contexts of our lives. Such systems are often built on
advanced machine learning (ML) algorithms that enable powerful
predictions – often at the expense of interpretability. As these
systems are introduced into more sensitive contexts of society, there is a
growing acceptance that they need to be capable of explaining their
behavior in human-understandable terms. Hence, much research
is conducted within the emerging domain of explainable artificial
intelligence (XAI) and interpretable machine learning (IML) on
developing models, methods, and interfaces that are interpretable to
human users – often through some notion of explanation.</p>
      <p>
        However, most works focus on computational problems while
limited research efort is reported concerning their user evaluation.
Previous surveys identified the need for more rigid empirical
evaluation of explanations [
        <xref ref-type="bibr" rid="ref17 ref2 ref5">2, 5, 17</xref>
        ]. The AI and ML communities often
strive for functional evaluation of their approaches with
benchmark data to demonstrate generalizability. While this is suitable to
demonstrate technical feasibility, it is also problematic since often
"there is no formal definition of a correct or best explanation" [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
Even if a formal foundation exists, it does not necessarily result in
practical utility for humans as the utility of an explanation is highly
dependent on the context and capabilities of human users.
Without proper human behavior evaluations, it is dificult to assess an
explanation method’s utility for practical use cases [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. We argue
that functional and behavioral evaluation approaches have their
legitimacy. Yet, since there is no consensus on evaluation methods,
the comparison and validation of diverse explanation techniques is
an open challenge [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ].
      </p>
      <p>In this work, we take an HCI perspective and focus on
evaluations with human subjects. We believe that the HCI community
should be the driving force for establishing rigorous evaluation
procedures that investigate how XAI can benefit users. Our work
is guided by three research questions:
• RQ-1: Which evaluation approaches have been proposed
and discussed across disciplines in the field of XAI?
• RQ-2: Which study design decisions have researchers made
in previous evaluations with human subjects?
• RQ-3: How can the proposed approaches and study designs
be integrated into a guiding taxonomy for human-centered
XAI evaluation?</p>
      <p>
        The contribution of this workshop paper is two-fold: First, we
introduce our methodology for taxonomy development and literature
review guided by RQ-1 and RQ-2. The review aims to provide an
overview of how evaluations are currently conducted and help
identify suitable best practices. As a second contribution, we present
a preliminary taxonomy of human evaluation approaches in XAI
and describe its dimensions. Taxonomies have been used in many
disciplines to help researchers and practitioners to understand and
analyze complex domains [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Our overarching goal is to
synthesize a human subject evaluation guideline for researchers and
practitioners of diferent disciplines in the field of XAI. With this
work, we put our review methodology and preliminary taxonomy
up for discussion with our fellow researchers.
2
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>FOUNDATIONS AND RELATED WORK</title>
    </sec>
    <sec id="sec-4">
      <title>Evaluating Explanations in Social Sciences</title>
      <p>
        Miller defines explanation as either a process or a product [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
On the one hand, an explanation describes the cognitive process
of identifying the cause(s) of a particular event. At the same time,
it is a social process between an explainer (sender of an
explanation) and an explainee (receiver of an explanation) with the goal
to transfer knowledge about the cognitive process. Lastly, an
explanation can describe the product that results from the cognitive
process and aims to answer a why-question. In our paper, we refer
to explanations from the product perspective. Psychologists and
social scientists investigated how humans evaluate explanations
for decades. Within their disciplines, explanation evaluation refers
to the process applied by an explainee for determining if an
explanation is satisfactory [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Scholars conducted experiments where
they presented participants with diferent types of explanations as
treatments. These experiments indicate that choosing one
explanation over another is often an arbitrary choice heavily influenced
by cognitive biases and heuristics [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The primary criteria of
explainees are whether the explanation helps them to understand
the underlying cause [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. For instance, humans are more likely to
accept explanations that are consistent with their prior beliefs.
Furthermore, they prefer explanations that are simpler (i.e., with fewer
causes), and more generalizable (i.e., that apply to more events).
Also, the efectiveness of an explanation depends on the current
information needs of the explainee. A suitable explanation for one
purpose may be irrelevant for another. Thus, for an explanation to
be efective, it is essential to know the intended context of use.
2.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Explainable Artificial Intelligence (XAI)</title>
      <p>
        Interpretability in machine learning is not a monolithic concept [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Instead, it is used to indirectly evaluate whether important
desiderata, such as fairness, reliability, causality, or trust, are met in a
particular context [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Some definitions of interpretability are rather
system-centric. Doshi-Velez and Kim [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] describe it as a model’s
"ability to explain or to present in understandable terms to a human."
Miller [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] takes a more human-centered perspective calling it "the
degree to which an observer can understand the cause of a decision".
Human understanding can be fostered either by ofering means of
introspection or through explanations [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A large variety of
methods exist for both approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The term interpretable machine
learning (IML) often refers to research on models and algorithms
that are considered as inherently interpretable while explainable
AI (XAI) often refers to the generation of (post-hoc) explanations
or means of introspection for black-box models [
        <xref ref-type="bibr" rid="ref27 ref33">27, 33</xref>
        ]. A model’s
black-box behavior may manifest itself in two ways: either from
complex architectures, as with deep neural networks, or from
proprietary models (that may otherwise be interpretable), as with the
COMPAS recidivism model [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. The lines between IML and XAI
are often seamless and the terms are often used interchangeably.
For instance, DARPA’s XAI program subsumes both terms with
the objective to "enable human users to understand, appropriately
trust, and efectively manage the emerging generation of artificially
intelligent partners" [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
2.3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Evaluating Explanations in XAI</title>
      <p>
        Multiple surveys of the ever-growing field of XAI exist. They
formalize and ground the concept of XAI [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], relate it to adjacent
concepts and disciplines [
        <xref ref-type="bibr" rid="ref1 ref16">1, 16</xref>
        ], categorize methods [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], or discuss
future research directions [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. All these surveys report a lack of
rigid evaluations. Adadi et al. found that only 5% of surveyed papers
evaluate XAI methods and quantify their relevance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Similarly,
Nunes and Jannach found that 78% of the analyzed papers on
explanations in decision support systems lacked structured evaluations
that go beyond anecdotal "toy examples" [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>
        Some works have addressed the design and conduction of
explanation evaluations in XAI. Gilpin et al. survey explainable methods
for deep neural networks and describe a categorization of evaluation
approaches at diferent stages of the ML development process [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Yang et al. provide a framework consisting of multiple levels of
explanation evaluation [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. Their definition of persuasibility
(measuring the degree of human comprehension) focuses on the human
and resonates with our notion of human subject evaluation. Our
work aims to elaborate on their generic strategy of "employing
users for human studies". Nunes and Jannach reviewed 217
publications spanning multiple decades and briefly report findings from
applied evaluation approaches [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Based on their survey they
derive a comprehensive taxonomy that guides the design of
explanations. However, their taxonomy omits aspects of evaluation.
Mueller identified 39 XAI papers that reported empirical
evaluations and qualitatively described chosen evaluation approaches
along 9 dimensions [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        While these works ofer valuable ideas, they are limited in their
scope and, thus, ofer little guidance for XAI user evaluations. Of
course, "there is no standard design for user studies that evaluate
forms of explanations" [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. However, we believe that a unified
taxonomy is needed that integrates the most common ideas related
to human subject evaluation and extends them with best practice
examples. Such an actionable format can provide great benefit for
researchers and practitioners by guiding them through the design
and reporting of structured XAI evaluations.
3
      </p>
    </sec>
    <sec id="sec-7">
      <title>METHODOLOGY</title>
      <p>In this section, we outline our method of taxonomy development
as well as the planned literature review. Our goal is to develop a
comprehensive taxonomy for human subject evaluations in XAI.
We seek to validate and iterate it through a structured literature
review (SLR). Figure 2 illustrates our proposed methodology and
the interplay between taxonomy and SLR.
3.1</p>
    </sec>
    <sec id="sec-8">
      <title>Taxonomy Development</title>
      <p>
        There are two approaches to constructing a taxonomy. Following
the conceptual-to-empirical approach, the researcher proposes a
classification based on a theory or model (deductive). In contrast,
the empirical-to-conceptual approach derives the taxonomy from
empirical cases (inductive). We follow the iterative process for
taxonomy development proposed by Nickerson et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Their
method unifies both approaches in an iterative process under a
shared meta-characteristic and defined ending conditions.
      </p>
      <p>In line with RQ-3, we defined our meta-characteristic as the
development of a taxonomy for human subject evaluation of black-box</p>
      <sec id="sec-8-1">
        <title>Taxonomy Development</title>
        <p>(Nickerson et al.)
Meta-Characteristic:
Taxonomy for human-subject evaluation of black-box
explanations that guides researchers and practitioners
with the design and reporting of future studies
Ending condition according to Nickerson et al.</p>
        <p>Determine
meta-characteristic
and ending conditions
Conceptual-to-empirical
approach</p>
        <p>Empirical-to-conceptual</p>
        <p>approach
Preliminary
taxonomy
Taxonomy
meeting
ending conditions</p>
      </sec>
      <sec id="sec-8-2">
        <title>Structured Literature Review</title>
        <p>(Kitchenham and Charters)
Exclusion Criteria:
EC-1: Not written in English; EC-2: Not related to
blackbox explanations; EC-3: Not reporting human subject
evaluation; EC-4: Full-text could not be retrieved; EC-5:
Not a scientific full- or short paper; EC-6: Is a duplicate
Inclusion Criterion IC-1: Reports setup and results of a
human subject evaluation in the XAI</p>
        <p>Publications identified
through Scopus</p>
        <p>(n=653)
Publications screened
based on abstract and</p>
        <p>full-text (n=653)
Publications analyzed
based on abstract and</p>
        <p>full-text (n=146)
Publications meeting
inclusion criteria
(n=133)</p>
        <p>Publications excluded
after screening</p>
        <p>(n=507)
Publications excluded
after detailed analysis
(n=13)
explanations that guides researchers and practitioners with the design
and reporting of future studies. We start by applying the
conceptualto-empirical approach. To follow this approach, one needs to
propose a classification based on a theory or model. We do this by
consolidating proposed categories for XAI evaluation in prior work
and connecting them with foundational literature on empirical
studies. The resulting taxonomy describes an ideal type, which
allows us to examine empirically how much current human subject
evaluations deviate from an ideal type.
3.2</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Structured Literature Review</title>
      <p>
        As part of the empirical-to-conceptual iteration, we aim to
validate and iterate the taxonomy using a structured literature review
(SLR). In line with RQ-2, the review’s objective is to capture how
researchers currently evaluate XAI methods and systems with
human subjects. Through this, we seek to find out how structured
and precise we can describe the field using our taxonomy. During
this process, we also aim to iterate the taxonomy. The planned
SLR follows established approaches proposed by Kitchenham and
Charters [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. In the following, we outline the proposed search
strategy.
      </p>
      <p>Source Selection: An exploratory search for XAI on Google Scholar
indicated that relevant work is dispersed across multiple publishers,
conferences, and journals. Thus, we use the Scopus database as a
source as it integrates publications from relevant publishers such
as ACM, IEEE, and AAAI.</p>
      <p>Search Query: Through our exploratory search, we obtained an
initial understanding of relevant keywords, synonyms, and related
concepts that helped us to construct a search query. We found that
diferent terms are used between the disciplines to describe the field
of XAI and human subject evaluation approaches. Early research
does not explicitly state the expressions XAI nor explainable
artificial intelligence. Thus, our search queries are composed of groups
and terms. Groups refer to a specific aspect of the research question
and limit the search scope. Terms have a similar semantic meaning
within the group domain or are often used interchangeably. We are
interested in the intersection of 3 groups that can be phrased using
diferent terms. Table 1 shows our used groups and terms.</p>
      <p>
        Study Selection Criteria: We filtered the search results by six
exclusion criteria (EC) and one inclusion criterion (IC). We are interested
in primary studies that report the setup and result of human subject
evaluations in the XAI context (IC-1). We limit the survey to
publications addressing the black-box explanation problem, according to
Guidotti et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] (EC-2). Furthermore, we exclude publications that
do not report human-grounded or application-grounded evaluations
according to Doshi-Velez and Kim [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (EC-3). We applied the
exclusion criteria in cascading order, i.e., if we excluded publications
due to one EC, we did not assess any following criteria.
      </p>
      <p>Study Analysis: So far, we conducted the search procedure for
Scopus in September 2019, which returned a total of 653 potentially
relevant publications. Both authors filtered the returned
publications by the inclusions and exclusion criteria to control for
interrater efects. We discussed difering assessments until we reached
consensus. We are currently in the process of analyzing the
publications that met the inclusion criterion.
4</p>
    </sec>
    <sec id="sec-10">
      <title>TAXONOMY OF HUMAN SUBJECT</title>
    </sec>
    <sec id="sec-11">
      <title>EVALUATION IN XAI</title>
      <p>In the following section, we describe relevant dimensions of
blackbox explanation evaluation with human subjects. We group
identiifed characteristics into task-related, participant-related, and study
design-related dimensions. The outlined taxonomy is a
preliminary result after the first iterations of the conceptual-to-empirical
approach based on propositions in prior work. Furthermore, the
taxonomy was validated and refined based on a small subset consisting
of 34 publications from the structured literature review following
the empirical-to-conceptual approach.
4.1</p>
    </sec>
    <sec id="sec-12">
      <title>Task Dimensions</title>
      <p>
        Mohseni and Ragan distinguish two types of human
involvement in the evaluation of explanations [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In the feedback
setting, participants provide feedback on actual explanations.
Experimenters determine the quality of the explanations through this
feedback. In contrast, in the feed-forward setting no explanations are
provided. Instead, humans are generating examples of reasonable
explanations serving as a benchmark for algorithmic explanations.
      </p>
      <p>
        Doshi-Velez and Kim distinguish two types of human subject
evaluations that difer in their level of task abstraction [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]:
Application-grounded evaluations conduct experiments within a real
application context. Typically, this requires a high level of participant
expertise. The quality of the explanation is assessed in measures
of the application context, typically with a test of performance.
Human-grounded evaluations conduct simplified or abstracted
experiments that aim to maintain the essence of the target application.
      </p>
      <p>
        Multiple types of user tasks have been proposed to elicit the
quality of explanations [
        <xref ref-type="bibr" rid="ref18 ref33 ref4">4, 18, 33</xref>
        ]. We suggest distinguishing them
by the information provided to the participant and the information
inquired in return. In verification tasks, participants are provided
with input, explanation, and output and asked for their
satisfaction with the explanation. Forced choice tasks extend this setting.
Here, participants are asked to choose from multiple competing
explanations. In the case of forward simulation tasks, participants
are presented with inputs as well as explanations and need to
predict the system’s output. Counterfactual simulation tasks, present
participants with an input, an explanation, an output, and an
alternative output (the counterfactual). Based on these, they predict
what input changes are necessary to obtain the alternative output.
In "Clever Hans" detection tasks, participants need to identify and
possibly debug flawed models, e.g., a naive or short-sighted
predictor [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. System usage tasks are characterized by participants
using the system and its explanations for its primary purpose, e.g.,
a decision-making situation. The quality of the explanation is
assessed in terms of decision quality. In annotation tasks, participants
provide a suitable explanation given input and output of a model.
      </p>
      <p>
        Explanations are provided to users with very diferent goals in
mind. For their efective evaluation, researchers need to ensure that
the intended explanation goal(s) are aligned with their intended
evaluation goal(s), and vice versa. Also, calibration of the
individual goals of participants with the intended explanation goal(s)
might be necessary (e.g., through a briefing before the task) [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ].
We distinguish 9 common explanation goals, which are derived
from [
        <xref ref-type="bibr" rid="ref24 ref30 ref32">24, 30, 32</xref>
        ]: transparency aims to explain how the system
works, scrutability aims to allow users to tell the system it is wrong,
trust aims to increase the user’s confidence in the system,
persuasiveness aims to convince the user to perform an action, satisfaction
aim to increase the ease of use or enjoyment, efectiveness aims to
help users make good decisions, eficiency aims to make decisions
faster, education aims to enable users to generalize and learn,
debugging aims to enable users to identify defects in the system. In
the case of multiple intended explanation goals, their dependencies
may be complementary, contradictory, or even unknown (e.g., the
impact of transparency on trust).
      </p>
      <p>
        Hofman et al. describe multiple levels of task evaluation to
assess a participant’s understanding of and XAI system. Furthermore,
they discuss suitable metrics for each level [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Tests of satisfaction
measure participants’ self-reported satisfaction with an
explanation and their perception of system understanding. On this level,
researchers can rarely be sure whether participants understand
the system to the degree that participants claim. Tests of
comprehension assess the participants’ mental models of the system and
tests their understanding, for example, through prediction tests and
generative exercises. Tests of performance measure the resulting
human-XAI system performance.
Intended Explanation Goal [
        <xref ref-type="bibr" rid="ref24 ref30 ref32">24, 30, 32</xref>
        ]
Study Approach
      </p>
      <p>Treat. Assignment</p>
      <p>
        Treat. Combination [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
      </p>
      <p>Information given to Participant
Qualitative
Quantitative
Mixed</p>
      <p>
        Input

 , ?





Participant Type [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]
(AI) Novice User
Domain Expert
AI Expert
      </p>
      <p>Within-subjects</p>
      <p>Between-subjects
Explanation</p>
      <p>
, … , 




?</p>
      <p>Output


?
 , 



Level of Expertise</p>
      <p>AI
low
low
high</p>
      <p>Domain
low
high
low</p>
      <p>
        Study Design Dimensions
Single Explanation
With and Without Explanation
Altern. Explanation
Altern. Explanation Interface
Participant Incentivation
[
        <xref ref-type="bibr" rid="ref25 ref28 ref29">28, 29, 25</xref>
        ]
Monetary
Non-Monetary
Number of Participants
Low
High
Participant Recruiting
Field Study
Lab Study
Online Study
Crowd-sourcing
      </p>
      <p>Participant Dimensions</p>
      <p>
        Participant Foresight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
Human-grounded
Application-grounded
      </p>
      <p>Intrinsic</p>
      <p>Extrinsic</p>
    </sec>
    <sec id="sec-13">
      <title>4.2 Participant Dimensions</title>
      <p>
        Mohseni et al. distinguish between several participant types: AI
novices who are usually end-users, data experts (including domain
experts), and AI experts [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This distinction is important as user
expertise strongly influences other participant-related dimensions.
For example, Doshi-Velez and Kim [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], referencing the work of
Neath and Surprenant [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], point out that user expertise determines
what kind of cognitive chunks participants apply to a situation. The
expertise of participants may determine the recruiting method
and number of participants. Recruiting dificulty is likely to
increase with the required level of participants’ expertise [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. One can
recruit novices in large numbers via crowd-sourcing. In contrast,
domain or AI experts are usually harder to identify and recruit. They
are often invited to a targeted online study, a lab study, or a field
study. According to Narayana et al., the user study task may have
dependencies with the level of participant foresight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. In an
intrinsic setting, the participant’s understanding of the context is
solely based on the provided information. Thus, all participants are
assumed to have equal knowledge about the context. Such types of
experiments are usually suitable for novices. In an extrinsic setting,
participants can additionally draw upon external facts, such as prior
experience, that may be relevant for assessing the quality of an
explanation, e.g., for spotting model flaws. Such a setting may be
more suitable for data experts. However, it also makes controlling
for participants’ knowledge more dificult.
      </p>
      <p>
        Incentivization of participants is another relevant dimension.
According to Sova and Nielsen, it should be chosen considering
study length, task demand, and participant expertise [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
Stadtmüller and Porst advise us to use a monetary incentive for
participants [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. However, several non-monetary incentives are known to
be efective as well (e.g., gifts for already paid employees) [
        <xref ref-type="bibr" rid="ref25 ref28">25, 28</xref>
        ].
Prost and Briel found that participants may take part in a study
because of study-related incentives (e.g., curiosity, sympathy, or
entertainment), personal-incentive (e.g., professional interest or a
promise made), or altruistic reasons (e.g., to benefit science, society,
or others) [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Esser argues that researchers should consider
incentives in their combination such that the benefits of participating
out-weigh the perceived cost [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
    </sec>
    <sec id="sec-14">
      <title>4.3 Study Design Dimensions</title>
      <p>
        The study design of evaluations may follow a qualitative,
quantitative, or mixed study approach. In experimental studies,
experimenters assign treatments to groups of participants. Applied to the
context of explanation evaluations, we can distinguish four
common types of treatments combinations in line with Nunes and
Jannach [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]: single treatment (i.e., no alternative treatment), with
and without explanation (i.e., no explanation is alternative
treatment), alternative explanation (i.e., varying information provided
in explanations between treatments with other aspects of user
interface fixed), alternative explanation interface (i.e., varying user
interfaces between treatments). Furthermore, we can distinguish
study designs by the treatment assignment: Between-subjects
designs study the diferences in understanding between groups of
participants, each usually assigned to one treatment. In contrast,
within-subject designs study diferences within individual
participants who are assigned to multiple treatments.
5
      </p>
    </sec>
    <sec id="sec-15">
      <title>LIMITATION AND FUTURE WORK</title>
      <p>
        Our preliminary taxonomy has limitations. The taxonomy is neither
collectively exhaustive nor mutually exclusive. Thus, it does not
meet the ending conditions of taxonomy development [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. We
aim to refine and iterate the taxonomy with the results from the
proposed structured literature review.
      </p>
      <p>
        Furthermore, human subject evaluations in XAI are typically
embedded in a broader context, which may create dependencies
and limit applicable evaluation approaches. Dependencies may
arise from the explanation design context, such as the form of
an explanation, its contents, or its underlying generation method.
Multiple taxonomies have been developed for guiding the design
of explanations [
        <xref ref-type="bibr" rid="ref24 ref7">7, 24</xref>
        ]. Nunes and Jannach proposed an elaborate
explanation design taxonomy [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. However, their taxonomy omits
aspects of evaluation. For now, we have abstained from relating our
preliminary human subject evaluation taxonomy with this prior
work, but plan to integrate them in later iterations.
6
      </p>
    </sec>
    <sec id="sec-16">
      <title>CONCLUSION</title>
      <p>
        In this work, we gave a brief overview of recent eforts on
explanation evaluation with human subjects in the growing field of XAI.
We proposed a methodology for developing a comprehensive
taxonomy for human subject evaluation that integrates the knowledge
from multiple disciplines involved in XAI. Based on ideas from
prior work, we presented a preliminary taxonomy following the
conceptual-to-empirical approach. Despite its limitations, we
believe our work is a starting point for rigorously evaluating the
utility of explanations for human understanding of XAI systems.
Researchers and practitioners developing XAI explanation facilities
and systems have been asked to "respect the time and efort involved
to do such evaluations" [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We aim to spark a discussion at the
workshop on how to support them along the way.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ashraf</given-names>
            <surname>Abdul</surname>
          </string-name>
          , Jo Vermeulen, Danding Wang,
          <string-name>
            <surname>Brian</surname>
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lim</surname>
            , and
            <given-names>Mohan</given-names>
          </string-name>
          <string-name>
            <surname>Kankanhalli</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Trends and Trajectories for Explainable, Accountable and Intelligible Systems: An HCI Research Agenda</article-title>
          .
          <source>In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (Montreal QC, Canada) (CHI '18)</source>
          . ACM, New York, NY, USA, Article
          <volume>582</volume>
          , 18 pages. https://doi.org/10.1145/3173574.3174156
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Adadi</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Berrada</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI)</article-title>
          .
          <source>IEEE Access</source>
          <volume>6</volume>
          (
          <year>2018</year>
          ),
          <fpage>52138</fpage>
          -
          <lpage>52160</lpage>
          . https://doi.org/10.1109/ACCESS.
          <year>2018</year>
          .2870052
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Or</given-names>
            <surname>Biran</surname>
          </string-name>
          and
          <string-name>
            <given-names>Courtenay</given-names>
            <surname>Cotton</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Explanation and justification in machine learning: A survey</article-title>
          .
          <source>In IJCAI-17 workshop on explainable AI (XAI)</source>
          , Vol.
          <volume>8</volume>
          . 1.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Finale</given-names>
            <surname>Doshi-Velez</surname>
          </string-name>
          and
          <string-name>
            <given-names>Been</given-names>
            <surname>Kim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards A Rigorous Science of Interpretability</article-title>
          .
          <source>CoRR abs/1702</source>
          .08608 (
          <year>2017</year>
          ). arXiv:
          <volume>1702</volume>
          .08608 http://arxiv.org/abs/ 1702.08608
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F. K.</given-names>
            <surname>Dosilovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brcic</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Hlupic</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Explainable artificial intelligence: A survey</article-title>
          .
          <source>In 2018 41st International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO)</source>
          .
          <volume>0210</volume>
          -
          <fpage>0215</fpage>
          . https://doi.org/ 10.23919/MIPRO.
          <year>2018</year>
          .8400040
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Hartmut</given-names>
            <surname>Esser</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Über die Teilnahme an Befragungen</article-title>
          .
          <source>ZUMA Nachrichten</source>
          <volume>10</volume>
          ,
          <issue>18</issue>
          (
          <year>1986</year>
          ),
          <fpage>38</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Friedrich</surname>
          </string-name>
          and
          <string-name>
            <given-names>Markus</given-names>
            <surname>Zanker</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A Taxonomy for Generating Explanations in Recommender Systems</article-title>
          .
          <source>AI Magazine</source>
          <volume>32</volume>
          ,
          <issue>3</issue>
          (Jun.
          <year>2011</year>
          ),
          <fpage>90</fpage>
          -
          <lpage>98</lpage>
          . https://doi.org/10.1609/aimag.v32i3.
          <fpage>2365</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Gilpin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bajwa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Specter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Kagal</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Explaining Explanations: An Overview of Interpretability of Machine Learning</article-title>
          .
          <source>In 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA)</source>
          .
          <volume>80</volume>
          -
          <fpage>89</fpage>
          . https://doi.org/10.1109/DSAA.
          <year>2018</year>
          .00018
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Riccardo</given-names>
            <surname>Guidotti</surname>
          </string-name>
          , Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and
          <string-name>
            <given-names>Dino</given-names>
            <surname>Pedreschi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A survey of methods for explaining black box models</article-title>
          .
          <source>Comput. Surveys</source>
          <volume>51</volume>
          ,
          <issue>5</issue>
          (aug
          <year>2018</year>
          ). https://doi.org/10.1145/3236009
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>David</given-names>
            <surname>Gunning</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Aha</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>DARPA's Explainable Artificial Intelligence (XAI) Program</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>40</volume>
          ,
          <issue>2</issue>
          (Jun.
          <year>2019</year>
          ),
          <fpage>44</fpage>
          -
          <lpage>58</lpage>
          . https://doi.org/10.1609/ aimag.v40i2.
          <fpage>2850</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Robert</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Hofman</surname>
            , Shane T. Mueller, Gary Klein, and
            <given-names>Jordan</given-names>
          </string-name>
          <string-name>
            <surname>Litman</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Metrics for Explainable AI: Challenges and Prospects</article-title>
          . CoRR abs/
          <year>1812</year>
          .04608 (
          <year>2018</year>
          ). arXiv:
          <year>1812</year>
          .04608 http://arxiv.org/abs/
          <year>1812</year>
          .04608
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Frank</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Keil</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Explanation and Understanding</article-title>
          .
          <source>Annual Review of Psychology 57</source>
          ,
          <issue>1</issue>
          (
          <year>2006</year>
          ),
          <fpage>227</fpage>
          -
          <lpage>254</lpage>
          . https://doi.org/10.1146/annurev.psych.
          <volume>57</volume>
          .102904.190100 arXiv:https://doi.org/10.1146/annurev.psych.
          <volume>57</volume>
          .102904.190100 PMID:
          <fpage>16318595</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kitchenham</surname>
          </string-name>
          and
          <string-name>
            <given-names>S</given-names>
            <surname>Charters</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Guidelines for performing Systematic Literature Reviews in Software Engineering</article-title>
          . Keele University and University of Durham,
          <source>Technical Report EBSE-2007-01</source>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Lapuschkin</surname>
          </string-name>
          , Stephan Wäldchen, Alexander Binder, Grégoire Montavon, Wojciech Samek, and
          <string-name>
            <surname>Klaus-Robert Müller</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Unmasking Clever Hans predictors and assessing what machines really learn</article-title>
          .
          <source>Nature communications 10</source>
          ,
          <issue>1</issue>
          (
          <year>2019</year>
          ),
          <fpage>1096</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Zachary</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lipton</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The Mythos of Model Interpretability</article-title>
          .
          <source>Queue</source>
          <volume>16</volume>
          ,
          <issue>3</issue>
          ,
          <string-name>
            <surname>Article 30</surname>
          </string-name>
          (
          <year>June 2018</year>
          ),
          <volume>27</volume>
          pages. https://doi.org/10.1145/3236386.3241340
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Tim</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Explanation in artificial intelligence: Insights from the social sciences</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>267</volume>
          (
          <year>2019</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          . https://doi.org/10.1016/j.artint.
          <year>2018</year>
          .
          <volume>07</volume>
          .007
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Tim</surname>
            <given-names>Miller</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Piers</given-names>
            <surname>Howe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Liz</given-names>
            <surname>Sonenberg</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <string-name>
            <surname>Explainable</surname>
            <given-names>AI</given-names>
          </string-name>
          :
          <article-title>Beware of Inmates Running the Asylum</article-title>
          .
          <source>In IJCAI 2017 Workshop on Explainable Artificial Intelligence (XAI)</source>
          . http://people.eng.unimelb.edu.au/tmiller/pubs/explanationinmates.pdf
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Sina</given-names>
            <surname>Mohseni</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eric D.</given-names>
            <surname>Ragan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning</article-title>
          . CoRR abs/
          <year>1801</year>
          .05075 (
          <year>2018</year>
          ). arXiv:
          <year>1801</year>
          .05075 http://arxiv.org/abs/
          <year>1801</year>
          .05075
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Sina</surname>
            <given-names>Mohseni</given-names>
          </string-name>
          , Niloofar Zarei, and
          <string-name>
            <given-names>Eric D.</given-names>
            <surname>Ragan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A Survey of Evaluation Methods and Measures for Interpretable Machine Learning</article-title>
          . CoRR abs/
          <year>1811</year>
          .11839 (
          <year>2018</year>
          ). arXiv:
          <year>1811</year>
          .11839 http://arxiv.org/abs/
          <year>1811</year>
          .11839
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Shane</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <surname>Robert R. Hofman</surname>
            ,
            <given-names>William J</given-names>
          </string-name>
          . Clancey, Abigail Emrey, and
          <string-name>
            <given-names>Gary</given-names>
            <surname>Klein</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI</article-title>
          . CoRR abs/
          <year>1902</year>
          .
          <year>01876</year>
          (
          <year>2019</year>
          ). arXiv:
          <year>1902</year>
          .
          <year>01876</year>
          http://arxiv.org/abs/
          <year>1902</year>
          .01876
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Menaka</surname>
            <given-names>Narayanan</given-names>
          </string-name>
          , Emily Chen, Jefrey He, Been Kim, Sam Gershman, and
          <string-name>
            <surname>Finale</surname>
          </string-name>
          Doshi-Velez.
          <year>2018</year>
          .
          <article-title>How do Humans Understand Explanations from Machine Learning Systems? An Evaluation of the Human-Interpretability of Explanation</article-title>
          . CoRR abs/
          <year>1802</year>
          .00682 (
          <year>2018</year>
          ). arXiv:
          <year>1802</year>
          .00682 http://arxiv.org/abs/
          <year>1802</year>
          .00682
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Ian</given-names>
            <surname>Neath</surname>
          </string-name>
          and
          <string-name>
            <given-names>Aimee</given-names>
            <surname>Surprenant</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Human Memory (2 edition ed</article-title>
          .). Thomson/Wadsworth, Australia ; Belmont, CA.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Robert</surname>
            <given-names>C Nickerson</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Upkar Varshney</surname>
            , and
            <given-names>Jan</given-names>
          </string-name>
          <string-name>
            <surname>Muntermann</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A method for taxonomy development and its application in information systems</article-title>
          .
          <source>European Journal of Information Systems</source>
          <volume>22</volume>
          ,
          <issue>3</issue>
          (
          <year>2013</year>
          ),
          <fpage>336</fpage>
          -
          <lpage>359</lpage>
          . https://doi.org/10.1057/ejis.
          <year>2012</year>
          .
          <volume>26</volume>
          arXiv:https://doi.org/10.1057/ejis.
          <year>2012</year>
          .26
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Ingrid</given-names>
            <surname>Nunes</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dietmar</given-names>
            <surname>Jannach</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Systematic Review and Taxonomy of Explanations in Decision Support and Recommender Systems</article-title>
          .
          <source>User Modeling and User-Adapted Interaction 27</source>
          ,
          <fpage>3</fpage>
          -
          <lpage>5</lpage>
          (
          <issue>Dec</issue>
          .
          <year>2017</year>
          ),
          <fpage>393</fpage>
          -
          <lpage>444</lpage>
          . https://doi.org/10. 1007/s11257-017-9195-0
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Rolf</given-names>
            <surname>Porst</surname>
          </string-name>
          and Christa von Briel.
          <year>1995</year>
          .
          <article-title>Wären Sie vielleicht bereit, sich gegebenenfalls noch einmal befragen zu lassen? Oder: Gründe für die Teilnahme an Panelbefragungen</article-title>
          . Vol.
          <year>1995</year>
          /04.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Forough</surname>
          </string-name>
          Poursabzi-Sangdeh, Daniel G. Goldstein,
          <string-name>
            <surname>Jake M. Hofman</surname>
          </string-name>
          , Jennifer Wortman Vaughan, and
          <string-name>
            <surname>Hanna</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Wallach</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Manipulating and Measuring Model Interpretability</article-title>
          . CoRR abs/
          <year>1802</year>
          .07810 (
          <year>2018</year>
          ). arXiv:
          <year>1802</year>
          .07810 http: //arxiv.org/abs/
          <year>1802</year>
          .07810
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Cynthia</given-names>
            <surname>Rudin</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead</article-title>
          .
          <source>Nature Machine Intelligence</source>
          <volume>1</volume>
          ,
          <issue>5</issue>
          (
          <year>2019</year>
          ),
          <fpage>206</fpage>
          -
          <lpage>215</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Deborah</given-names>
            <surname>Hinderer</surname>
          </string-name>
          Sova and
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Nielsen</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>How to Recruit Participants for Usability Studies</article-title>
          . https://www.nngroup.com/reports/how-to
          <article-title>-recruitparticipants-usability-studies/</article-title>
          ,
          <source>accessed December 20th</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Sven</given-names>
            <surname>Stadtmüller</surname>
          </string-name>
          and
          <string-name>
            <given-names>Rolf</given-names>
            <surname>Porst</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Zum Einsatz von Incentives bei postalischen Befragungen</article-title>
          . Vol.
          <volume>14</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Nava</given-names>
            <surname>Tintarev</surname>
          </string-name>
          and
          <string-name>
            <given-names>Judith</given-names>
            <surname>Masthof</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A Survey of Explanations in Recommender Systems</article-title>
          .
          <source>In Proceedings of the 2007 IEEE 23rd International Conference on Data Engineering Workshop (ICDEW '07)</source>
          . IEEE Computer Society, Washington, DC, USA,
          <fpage>801</fpage>
          -
          <lpage>810</lpage>
          . https://doi.org/10.1109/ICDEW.
          <year>2007</year>
          .4401070
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Nadya</surname>
            <given-names>Vasilyeva</given-names>
          </string-name>
          , Daniel A Wilkenfeld, and
          <string-name>
            <given-names>Tania</given-names>
            <surname>Lombrozo</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Goals Afect the Perceived Quality of Explanations.</article-title>
          .
          <source>In CogSci.</source>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Danding</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qian Yang</surname>
            ,
            <given-names>Ashraf</given-names>
          </string-name>
          <string-name>
            <surname>Abdul</surname>
          </string-name>
          , and Brian Y Lim.
          <year>2019</year>
          .
          <article-title>Designing Theory-Driven User-Centric Explainable AI</article-title>
          .
          <source>In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>601</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Fan</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mengnan</given-names>
            <surname>Du</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xia</given-names>
            <surname>Hu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Evaluating Explanation Without Ground Truth in Interpretable Machine Learning</article-title>
          . CoRR abs/
          <year>1907</year>
          .06831 (
          <year>2019</year>
          ). arXiv:
          <year>1907</year>
          .06831 http://arxiv.org/abs/
          <year>1907</year>
          .06831
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>