<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Efect of Explanation Styles on User's Trust</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Retno Larasati</string-name>
          <email>retno.larasati@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna De Liddo</string-name>
          <email>anna.deliddo@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Enrico Motta</string-name>
          <email>enrico.motta@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Knowledge Media Institute, The Open University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>This paper investigates the efects that diferent styles of textual explanation have on explainee's trust in an AI medical support scenario. From the literature, we focused on four diferent styles of explanation: contrastive, general, truthful, and thorough. We conducted a user study in which we presented explanations of a ifctional mammography diagnosis application system to 48 nonexpert users. We carried out a between-subject comparison between four groups of 11-13 people each looking at a diferent explanation style. Our findings suggest that contrastive and thorough explanations produce higher personal attachment trust scores compared to general explanation style, while truthful explanation shows no difference compared to the rest of explanations. This means that users who received contrastive and thorough explanation types found the explanation given significantly more agreeable and suiting their personal taste. These findings, even though not conclusive, confirm the impact of explanation style on users trust towards AI systems and may inform future explanation design and evaluation studies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        One of the main arguments motivating Explainable Artificial
Intelligence research is that the explicability of AI systems can improve
people’s trust and adoption of AI solutions [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Still, the
relationships between trust and explanation is complex, and it is not
always the case that explicability improves users’ trust. Trust in
AI systems is claimed to be enhanced by transparency [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and
understandability [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In order to gain understandability, an AI
system should provide explanations that are meaningful to the
explainee(someone who received explanation). Providing meaningful
explanations could then support users to appropriately calibrate
trust, by improving trust (when they tend to down-trust the system)
and mitigating over-trust issues [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. Previous research has shown
that a key role in calibrating trust can be played by the way in which
explanation is expressed and presented to the users. Explanation
style and modalities afect users’ trust toward algorithmic systems
sometime improving sometime reducing trust [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ][
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. This paper
aims to investigate the relation between explanation and trust by
exploring diferent explanation styles. We first conducted a
literature review in psychology, philosophy, and information systems,
to understand what are the characteristics of meaningful
explanations. We then designed several styles of explanation based on
these characteristics. Since we are interested in assessing the efects
of explanation styles on users’ trust, we also defined a variety of
trust components to measure users’ trust levels. Our proposed trust
measurement was gathered from the literature in human factors
and HCI research. Finally we carried out a user study to see if
any specific explanation style diferently afects users’ trust. Our
contribution is twofold:
(1) we provide evidence which confirms the efect of explanation
styles on diferent trust factors;
(2) we propose a reliable human-AI trust measurement
(Cronbach’s α =0.88) to investigate explanation and trust in
healthcare.
      </p>
      <p>Thie rest of the paper is organized as follows: Section 2 introduces
the context of this research and summarises the relevant literature.
Section 3 describes the methodology of this study. Section 4 and 5
presents and analyse the results from the study. Finally, Section 6
discusses the limitation of this work and outlines the next steps of
the research.</p>
    </sec>
    <sec id="sec-3">
      <title>BACKGROUND AND RELATED WORK</title>
    </sec>
    <sec id="sec-4">
      <title>Explanation</title>
      <p>
        Explanation can be seen as an act or a product and can be
categorised as good or bad. A good explanation is an explanation that
feels right because ofers a phenomenologically familiar sense of
understanding [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this paper, we focus on meaningful
explanation, to stress our interest and focus on the explanation’s capability
to improve understanding and sense-making of AI and algorithmic
results. As such good explanations are not explanations that
necessarily improve trust, and can afect user’s trust both ways, by either
improving or moderating trust.
      </p>
      <p>
        We might ask what is meaningful explanation? There is no single
definition of meaningful explanation. Guidotti et al. defined
meaningful explanation as explanation that is faithful and interpretable
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Thirumuruganathan et al. defined meaningful explanation as
explanation that is personalised based on users’ demographic [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
Regulators have also mentioned meaningful explanation. GDPR
Articles 13–15 state that users have the right to receive ‘meaningful
information about the logic involved’ in automated decisions, but
it fails to provide any specific definition of what is to be considered
’meaningful information’. In this paper, we will refer to meaningful
explanation as explanation that is understandable.
      </p>
      <p>
        In cognitive psychology, explanation can be classified into
different types: i. Causal explanation, which tells you what causes
what, ii. Mechanical explanation, which tells you how a certain
phenomenon comes about, and iii. Personal explanation which tells
you what causes what in the context of personal reasons or beliefs
[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. Approaching these definitions from an explainable AI and AI
reasoning angle, we could say that causal and mechanical
explanation could be the same, because the causal explanation of an AI
system is mechanical by definition. For instance, if we ask why the
AI system gives us a certain prediction, the answer will consist of
an illustration of the AI’s mechanical process, which produced that
prediction result. Personal explanation might also not be relevant,
since all AI "personal" explanations are defined in terms of what
causes what in the context of a specific AI reasoning mechanism.
Therefore, in what follows we will focus on causal explanation.
      </p>
      <p>
        Hilton proposed that causal explanation proceeds through the
operation of counterfactual and contrastive criteria [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Lipton
suggested that "to explain why P rather than Q, we must cite a causal
diference between P and not-Q, consisting of a cause of P and the
absence of a corresponding event in the history of not-Q” [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Miller
quoted Lipton and argued that everyday explanations, or human
explanations, are “sought in response to particular counterfactual
cases. [...]people do not ask why event P happened, but rather why
event P happened instead of some event Q” [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        Causal explanation happens through several processes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. First,
there is information collection: a person gathers the information
available. Second, a causal diagnosis takes place: a person tries
to identify a connection between two events/instances based on
the information. Third, there is causal selection, a person dignifies
a set of conditions as "the explanation". This selection process is
influenced by the information gathered and the domain knowledge
of a person [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. This means that what people consider acceptable
and understandable is selected from the information provided and
depends on people’s own domain knowledge or role. According to
Lambrozo, explanations that are simpler are judged more likely to be
believed and more valuable [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and another study also highlighted
that users prefer a combination of simple and broad explanations
[
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        As mentioned previously, explanation can be seen as an act or
can be seen as a product. Explanation as an act involves the
interaction between one or more explainer and explainee [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. According
to Hilton, explanation is understandable only when it involves
explainer and explainee engaging in information exchange through
dialogue, visual representation, or other communication modalities
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. This statement implies that static explanations could be harder
to understand because they could be less engaging and would not
involve a dynamic interchange between explainer and explainee. To
achieve meaningful explanation, a social (interactive) characteristic
of explanation needs to be taken into account.
      </p>
      <p>
        Previous research also showed that participants place the highest
trust in explanations that are sound and complete [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Soundness
here means nothing but the truth, how truthful each element in an
explanation is with respect to the underlying system. Completeness
here means the whole truth, the extent to which an explanation
describes all of the underlying system. Completeness is argued
to positively afect user understandability [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Even though both
of Kulesza’s studies used explanation in the case of a music
recommender system, we think that being truthful (soundness) and
thorough (completeness) are key characteristics of explanations to
be further explored. Building on the literature reviewed above, we
therefore distilled 6 key characteristics of meaningful explanation,
that are defined in Table 1.
2.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Explanation and User’s Trust</title>
      <p>
        There is arguably a relation between explanation and users’ trust.
According to the Defense Advanced Research Projects Agency
(DARPA), Explainable AI is essential to enable human users to
understand and appropriately trust a machine learning system [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Previous studies proposing diferent types of explanation [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ][
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
further cemented the claim that explanations improves user trust
[
        <xref ref-type="bibr" rid="ref30">30</xref>
        ][
        <xref ref-type="bibr" rid="ref24">24</xref>
        ][
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        However, users’ trust could be misplaced and lead to over-reliance
or over-trust. In a healthcare scenario, a doctor could unknowingly
trust a technologically complex laboratory diagnostic test that
incorrectly calibrated and misdiagnosed patients [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Previous research
suggests that giving explanation could help users to moderate their
trust level [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ], either by providing explanation as system’s
accuracy [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ][
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] or as system’s confidence level [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. On one hand,
these findings are not applied to healthcare. Hence, while system’s
accuracy and system’s confidence level might be highly afecting
users’ trust in dating app [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], or context aware app [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], it is
unclear if that would be the case in a healthcare scenario. On the
other hand, in the healthcare/medical domain, Bussone et al. found
that a high system’s confidence level had only a slight efect on
over-reliance [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        There are a number of ways to present an explanation. For
example, a study mentioned above, used accuracy level as explanation. It
is important to know, what kind of style we are going to present our
explanation. Research found that explanation style and modalities
afect users trust toward algorithmic systems, with the result that
this can either improve or decrease [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ][
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>In addition, in each of the reviewed studies trust was measured
diferently, hence the results are hard to compare and do not provide
a clear picture of the extent to which diferent styles of
explanation afect diferent types of trust. To better understand users’ trust
towards an AI medical system, a more comprehensive trust
measurement instrument is needed and will be explored in the next
section.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Trust Measurement</title>
      <p>In general, there is quite a large literature presenting scales for
measuring trust. This paper will focus on identifying an appropriate
scale for the assessment of human trust in a machine prediction
system, which can be contextualised to a healthcare scenario.</p>
      <p>
        Some of the trust measurements reviewed from the automation
literature are highly specific to particular application contexts. For
example, the scale developed by Schaefer [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] refers specifically to
the context of human reliance on a robot. The questions that are
asked to users to measure trust are, for example: "Does it act as part
of a team?" and "Is it friendly?". Another example of specific trust
measurement is the scale developed by Dzindolet, et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It was
created in the context of aerial terrain photography, showing images
to detect camouflaged soldiers. The questions asked to measure
trust in this case are for example: "How many errors do you think
you will make during the 200 trials?". As these questions are very
specific to the task and the technical knowledge of the users in the
specific application context, it would be hard to translate them to a
healthcare scenario.
      </p>
      <p>
        Madsen and Gregor [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] developed and tested a more generic
human-computer trust measurement instrument, with the focus on
trust in an intelligent decision aid. A validity analysis conducted
of this instrument showed high Cronbach’s alpha results, which
makes this scale promising to be tested in a diferent application
ifeld. Trust factors here are divided in two groups, cognitive based
trust and afect based trust. Madsen and Gregor [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] conceptualise
trust as consisting of five main factors: perceived reliability,
perceived technical competence, perceived understandability, faith,
and personal attachment. Perceived Technical competence means
that the system is perceived to perform the tasks accurately and
correctly, based on the input information. Perceived Understandability
means that the user can form a mental model and predict future
system behaviours. Perceived Reliability means that the system is
perceived to be consistently functioning. Faith means that the user
is confident in the future ability of the system to perform, even in
situations in which has never used the system before. Finally,
personal attachment means that users find using the system agreeable,
preferable, and that suits their personal taste.
      </p>
      <p>
        Some of these factors overlap with the trust factors identified
by McKnight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. McKnight provides an understanding of trust
in technology in a wider societal context. McKnight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] defines
trust as consisting of three main components: propensity to trust
general technology, institution-based trust in technology, and trust
in specific technology. In the context of this paper we only focus
on trust in a specific technology. McKnight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] defines trust in
a specific technology as a person’s relationship with a particular
technology. Even if the study does not specifically target decision
systems, the paper goes into a large literature and looks at diferent
object of trust, trust attributes, and their empirical relationships,
thus proposing a scale of trust which demonstrated good reliability
with Cronbach’s alpha &gt; 0.89. In the proposed scale, trust with a
specific technology was analyzed into three factors: perceived
functionality, perceived helpfulness, and perceived reliability. Perceived
functionality is users’ perceived capability of the system to
properly accomplish its main function. Perceived helpfulness is users’
perception of the technology providing adequate, efective, and
responsive help. Finally, perceived reliability means that the system
is perceived to operate continually or responding predictably to
inputs.
      </p>
      <p>In our study we adopt a merged and modified version of the
9 trust items proposed by Madsen and Gregor and by McKnight.
From the total 9 trust items, that have been described above, we
merged items that overlapped in meaning and modified some of
their descriptions into the final 6 trust metrics: perceived
understandability, perceived reliability, perceived technical competence,
faith, personal attachment, and helpfulness (See Table 2).
3</p>
    </sec>
    <sec id="sec-7">
      <title>METHODOLOGY</title>
      <p>We aimed to test to what extent diferent types of textual
explanations afect diferent factors of users’ trust. In section 2.1 we have
identified 6 characteristics of meaningful explanation: contrastive,
truthful, general, thorough, social/interactive, and
role/domaindependent explanations (see Table 1). We used these characteristics
to design distinctive textual explanations, and then presented them
to users. Since we focus on a healthcare scenario, we used a
dramatising vignette to probe participants responses. We asked them to
read the explanation after reading the vignette and then run an
online survey asking them to rate diferent explanation types. To elicit
feedback on the explanation types we used the trust measurement
mentioned above.</p>
      <p>We designed a between-subjects study, in which diferent groups
of users were each presented with a diferent explanation type.
When designing the explanations, we focused on 4 out of the 6
explanation characteristics: contrastive, general, truthful, and
thorough. Social/interactive and role/domain-dependent characteristics
Characteristic
contrastive
general
Presented Explanation
"From the screen image, Malignant lesions are
present. Benign cases and fluid cyst looks
hollow and have a round shape. Your spots are not
hollow and and have irregular shapes.
Therefore, your spots are detected as Malignant."
"Based on your screen image, your spots are
detected as Malignant. 19 in 20 similar images
are in Malignant class."
"Using 5,600 of ultrasound images in our
database, your image have 95% similarities with
Malignant cases."
"Malignant lesions are present at 2 sites, 30mm
and 5mm. Non homogeneous. Non parallel. Not
circumscribed. Your risk of breast cancer as;
3050 years old, cyst history, woman is increased
20%"
were ignored at this stage for simplicity. In fact, these
explanation styles could not be expressed with a textual description, and
needed work on the UX design of the explanation type in order to
be realized. Therefore the assessment of the efects of these two
characteristics was left for future study. The AI system’s diagnosis
tool described in the dramatising vignette was a fictional AI system
for mammography diagnosis, used in a self managed health
scenario. With the system users could upload images of self-scanned
mammograms and then received a diagnosis result with an attached
textual explanation.
3.1</p>
    </sec>
    <sec id="sec-8">
      <title>Explanation Design</title>
      <p>In order to design the explanation, we first tried to look at breast
cancer diagnosis report and several screening reports including
ultrasound. Next, we designed the possible textual explanations
based on each characteristic definition in a small-scale informal
design phase. We then consulted the designed explanations with
researcher outside this study and medical professional. The
explanations were identical from a UI perspective, with one graphic and
followed by the diagnosis and the explanation text. The explanation
texts were designed to stress the four explanation characteristics:
contrastive, truthful, general, thorough. We also tried to present a
balanced level of system’s capability, for example in general style:
"19 in 20 similar images" and in truthful style: "95% similarities". The
explanation text presented to the participants can be seen in Table
3 and how we presented it can be seen in Figure 1.
3.2</p>
    </sec>
    <sec id="sec-9">
      <title>Data Collection and Analysis</title>
      <p>The participants were recruited on Mechanical Turk, with a survey
set up using Google Form. Our target was initially 80 participants,
with 40 participants from the general public and 40 participants
from worker in the healthcare field. We choose the option of "master
worker" and added one check-in question in the survey, to maximise
participation quality and check if the participant read the vignette
carefully. The Mechanical Turk hits were up for a week, and in the
end, we got 48 participants (only 8 with some medical expertise).
Participants were randomly assigned to 1 of the 4 conditions, with
each condition being a diferent explanation type. The number of
participants for each condition are not identical, with n1 = 12,
n2 = 12, n3 = 11, and n4 = 13.</p>
      <p>We asked participants to rate the AI system after having read
the dramatizing vignette and to reflect on the 6 trust’s components
while rating the explanation using a 7-points Likert scale. Following
a between-subject comparison of the results we were able to identify
which explanation (if any) afects which of the 6 components of
trust, and to what extent. The overall aim of the study was to give
us insights on how diferent styles of linguistics explanations afect
specific aspects of users’ trust. We also asked participants if they
would have liked the presented explanation to be included in the
AI system and explain why.</p>
      <p>To analyse the data, we used ANOVA tests, followed by Tukey’s
posthoc paired tests, to see the relative efects of diferent
explanation types. The ANOVA test tells us whether there is an overall
diference between the groups, but it does not indicate which
specific groups difered. The Tukey’s post-hoc tests can confirm where
the diference occurred between specific groups. In addition, we
evaluated the trust measurement instrument, by using Cronbach’s
Alpha.
4</p>
    </sec>
    <sec id="sec-10">
      <title>RESULTS</title>
      <p>From the online survey data, we ran two ANOVA tests, to check
the explanation styles and the trust factors. In the first ANOVA test,
we compared the 4 explanations types in relation to an average
trust factor (calculated as median value between the 6 trust scores).
We found that diferent styles of explanation significantly afect
average trust values (pvalue=0.0033, α =0.05). We then ran a Tukey’s
posthoc test, and found that general explanation show significantly
lower trust scores compared to the rest of the explanation styles;
contrastive, truthful, and thorough (α =0.05). The Tukey’s posthoc
test analysis can be seen in Fig 2.
In the second ANOVA test, we compared the four explanation
styles for each trust factor, we therefore ran 6 comparisons and
found that Personal Attachment was the only trust factor
showing significant diference (pvalue=0.02158, α =0.05). We then ran a
Tukey’s posthoc test for Personal Attachment, to identify where
the specific diference occurred, and found that contrastive and
thorough explanation styles shows significant diference compared
to general explanation style (α =0.05). The Tukey’s posthoc test
analysis can be seen in Fig 3.
As mentioned above, other than trust scaling, we also asked
participants if they would like the explanation style presented to
them to be included in the app for self managed health. We can see
in Fig 4, contrastive, truthful, and thorough explanation styles are
rated quite high (6 = very), while the general explanation style is
rated lower (5 = moderately). This assessment is consistent with the
explanation style-trust analysis we did. In the analysis, it shows that
general explanation is the least performing explanation in afecting
personal attachment.</p>
      <p>We also asked why participants preferred or not to receive the
explanation given to them. By qualitative analysing the 25 answers
from thorough and contrasting style groups, users reported the
presence of a clear rationale, and the use of lay terms, as the two
distinctive factors motivating the high trust rating. In turn, the need
of a rationale for the AI result was also explicitly mentioned as a
way to improve general explanation (by 4 out of 11 people in the
general explanation style group mentioned rationale as a need).</p>
      <p>The trust measurement was tested using the overall data from
48 participants. The reliability of the overall measurement was
determined by Cronbach’s Aplha. We found that the alpha is quite
high, α =0.88. This is an encouraging result which may inform
further use, testing and validation of the proposed human-AI trust
measure in other healthcare applications.
5</p>
    </sec>
    <sec id="sec-11">
      <title>DISCUSSION</title>
      <p>Our study confirms previous research indicating that diferent styles
of explanation significantly afect specific trust factors. In particular
we found that Personal Attachment (pvalue=0.02158) was
significantly afected by diferent textual explanation styles, and was
highly rated by the groups that were presented with thorough and
contrastive explanation styles. This means that among the
participant, thorough and contrastive styles suited their taste more,
compared to the general explanation style.</p>
      <p>This finding was corroborated by the additional comparison
of the 4 explanations by average trust ratings, which showed that
general style explanation was significantly rated lower than the rest
of the explanation styles. Overall preferability scores also confirmed
that general style explanation was rated the lowest.</p>
      <p>Participants seemed to prefer thorough and contrastive styles
explanation because of the rationale provided, and because of the
layperson language used to provide the explanation. The need of
rationale was also suggested as a way to improve general explanation
style.</p>
      <p>However, further investigations about the extent to which
explanation afects trust judgement need to be conducted. The current
results are not conclusive and suficient to develop an explanation
style and trust relation model. Additional studies to explore the
explanation mediums and interaction types are also necessary.
6</p>
    </sec>
    <sec id="sec-12">
      <title>LIMITATIONS AND FUTURE WORK</title>
      <p>This preliminary study has several limitations that should be noted.
This is an exploratory study of quite a broad topic and we only
conducted one online survey with low number of participants. The
fact that some explanation styles did not show significantly diferent
efects on users trust judgements could be caused by the small
sample size. Future studies with a bigger sample size and a baseline
group are needed to determine the extent of which explanation
afects trust.</p>
      <p>We also acknowledge that trust is dificult to measure. Even
though our trust measurement has shown high internal consistency,
we have not fully investigated the validity of the measurement in
other cases/fields. Moreover, in this experiments, we only measured
user’s trust as a self reported measure. Our experimental design,
and the use of a probing method, may have also possibly
influenced participants’ reflection and self reporting. Further research
is needed to carefully determine whether this was the case.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Achinstein</surname>
          </string-name>
          .
          <year>1983</year>
          .
          <article-title>The nature of explanation</article-title>
          . Oxford University Press on Demand.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Stavros</given-names>
            <surname>Antifakos</surname>
          </string-name>
          , Nicky Kern, Bernt Schiele, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Schwaninger</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Towards improving trust in context-aware systems by displaying system conifdence</article-title>
          .
          <source>In Proceedings of the 7th international conference on Human computer interaction with mobile devices &amp; services. ACM</source>
          ,
          <fpage>9</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Bussone</surname>
          </string-name>
          , Simone Stumpf, and
          <string-name>
            <surname>Dympna O'Sullivan</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The role of explanations on trust and reliance in clinical decision support systems</article-title>
          .
          <source>In 2015 International Conference on Healthcare Informatics. IEEE</source>
          ,
          <fpage>160</fpage>
          -
          <lpage>169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Pat</given-names>
            <surname>Croskerry</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Clinical cognition and diagnostic error: applications of a dual process model of reasoning</article-title>
          .
          <source>Advances in health sciences education 14</source>
          ,
          <issue>1</issue>
          (
          <year>2009</year>
          ),
          <fpage>27</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Finale</given-names>
            <surname>Doshi-Velez</surname>
          </string-name>
          , Mason Kortz, Ryan Budish, Chris Bavitz, Sam Gershman,
          <string-name>
            <surname>David O'Brien</surname>
            , Stuart Schieber, James Waldo, David Weinberger,
            <given-names>and Alexandra</given-names>
          </string-name>
          <string-name>
            <surname>Wood</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Accountability of AI under the law: The role of explanation</article-title>
          .
          <source>arXiv preprint arXiv:1711.01134</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Mary</surname>
            <given-names>T</given-names>
          </string-name>
          <string-name>
            <surname>Dzindolet</surname>
          </string-name>
          ,
          <article-title>Scott A Peterson, Regina A Pomranky, Linda G Pierce,</article-title>
          and Hall P Beck.
          <year>2003</year>
          .
          <article-title>The role of trust in automation reliance</article-title>
          .
          <source>International journal of human-computer studies 58</source>
          ,
          <issue>6</issue>
          (
          <year>2003</year>
          ),
          <fpage>697</fpage>
          -
          <lpage>718</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Riccardo</given-names>
            <surname>Guidotti</surname>
          </string-name>
          , Anna Monreale, Salvatore Ruggieri, Dino Pedreschi, Franco Turini, and
          <string-name>
            <given-names>Fosca</given-names>
            <surname>Giannotti</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Local rule-based explanations of black box decision systems</article-title>
          . arXiv preprint arXiv:
          <year>1805</year>
          .
          <volume>10820</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>David</given-names>
            <surname>Gunning</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Explainable artificial intelligence (xai)</article-title>
          .
          <source>(</source>
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Jonathan</surname>
            <given-names>L Herlocker</given-names>
          </string-name>
          ,
          <article-title>Joseph A Konstan,</article-title>
          and John Riedl.
          <year>2000</year>
          .
          <article-title>Explaining collaborative filtering recommendations</article-title>
          .
          <source>In Proceedings of the 2000 ACM conference on Computer supported cooperative work. ACM</source>
          ,
          <volume>241</volume>
          -
          <fpage>250</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Denis</surname>
            <given-names>J</given-names>
          </string-name>
          <string-name>
            <surname>Hilton</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Conversational processes and causal explanation</article-title>
          .
          <source>Psychological Bulletin</source>
          <volume>107</volume>
          ,
          <issue>1</issue>
          (
          <year>1990</year>
          ),
          <fpage>65</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Andreas</surname>
            <given-names>Holzinger</given-names>
          </string-name>
          , Chris Biemann, Constantinos S Pattichis, and Douglas B Kell.
          <year>2017</year>
          .
          <article-title>What do we need to build explainable AI systems for the medical domain</article-title>
          ?
          <source>arXiv preprint arXiv:1712.09923</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>René</surname>
            <given-names>F</given-names>
          </string-name>
          <string-name>
            <surname>Kizilcec</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>How much information?: Efects of transparency on trust in an algorithmic interface</article-title>
          .
          <source>In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>2390</volume>
          -
          <fpage>2395</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Todd</surname>
            <given-names>Kulesza</given-names>
          </string-name>
          , Simone Stumpf, Margaret Burnett, and
          <string-name>
            <given-names>Irwin</given-names>
            <surname>Kwan</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Tell me more?: the efects of mental model soundness on personalizing an intelligent agent</article-title>
          .
          <source>In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Todd</surname>
            <given-names>Kulesza</given-names>
          </string-name>
          , Simone Stumpf, Margaret Burnett, Sherry Yang,
          <string-name>
            <given-names>Irwin</given-names>
            <surname>Kwan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Weng-Keen Wong</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Too much, too little, or just right? Ways explanations impact end users' mental models</article-title>
          .
          <source>In 2013 IEEE Symposium on Visual Languages and Human Centric Computing. IEEE</source>
          ,
          <fpage>3</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Brian</surname>
            <given-names>Y</given-names>
          </string-name>
          <string-name>
            <surname>Lim and Anind K Dey</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Design of an intelligible mobile contextaware application</article-title>
          .
          <source>In Proceedings of the 13th international conference on human computer interaction with mobile devices and services. ACM</source>
          ,
          <volume>157</volume>
          -
          <fpage>166</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Lipton</surname>
          </string-name>
          .
          <year>1990</year>
          .
          <article-title>Contrastive explanation</article-title>
          .
          <source>Royal Institute of Philosophy Supplements</source>
          <volume>27</volume>
          (
          <year>1990</year>
          ),
          <fpage>247</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Zachary</surname>
            <given-names>C</given-names>
          </string-name>
          <string-name>
            <surname>Lipton</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <source>The Doctor Just Won't Accept That! arXiv preprint arXiv:1711.08037</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Tania</given-names>
            <surname>Lombrozo</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The structure and function of explanations</article-title>
          .
          <source>Trends in cognitive sciences 10</source>
          ,
          <issue>10</issue>
          (
          <year>2006</year>
          ),
          <fpage>464</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Madsen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shirley</given-names>
            <surname>Gregor</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Measuring human-computer trust</article-title>
          .
          <source>In 11th australasian conference on information systems</source>
          , Vol.
          <volume>53</volume>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <volume>6</volume>
          -
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Bertram</surname>
            <given-names>F</given-names>
          </string-name>
          <string-name>
            <surname>Malle</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>How the mind explains behavior: Folk explanations, meaning, and social interaction</article-title>
          . Mit Press.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>D</given-names>
            <surname>Harrison Mcknight</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michelle Carter</surname>
          </string-name>
          , Jason Bennett Thatcher, and Paul F Clay.
          <year>2011</year>
          .
          <article-title>Trust in a specific technology: An investigation of its components and measures</article-title>
          .
          <source>ACM Transactions on Management Information Systems (TMIS) 2</source>
          ,
          <issue>2</issue>
          (
          <year>2011</year>
          ),
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Tim</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Explanation in artificial intelligence: Insights from the social sciences</article-title>
          .
          <source>Artificial Intelligence</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Andrea</surname>
            <given-names>Papenmeier</given-names>
          </string-name>
          , Gwenn Englebienne, and
          <string-name>
            <given-names>Christin</given-names>
            <surname>Seifert</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How model accuracy and explanation fidelity influence user trust</article-title>
          .
          <source>arXiv preprint arXiv:1907</source>
          .
          <volume>12652</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Alun</given-names>
            <surname>Preece</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Asking 'Why'in AI: Explainability of intelligent systemsperspectives and challenges</article-title>
          .
          <source>Intelligent Systems in Accounting, Finance and Management</source>
          <volume>25</volume>
          ,
          <issue>2</issue>
          (
          <year>2018</year>
          ),
          <fpage>63</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Pearl</given-names>
            <surname>Pu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Li</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Trust building with explanation interfaces</article-title>
          .
          <source>In Proceedings of the 11th international conference on Intelligent user interfaces. ACM</source>
          ,
          <volume>93</volume>
          -
          <fpage>100</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Stephen J Read and Amy</surname>
          </string-name>
          Marcus-Newhall.
          <year>1993</year>
          .
          <article-title>Explanatory coherence in social explanations: A parallel distributed processing account</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>65</volume>
          ,
          <issue>3</issue>
          (
          <year>1993</year>
          ),
          <fpage>429</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Tulio</surname>
          </string-name>
          <string-name>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Model-agnostic interpretability of machine learning</article-title>
          .
          <source>arXiv preprint arXiv:1606.05386</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Kristin</given-names>
            <surname>Schaefer</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The perception and measurement of human-robot trust</article-title>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Saravanan</surname>
            <given-names>Thirumuruganathan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mahashweta Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Shrikant Desai</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sihem</surname>
            <given-names>AmerYahia</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gautam Das</surname>
            , and
            <given-names>Cong</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Maprat: Meaningful explanation, interactive exploration and geo-visualization of collaborative ratings</article-title>
          .
          <source>Proceedings of the VLDB Endowment 5</source>
          ,
          <issue>12</issue>
          (
          <year>2012</year>
          ),
          <fpage>1986</fpage>
          -
          <lpage>1989</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Eric</surname>
            <given-names>S</given-names>
          </string-name>
          <string-name>
            <surname>Vorm</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Assessing Demand for Transparency in Intelligent Systems Using Machine Learning</article-title>
          .
          <source>In 2018 Innovations in Intelligent Systems and Applications (INISTA)</source>
          .
          <source>IEEE</source>
          , 1-
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Danding</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qian Yang</surname>
            ,
            <given-names>Ashraf</given-names>
          </string-name>
          <string-name>
            <surname>Abdul</surname>
          </string-name>
          , and Brian Y Lim.
          <year>2019</year>
          .
          <article-title>Designing Theory-Driven User-Centric Explainable AI</article-title>
          .
          <source>In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>601</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>Sam</given-names>
            <surname>Wilkinson</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Levels and kinds of explanation: lessons from neuropsychiatry</article-title>
          .
          <source>Frontiers in psychology 5</source>
          (
          <year>2014</year>
          ),
          <fpage>373</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Ming</surname>
            <given-names>Yin</given-names>
          </string-name>
          , Jennifer Wortman Vaughan, and
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Wallach</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Understanding the Efect of Accuracy on Trust in Machine Learning Models</article-title>
          .
          <source>In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>279</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>