<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Counterfactual Explanations for eXplainable AI (XAI)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Greta Warren</string-name>
          <email>greta.warren@ucdconnect.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight SFI Centre for Data Analytics, University College Dublin</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science, University College Dublin</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>Counterfactual explanation has become a popular and promising method of explaining black-box AI systems and their decisions in recent years. However, a lack of rigorous psychological research means that little is known about what constitutes a 'good' counterfactual explanation, or how they facilitate user understanding of the underlying system. My doctoral research aims to examine how these sorts of explanations are understood and evaluated by users, identify desirable characteristics of counterfactual explanations, and investigate how current state-of-the-art counterfactual explanation techniques satisfy these criteria. These insights will guide the development of a novel explanation method designed to meet the psychological requirements of users. In order to address these research questions, to date I have conducted three large-scale, well-controlled user studies using materials drawn from an existing case-base. These studies have yielded novel findings about the impact of counterfactual explanation on users objective understanding and subjective judgments of an AI system. Based on these results, we have proposed an extension of a case-based counterfactual method that produces psychologically-valid explanations, which is to our knowledge, the first method designed with this specific criterion in mind.</p>
      </abstract>
      <kwd-group>
        <kwd>XAI</kwd>
        <kwd>counterfactual</kwd>
        <kwd>contrastive</kwd>
        <kwd>CBR</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Explaining opaque AI systems and their decisions using contrastive counterfactual examples
has gained considerable traction in recent years (see [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] for reviews). To this end, concepts
from case-based reasoning (CBR) such as Nearest Unlike Neighbours (NUNs [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]) have inspired
such approaches to explanation-by-example, by providing information about how an alternative
system decision could have been made, had some aspect of the input data been diferent [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
For example, after rejection for a bank loan, a counterfactual explanation may inform the
applicant: “had your salary been €10,000 higher, your application would have been approved”.
Counterfactual explanations have been proposed to appeal to important characteristics of
human explanation and causal reasoning [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ], as well as ofering potential for recourse [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
However, although there has been a surge in the number of methods proposed for generating
counterfactual explanations computationally, there is limited evidence to show that the outputs
of these methods meet the psychological criteria of a ‘good’ explanation, while a lack of
controlled user studies to evaluate their impact on user understanding and perceptions of the
CEUR
system jeopardises their real-world utility. Furthermore, many existing studies rely on users’
subjective satisfaction, trust, or fairness judgments, which may not necessarily reflect the depth
of their understanding of the system’s causal mechanisms [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>My doctoral research seeks to address these issues by examining how counterfactual
explanations are evaluated by human users using both objective and subjective measures, and
identifying psychological desiderata of these sorts of explanations. This is achieved by
conducting large-scaled, controlled user studies with materials drawn from an existing case-base.
These insights will guide an analysis of existing computational methods to assess how well they
meet these psychological criteria, as well as the design of a novel, psychologically-grounded
case-based approach to counterfactual explanation.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Research Plan</title>
      <sec id="sec-2-1">
        <title>2.1. Research Objectives</title>
        <p>
          Counterfactual explanations have received significant attention in recent years as a means of
elucidating decisions made by black box AI systems to users. Over 100 methods have been
proposed to generate counterfactual explanations [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], and are commonly compared to the state
of the art with reference to proximity [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], sparsity [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], and plausibility [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. However, it is striking
that so few of these methods are evaluated with respect to the primary stakeholders (i.e.,
endusers [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]). Moreover, these quantitative metrics are based on researchers’ intuitions about what
constitutes a ’good’ explanation, however, it is unclear how (and if) they map to longstanding
psychological and philosophical definitions of explanatory power [
          <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
          ]. Indeed, although there
is a rich body of literature surrounding human explanation [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], and counterfactual reasoning
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], relatively little is known about counterfactual explanations beyond the context of XAI and
how they are understood.
        </p>
        <p>The core objectives of my research are to examine how counterfactual explanations impact
users’ understanding and perceptions of an AI system, and identify the optimal characteristics
of these explanations, in order to guide the design of a novel, user-centric counterfactual method
that produces psychologically-valid explanations. Specifically, I investigate how counterfactual
explanations of AI predictions improve users’ objective accuracy in a prediction task, and
subjective judgments of explanation satisfaction and trust in the system. In addition, I examine
how focusing on certain feature-types appears to increase user accuracy, and hence, help users
more readily understand the AI system. These insights will guide both the development of a
counterfactual explanation method that meets users’ psychological requirements, as well as
shed new light on counterfactual explanation in human cognition. The key research questions I
have identified are:
• What are the optimal characteristics of a counterfactual explanation from a psychological
perspective?
• Which counterfactual methods produce the best explanations in terms of computational
metrics (e.g., sparsity, proximity, plausibility)?
• How can a counterfactual method produce explanations that meet users’ psychological
criteria of explanations?</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Approach / Methodology</title>
        <p>
          User Studies. In order to investigate the efects of counterfactual explanations on users’
understanding and evaluations of an AI system, we conducted a series of user studies designed
to assess the impact of counterfactual explanation on users’ task accuracy and subjective
judgments. We compare these efects to those of causal explanations and a control condition
(in which participants receive only descriptions of the system’s decisions). Participants in the
studies were presented with materials in the form of case-instances, each consisting of five
features used to predict blood alcohol content: gender (male/female), weight (in kg), amount
of alcohol consumed by the person (in units), duration of drinking period (in minutes), and
stomach-fullness (full/empty). Users were shown the output of a simulated AI system presented
as an application, designed to predict whether someone is over the legal blood alcohol content
limit to drive. Materials were selected from a case-base of instances of normally-distributed
values of the feature-set. In the training phase of the experiments, participants were shown
examples of tabular data for diferent individuals, and asked to make a judgment about whether
each individual was under or over the limit on each screen. After giving their response, feedback
was given on the next page, along with an explanation, the content of which was dependent on
the experimental condition (see Figure 1 for a sample of the material used in the counterfactual
condition). Upon completing the training phase, participants began the testing phase, in which
they were shown more example instances referring to individuals and again asked to judge
if each individual was over or under the legal limit to drive. For each instance, participants
were asked to consider a specific feature in making their prediction; for instance, “Given this
person’s WEIGHT, please make a judgment about their blood alcohol level.” After submitting
their response, no feedback or explanation was given. In addition to measuring task accuracy,
participants were also asked to provide judgments of explanation satisfaction and trust, measured
using the DARPA Explanation Satisfaction and Trust scales [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] respectively, allowing us to
evaluate explanation quality using both objective and subjective measures, which may not
necessarily correspond with one another.
        </p>
        <p>
          Towards a Psychologically-valid Counterfactual Method. A key result from the user
studies discussed above was that users were significantly more accurate when making
predictions about categorical features (stomach fullness and gender) than continuous features (units,
weight and drinking duration). This finding is supported by evidence from the counterfactual
reasoning literature that people do not spontaneously change continuous variables when
generating counterfactuals for past events [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. In light of this, we conducted an analysis of NUNs
with categorical feature diferences in a number or popular UCI datasets, observing that they
are exceedingly rare. Hence, we developed a variation of Keane and Smyth’s [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] case-based
counterfactual method, which applies post-hoc transformations to the feature diferences in
order to produce counterfactual explanations more intuitively understandable to end-users (see
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] for more detail).
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Progress Summary</title>
      <p>
        To date, I have conducted three large-scale, well-controlled user studies (total N = 474) which
have revealed novel insights into how counterfactual and causal explanations are understood
and perceived by users. While counterfactual explanations are judged as more satisfying and
trustworthy than causal explanations, they appear to be only slightly more efective in improving
objective performance in a prediction task. This disconnect between objective and subjective
measures suggests that it is critical to examine how explanations aid user understanding rather
than merely improve subjective perceptions. Furthermore, users appear to understand the
impact of categorical features on the system’s decision more readily than that of continuous
features, a distinction that current computational methods do not account for. Findings from the
ifrst user study were presented at the Cognitive Aspects of Knowledge Representation workshop
at IJCAI’22 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], with preliminary results presented at CogSci’21. Findings from the complete
series of user studies are currently being prepared for submission to a top-tier conference.
      </p>
      <p>
        Based on the finding that counterfactuals that change categorical features are more readily
understood than those focusing on continuous features, we developed a counterfactual method
that accounts for this feature-type distinction. An analysis of common UCI datasets suggests
that sparse counterfactuals with categorical feature-changes are relatively rare, and so our
method adapts Keane and Smyth’s [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] case-based technique to transform feature-diferences
into categorical versions, without significant decrement to performance in terms of coverage
and proximity of the counterfactuals produced. To our knowledge, this is the first counterfactual
method designed to meet identified psychological requirements for explanation by users, and
will be presented at ICCBR’22 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>The main focus of my research at present is the design of a second series of psychological
experiments examining the role of simplicity (or sparsity) in counterfactual explanation, and how
it impacts user understanding and subjective judgments. In tandem, I am working on the
implementation and evaluation of popular counterfactual computational methods, in order to identify
those methods which are most successful (i.e. have the best average performance) in generating
counterfactuals that meet given criteria of an explanation over a set of representative problems.
These criteria include conventional metrics (e.g., proximity, sparsity, plausibility) as well as novel
properties derived from user testing (such as whether a counterfactual makes continuous or
categorical feature-changes). The final phase of my Ph.D. research will involve synthesising the
insights from these two strands of work in order to develop a novel, psychologically-grounded
method for counterfactual explanation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Keane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Kenny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Delaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          ,
          <article-title>If only we had better counterfactual explanations: Five key deficits to rectify in the evaluation of counterfactual xai techniques</article-title>
          ,
          <source>IJCAI-21</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Karimi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schölkopf</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Barthe, I. Valera, A survey of algorithmic recourse: Definitions, formulations, solutions, and prospects</article-title>
          , volume
          <volume>1</volume>
          , Association for Computing Machinery,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nugent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Doyle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <article-title>Gaining insight through case-based explanation</article-title>
          ,
          <source>Journal of Intelligent Information Systems</source>
          <volume>32</volume>
          (
          <year>2009</year>
          )
          <fpage>267</fpage>
          -
          <lpage>295</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wexler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pushkarna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bolukbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wattenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Viégas</surname>
          </string-name>
          , J. Wilson,
          <article-title>The what-if tool: Interactive probing of machine learning models</article-title>
          ,
          <source>IEEE transactions on visualization and computer graphics 26</source>
          (
          <year>2019</year>
          )
          <fpage>56</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>Explanation in artificial intelligence: Insights from the social sciences</article-title>
          ,
          <source>Artificial Intelligence</source>
          <volume>267</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.artint.
          <year>2018</year>
          .
          <volume>07</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <article-title>Counterfactuals in explainable artificial intelligence (xai): Evidence from human reasoning</article-title>
          , volume
          <volume>2019</volume>
          <source>-Augus</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6276</fpage>
          -
          <lpage>6282</lpage>
          . doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2019</year>
          / 876.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Buçinca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Z.</given-names>
            <surname>Gajos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Glassman</surname>
          </string-name>
          ,
          <article-title>Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems</article-title>
          ,
          <year>2020</year>
          , pp.
          <fpage>454</fpage>
          -
          <lpage>464</lpage>
          . doi:
          <volume>10</volume>
          .1145/3377325. 3377498.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wachter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mittelstadt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <article-title>Counterfactual explanations without opening the black box: Automated decisions and the gdpr</article-title>
          ,
          <source>Harvard Journal of Law &amp; Technology</source>
          <volume>31</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Keane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          ,
          <article-title>Good counterfactuals and where to find them: A case-based technique for generating counterfactuals for explainable ai (xai</article-title>
          ), Springer, Cham,
          <year>2020</year>
          , pp.
          <fpage>163</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Keil</surname>
          </string-name>
          , Explanation and understanding,
          <source>Annual Review of Psychology</source>
          <volume>57</volume>
          (
          <year>2006</year>
          )
          <fpage>227</fpage>
          -
          <lpage>254</lpage>
          . doi:
          <volume>10</volume>
          .1146/annurev.psych.
          <volume>57</volume>
          .102904.190100.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>R. M. Byrne</surname>
          </string-name>
          , Counterfactual thought,
          <source>Annual review of psychology 67</source>
          (
          <year>2016</year>
          )
          <fpage>135</fpage>
          -
          <lpage>157</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Hofman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Mueller</surname>
          </string-name>
          , G. Klein,
          <string-name>
            <given-names>J.</given-names>
            <surname>Litman</surname>
          </string-name>
          ,
          <article-title>Metrics for Explainable AI: Challenges and Prospects</article-title>
          ,
          <source>Technical Report December</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kahneman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tversky</surname>
          </string-name>
          ,
          <article-title>The simulation heuristic</article-title>
          , in: D.
          <string-name>
            <surname>Kahneman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Slovic</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Tversky (Eds.),
          <source>Judgment Under Uncertainty: Heuristics and Biases</source>
          , Cambridge University Press, New York,
          <year>1982</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Warren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          , M. T. Keane, “
          <article-title>better” counterfactuals, ones people can understand: Psychologically-plausible case-based counterfactuals using categorical features for explainable ai (xai)</article-title>
          , in: To appear
          <source>in ICCBR'22</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G.</given-names>
            <surname>Warren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Keane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <article-title>Features of explainability: How users understand counterfactual and causal explanations for categorical and continuous features in xai</article-title>
          ,
          <source>in: IJCAI-22 Workshop on Cognitive Aspects of Knowledge Representation</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>