<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Xiv.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Fostering Human-AI interaction: development of a Clinical Decision Support System enhanced by eXplainable AI and Natural Language Processing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Laura Bergomi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical, Computer and Biomedical Engineering, University of Pavia</institution>
          ,
          <addr-line>Pavia</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <volume>00711</volume>
      <abstract>
        <p>Artificial Intelligence (AI) is increasingly integrated into Decision Support Systems (DSS). The explainability of AI-based systems becomes crucial in sensitive and critical domains, such as healthcare, where ethical considerations and reliability are paramount concerns. In the clinical setting, it is important to evaluate how humans and AI can collaborate on cognitive tasks. Collaboration protocols (HAI-CP) allow for the investigation of the usefulness of AI models and their impact on users (both positive and negative). Although research on the application of these methods is blooming, there is little understanding of the impact on clinical decision-making, especially for eXplainable AI (XAI) systems, due to the lack of user studies. Therefore, the goal of this proposal is to develop a clinical DSS enhanced by XAI and Natural Language Processing (NLP): their synergy can add value to the interaction between users and AI, fostering a more linguistically natural, comprehensible, trustworthy, and supporting interfacing, that blends into the existing workflows. This proposal explores potential solutions to tailor natural language explanations and data visualizations to the end-user, improving the comprehensibility of the reasons behind a decision, and increasing the user's confidence in the decision; investigates and tests possible strategies to “get the patient-in-the-loop”; explores uncertainty quantification and counterfactual approaches, and finally assesses the impact on naturalistic (i.e., real-world) decision-making and long-term efects and biases.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Clinical decision making</kwd>
        <kwd>Explainable artificial intelligence</kwd>
        <kwd>Natural language interaction</kwd>
        <kwd>Human-AI collaboration protocol</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Artificial Intelligence (AI) is increasingly integrated into Decision Support Systems (DSS). To
understand how useful and usable an AI-based system actually is, it is important to assess
its impact on the decision support workflow. For a DSS to meet the basic requirement of
Art. 22 para 1 GDPR (which prohibits “decision based solely on automated processing”), it is
important to provide that the human-in-the-loop has substantial evaluative power and so the
last word on the outcome of a decision [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To ensure that a human can make a thoughtful
and reliable decision, the explainability of AI-based systems becomes crucial. In sensitive and
critical domains, such as healthcare, where ethical considerations and reliability are paramount
concerns, AI-based DSSs should provide high-quality explanations. In this context, eXplainable
AI (XAI) has become increasingly important as a counterbalancing force to the widespread
adoption of complex black box models, that leave users, and even developers, in the dark as to
how results were obtained [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. XAI refers to the development of AI systems that can provide
clear, understandable, and interpretable explanations. XAI methods can be described based
on several categorizations; what remains essential is to adapt and test their explanations in
Human-Artificial Intelligence collaboration protocol (HAI-CP). An HAI-CP is “the instance of
a process schema that stipulates the use of AI tools by competent practitioners to perform a
certain task or do a certain job” [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], thus allowing to study the interaction of several parameters
involving at least the following dimensions : Afordance (functionalities and task automation),
Fit (fitting into the existing work practice), Optimization (learning phase tuning), Output (type
of result returned) and Target (the characteristics of the intended user).
      </p>
      <p>
        Despite the application of AI methods in healthcare being a highly active field of research
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], how AI recommendations afect clinical decision-making is still poorly understood due to
the lack of user studies [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This motivates the present proposal to investigate extensively how
AI systems can be integrated into medical practice, being an aid that is easily understood, but
also naturally interfaceable (enabling natural language interaction, NLI). The latter issue can be
supported by using Natural Language Processing (NLP) techniques to generate explanations
in natural language, that are tailored to the end-user. The type of end-user has intentionally
not been specified, suggesting that it may be not only a medical professional, but also the
patient (one of the original contributions of the research). GDPR (Recital 71) remarks that
providing information about the existence of automated processes, the logic behind them, and
their potential consequences should nevertheless be provided voluntarily as a good practice to
ensure fairness and transparency. This means patients should have access to explanations for
clinical decisions, allowing them to engage in discussions about their care and understand the
reasoning behind treatments or examinations.
      </p>
      <p>This proposal is part of the Italian project PRIN PNRR 2022 "InXAID - Interaction with
eXplainable Artificial Intelligence in (medical) Decision-making" (CUP: H53D23008090001
funded by the European Union - Next Generation EU), whose aim is to explore a model-agnostic
perspective on the development and evaluation of AI-based (medical) DSSs. Our focus is on
the interaction between an AI system and the human user, i.e., understanding to what extend
the advice (and the way it is presented) can influence, be understood, and used by users; the
interaction features and efects.</p>
      <p>The rest of the paper is organized as follows. A brief review of related work is provided in
Section 1.1; Section 2 presents the goals of this research, followed by the planned approaches
and methods to achieve them in Section 3. Expected results of our research are proposed in
Section 4, and finally, the conclusions are described in Section 5 ofering some questions I would
like feedback on.</p>
      <sec id="sec-1-1">
        <title>1.1. Related works</title>
        <p>
          A human-in-the-loop approach has been advocated as essential for proper evaluation of AI for
healthcare [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]: explanations are ultimately directed to a human (expert) interacting with the
AI system and should be optimized for this. Following this direction, Gaube et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] compare
the detection of abnormalities in chest radiology images by physicians with the support of an
AI or a second human agent; Tschandl et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] evaluate, both with and without the use of AI,
clinicians skin cancers recognition. Cabitza et al. [
          <xref ref-type="bibr" rid="ref3 ref9">3, 9</xref>
          ] propose experimental results of applying
XAI to knee MRI and ECG studies, intending to compare the efectiveness of HAI-CP (testing
diferent orders of presentation (human first vs. AI first) and availability of explanations (yes
vs. no)). In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] the support of the AI system includes both a proposed diagnosis and a textual
explanation to back the former one.
        </p>
        <p>
          Some surveys (e.g. [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]) have studied the use of natural language techniques in creating
explanations, e.g. the synergy of NLP and XAI methods. Sokol and Flach [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] argue that
natural language explanations give the process a natural feeling, increasing the reliability
of explanations and helping to gain acceptance from a wider range of users. Despite these
considerations, a small part of XAI’s works uses natural language presentation methods [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
Works in the literature concerning the involvement of humans in clinical decision-making, such
as those cited so far, point to the clinician as the user (i.e., the human-in-the-loop). Only a few
examples involve patients, e.g., Donatello et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] address the challenges of a system that
supports patients in following a healthy behavior by presenting an XAI system that supports
the monitoring of users’ behaviors and persuades them to follow a healthy lifestyle.
        </p>
        <p>
          Ultimately, some works (e.g. [
          <xref ref-type="bibr" rid="ref14 ref15 ref3">3, 14, 15</xref>
          ]) emphasize the need to examine biases in AI and XAI
support and focus on AI advice efects.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Research goals</title>
      <p>
        Current XAI research often overlooks the importance of presentation techniques, leading to
challenges for researchers and practitioners in selecting appropriate methods for explainability.
The lack of comprehensive studies adds complexity and potential errors to the process [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
This research will focus primarily on two main dimensions: the output and the target. In the
ifrst case (i.e., the output) the goal is to compare the presentation techniques of XAI methods,
studying what type of output the end-user prefers and understands. The output of an XAI
system can be multimodal (e.g., classes, confidence scores, category lists, visual or textual
explanations, etc.). A better understanding of this output allows users to increase the overall
acceptability of the system. In the second case (i.e., the target) the goal is to tailor explanations
to the intended end-user in terms of profiling (expertise, experience, role, etc.), content, and
form. Going towards dialogues that directly engage the user in the explanation process, we can
ofer rich and personalized interactions that mimic how humans explain their decisions.
• What factors determine the choice of how to represent the explanations obtained from
      </p>
      <p>AI-based methods?
• Does XAI support, both in terms of visual aids and textual explanations, have a significant
efect on taking better clinical decisions, and reducing errors?</p>
      <p>Clinicians may lack the computer science expertise to comprehend algorithmic
decisionmaking processes, highlighting the need for clear communication. As a result, it becomes
imperative for clinicians, as decision-makers, to ensure that decisions made through these
processes can be clearly communicated to patients, who may have little technical knowledge.
• What is the impact of asking decision-makers to explain their AI-based decisions?
• What is the best way to “get the patient-in-the-loop”?</p>
      <p>Regarding the assessment of impact on decision-making, we are asking about counterfactual
explanations and biases in AI and XAI support:
• What is the impact of providing decision-makers with similar cases, or providing
counterfactual outcomes, or playing the devil’s advocate role?
• Does the AI advice afect the users’ performance and proficiency-building processes, both
in the short and in the long run? Is this influence also relevant according to the user’s
experience, expertise, and role?
• Are explanations always beneficial or can they induce paradoxically harmful efects?</p>
    </sec>
    <sec id="sec-3">
      <title>3. Planned approaches and methods</title>
      <p>
        This section outlines the main phases of the research (as shown in the Gantt diagram in Figure 1),
the approaches, and methods thought to be used as a starting point in each phase and from which
to develop extensions and insights. At each stage, there is the analysis of State-of-Art methods
and the implementation of HAI-CPs to test hypotheses and collect reliable and naturalistic
(i.e., real-world) decision-making results. Particular attention will be paid to interpreting and
ranking the relative strength of claims about the efectiveness of design solutions and their
superiority concerning other possible alternatives [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. The inXAID project will ensure that the
frameworks, metrics and methods developed, as well as the results of the experiments performed
throughout the duration of the research activity, will be disseminated to the appropriate target
communities and audiences. A research period abroad is also planned.
      </p>
      <p>Research activity topic
RP 1: Tailoring XAI and NLP explanations
XAI / NLP: deriving the explanation
XAI / NLP: data presentation to the end-user
Combination of XAI and NLP
RP 2: Get the patient-in-the-loop
Review: patient-AI collaborations
"Patient-out-the-loop"
"Patient aware-of-the-loop"
"Patient-in-the-loop"
RP 3: Assessment of impact on decision-making
Uncertainty Quantification (UQ) assessment
Counterfactual explanations methods
Reliance patterns and biases
Mitigation strategies</p>
      <p>Further phases
Dissemination
PhD dissertation writing</p>
      <p>I year* II year* III year
bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim. bim.
1 2 3 4 5 6 1 2 3 4 5 6 1 2 3 4 5 6
SoA TI
SoA TI</p>
      <p>RA
ES 1 RA
TI RA</p>
      <p>SoA RP</p>
      <p>TI ES 2 RA</p>
      <p>ES 3</p>
      <p>RA
TI ES 4 RA</p>
      <p>SoA ES 5 RA
SoA TI RA</p>
      <p>ES 6
SoA
SoA</p>
      <p>RA
RA</p>
      <sec id="sec-3-1">
        <title>Research phase 1: Tailoring XAI and NLP explanations. In this phase, the aim is to</title>
        <p>
          investigate (i) techniques for deriving the explanation and (ii) presentation to the end-user.
The main XAI techniques to test are related to feature importance analysis (e.g. SHAP [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ],
T-EBAnO [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]) and the use of surrogate models (e.g. LIME [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], AraucanaXAI [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]). Comparing
diferent models, we can assess which type of data visualization is preferred by users. Among
the existing NLP models, we analyze transformers models for two pivotal reasons: they rely on
the attention mechanism and they are exceptionally efective for common natural language
understanding (NLU) and natural language generation (NLG) tasks [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. A textual explanation
may be presented in diferent ways to the end-user. Some techniques to test include: saliency,
visualization of the importance scores, showing input-output word alignment, highlighting
words in input text, or displaying extracted relations or word clouds; rewrite the explanations by
changing linguistic register. A possible comparison to explore concerns, on the one hand, is the
self-explaining approach, which generates the explanation at the same time as the prediction;
on the other hand, the post-hoc approach, which requires that an additional operation be
performed after the return of the prediction. The final step regards the combination of XAI and
NLP techniques; a possible simple solution could be rendering XAI explanations through NLG
or using XAI techniques to explain image, text, and graph classification models (e.g.,
Layerwise relevance propagation [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]); or incorporating feedback from human users to improve the
explanations generated by the model. This could involve allowing users to provide feedback on
the explanations or incorporating user preferences into the explanation process. In the example
of [
          <xref ref-type="bibr" rid="ref3 ref9">3, 9</xref>
          ], it is also important to test alternative fitting of AI advice in existing practice (e.g.,
human first vs. AI first, availability of explanations only in critical cases, AI advice on request).
Research phase 2: Get the patient-in-the-loop. In this phase, the aim is to investigate
diferent strategies to involve the patient in the clinical decision-making process. Firstly, it
is necessary to review the existing work in the literature regarding surveys on patient-AI
collaborations. Subsequently the idea is to investigate diferent ways of including patients,
possibly going through step: (i) give clinicians help in explaining their decision
(patient-out-theloop); (ii) analyze how patients’ behavior changes knowing that the doctor’s decision about their
health status is based on AI advice (patient aware-of-the-loop); (iii) use conversational chatbots
to interact with patients and, thus, collect important information (patient-in-the-loop). In this
direction, Bennettot et al. [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] provide an example of how XAI can benefit healthcare, specifically
in monitoring diabetic patients. They discuss a virtual coaching system that ofers guidance
on healthy behaviors based on patient-reported data. When the system detects undesirable
behaviors, it generates tailored explanations for clinicians and patients. This approach enables
clinicians to adjust treatment decisions as needed, while patients gain confidence, feel supported,
and become more engaged in their care decisions. At this level, also for ethical considerations,
the inclusion of a professional figure such as a behavioral psychologist, is essential.
Research phase 3: Assessment of impact on decision-making. In this phase, the aim is
to assess how, how much, and when, an AI-based system influences the cognitive processes
involved in clinicians’ interpretation of AI support. An explanation, whether true or fictitious,
could be correlated by Uncertainty Quantification (UQ) assessment. In this context, we can
apply diferent approaches, also integrated in NLP models [
          <xref ref-type="bibr" rid="ref18 ref23">18, 23</xref>
          ]. On the other hand, some
methods explain the model by providing information on feature-perturbed versions of the
analyzed instance. These methods fall into the counterfactual explanations methods [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Some
common approaches to test could be [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]: add or delete information to what users know about
the facts; create counterfactuals that imagine how the outcome could have been better or worse;
construct explanations by ensuring they identify cause-efect or reason-action relations between
events; semi-factual about how the outcome could have been the same “even if” the action had
been diferent. Identifying the cognitive processes involved in physicians’ interpretation of AI
support also precludes analysis of how these same processes may change over time. It can be
useful to realize a mapping of reliance patterns and biases afecting human decision-making
in the short and long run (e.g., preference for usability over performance, more attention to
false negatives than false positives, deskilling/ upskilling etc.); moreover, summarize promising
mitigation strategies and research directions to support human critical thinking (e.g., delay
showing the AI’s prediction and/or explanations, give arguments for non-predicted outcomes,
enable to actively explore the data, etc.).
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Expected results</title>
      <p>
        The main objective of this research proposal is to make a significant contribution towards
advancing the methodological leveraging of XAI methods and natural language explanations,
fostering a more linguistically natural, comprehensible, trustworthy, and supporting interfacing
among AI systems and human users, that fits into the existing work practices. On the clinical
front, the project’s main objective is to enhance decision-making processes by expanding
output solutions and taking into consideration aspects that are currently unexplored. These
aspects include the contribution of socio-technical elements, experience, expertise, habits,
and environmental and temporal information towards cognitive processes. It is expected
that “getting the patient-in-the-loop” of decision-making can be an original strength for AI
collaborative systems; so that it can also be supportive to the patient, improving well-being.
Preliminary results and contributions to date are related to the test of ALFABETO [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] (whose
aim is to aid clinicians during COVID-19 patients’ hospital admission through the application
of machine learning approaches exploiting clinical and chest x-ray features) in a clinical survey.
In this framework, diferent predictions and explanations are proposed to clinicians: from an
interpretable model (i.e., ALFABETO original Bayesian network) and from a black box model
(i.e., Gradient boosting), in this latter case, with the explanations of two diferent XAI approaches
(SHAP and Araucana XAI).
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Research challenges and future direction</title>
      <p>
        The research is still at an early stage, and thus, there are several challenges that we should
cope with. In XAI, some works have claimed that explainability may come at the price of losing
predictive performance. Studying such possible trade-ofs is an important research area, but
one that cannot advance until standardized metrics are developed for evaluating the quality
of explanations [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Indeed, an open debate in the NLG community is about finding the right
way to measure the goodness of generated explanations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The main issues revolve around
whether to rely only on automatic metrics (e.g., ROUGE, BLEU), or instead, how to properly
perform human evaluations. That said, although human evaluation remains the gold standard
for overall system quality assessment, using it at every stage of the development process would
be too costly and slow. Another challenge may be the interactive integration of desired behavior,
fairness, correctness, and reliability; as well as verifying and integrating knowledge in each step
of decision-making process, instead of XAI producing a single description of a static system
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>From these general considerations, and more, arise the following questions, on which I would
appreciate feedback by the DC mentors, as I believe they can help improve my research path:
• What metrics would be better to investigate (automatic or human) to evaluate the goodness
of explanations, especially in natural language?
• How to engage the user interactively, hoping for fruitful and continuous use of the
provided DSS?
• What other human actors (besides clinicians and patients) could we include?
• What aspects should we not forget to consider? (e.g., from an ethical and legal perspective)
• What approaches and potential collaborations should we consider to address the
challenges of developing a reliable and explainable DSS using clinical data?
I would like to express my willingness to receive constructive criticism and suggestions on this
work and to clarify any points that may need further explanation.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>I would like to express my gratitude to my supervisor, Enea Parimbelli, for his invaluable
assistance in developing this project and its implementation in the future. Authors acknowledge
funding support provided by the Italian project PRIN PNRR 2022 InXAID - Interaction with
eXplainable Artificial Intelligence in (medical) Decision-making. CUP: H53D23008090001 funded
by the European Union - Next Generation EU.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Schneeberger</surname>
          </string-name>
          , et al.,
          <article-title>The european legal framework for medical ai</article-title>
          ,
          <source>in: Machine Learning and Knowledge Extraction</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>209</fpage>
          -
          <lpage>226</lpage>
          . doi:https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -57321-8_
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Combi</surname>
          </string-name>
          , et al.,
          <article-title>A manifesto on explainability for artificial intelligence in medicine</article-title>
          ,
          <source>Artificial Intelligence in Medicine</source>
          <volume>133</volume>
          (
          <year>2022</year>
          ). doi:https://doi.org/10.1016/j.artmed.
          <year>2022</year>
          .
          <volume>102423</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          , et al.,
          <article-title>Rams, hounds and white boxes: Investigating human-ai collaboration protocols in medical diagnosis</article-title>
          ,
          <source>Artificial Intelligence in Medicine</source>
          <volume>138</volume>
          (
          <year>2023</year>
          )
          <article-title>102506</article-title>
          . doi: https://doi.org/10.1016/j.artmed.
          <year>2023</year>
          .
          <volume>102506</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V. L.</given-names>
            <surname>Patel</surname>
          </string-name>
          , et al.,
          <source>The Coming of Age of Artificial Intelligence in Medicine, Artificial intelligence in medicine 46</source>
          (
          <year>2009</year>
          )
          <fpage>5</fpage>
          -
          <lpage>17</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.artmed.
          <year>2008</year>
          .
          <volume>07</volume>
          .017.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          , et al.,
          <article-title>Quod erat demonstrandum? - towards a typology of the concept of explanation for the design of explainable ai</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>213</volume>
          (
          <year>2023</year>
          ). doi:https://doi.org/10.1016/j.eswa.
          <year>2022</year>
          .
          <volume>118888</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>A. Holzinger,</surname>
          </string-name>
          <article-title>Interactive machine learning for health informatics: when do we need the human-in-the-loop? 3 (</article-title>
          <year>2016</year>
          )
          <fpage>119</fpage>
          -
          <lpage>131</lpage>
          . doi:
          <volume>10</volume>
          .1007/s40708-016-0042-6.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gaube</surname>
          </string-name>
          , et al.,
          <article-title>Do as ai say: susceptibility in deployment of clinical decision-aids</article-title>
          ,
          <source>NPJ digital medicine 4</source>
          (
          <year>2021</year>
          )
          <article-title>31</article-title>
          . doi:
          <volume>10</volume>
          .1038/s41746-021-00385-9.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Tschandl</surname>
          </string-name>
          , et al.,
          <article-title>Human-computer collaboration for skin cancer recognition</article-title>
          ,
          <source>Nature Medicine</source>
          <volume>26</volume>
          (
          <year>2020</year>
          )
          <fpage>1229</fpage>
          -
          <lpage>1234</lpage>
          . doi:
          <volume>10</volume>
          .1038/s41591-020-0942-0.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          , et al.,
          <article-title>Painting the black box white: Experimental findings from applying XAI to an ECG reading setting</article-title>
          ,
          <source>Machine Learning and Knowledge Extraction</source>
          <volume>5</volume>
          (
          <year>2023</year>
          )
          <fpage>269</fpage>
          -
          <lpage>286</lpage>
          . doi:
          <volume>10</volume>
          .3390/make5010017.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cambria</surname>
          </string-name>
          , et al.,
          <article-title>A survey on xai and natural language explanations</article-title>
          ,
          <source>Information Processing &amp; Management</source>
          <volume>60</volume>
          (
          <year>2023</year>
          )
          <article-title>103111</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.ipm.
          <year>2022</year>
          .
          <volume>103111</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Qian</surname>
          </string-name>
          , et al.,
          <article-title>Xnlp: A living survey for xai research in natural language processing</article-title>
          ,
          <source>in: 26th International Conference on Intelligent User Interfaces-Companion</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>80</lpage>
          . doi:
          <volume>10</volume>
          .1145/3397482.3450728.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sokol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Flach</surname>
          </string-name>
          ,
          <article-title>Conversational explanations of machine learning predictions through class-contrastive counterfactual statements</article-title>
          ,
          <source>in: Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>5785</fpage>
          -
          <lpage>5786</lpage>
          . doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2018</year>
          /836.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Donadello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dragoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Eccher</surname>
          </string-name>
          ,
          <article-title>Persuasive explanation of reasoning inferences on dietary data</article-title>
          , in: PROFILES/SEMEX@ISWC, volume
          <volume>2465</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bertrand</surname>
          </string-name>
          , et al.,
          <article-title>How cognitive biases afect xai-assisted decision-making: A systematic review</article-title>
          ,
          <source>in: Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society</source>
          , ACM,
          <year>2022</year>
          , pp.
          <fpage>78</fpage>
          -
          <lpage>91</lpage>
          . doi:
          <volume>10</volume>
          . 1145/3514094.3534164.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          ,
          <article-title>Biases afecting human decision making in AI-supported second opinion settings</article-title>
          ,
          <source>in: Modeling Decisions for Artificial Intelligence</source>
          , Springer International Publishing,
          <year>2019</year>
          , pp.
          <fpage>283</fpage>
          -
          <lpage>294</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>030</fpage>
          -26773-5_
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Famiglini</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Evidence-based</surname>
            <given-names>XAI</given-names>
          </string-name>
          :
          <article-title>An empirical approach to design more efective and explainable decision support systems</article-title>
          ,
          <source>Computers in Biology and Medicine</source>
          <volume>170</volume>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .1016/j.compbiomed.
          <year>2024</year>
          .
          <volume>108042</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lundberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-I.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>A unified approach to interpreting model predictions</article-title>
          ,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv. 1705.07874.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ventura</surname>
          </string-name>
          , et al.,
          <article-title>Trusting deep learning natural-language models via local and global explanations</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>64</volume>
          (
          <year>2022</year>
          )
          <fpage>1863</fpage>
          -
          <lpage>1907</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10115-022-01690-9.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          , et al.,
          <article-title>"why should i trust you?": Explaining the predictions of any classifier</article-title>
          ,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .48550/ arXiv.1602.04938.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Parimbelli</surname>
          </string-name>
          , et al.,
          <article-title>Why did AI get this one wrong? - tree-based explanations of machine learning model predictions</article-title>
          ,
          <source>Artificial Intelligence in Medicine</source>
          <volume>135</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .1016/j.artmed.
          <year>2022</year>
          .
          <volume>102471</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bennetot</surname>
          </string-name>
          , et al.,
          <source>A practical guide on explainable AI techniques applied on biomedical use case applications</source>
          ,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2111.14260.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bach</surname>
          </string-name>
          , et al.,
          <article-title>On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation</article-title>
          ,
          <source>PloS one 10</source>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0130140</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Tanneru</surname>
          </string-name>
          , et al.,
          <article-title>Quantifying uncertainty in natural language explanations of large language models</article-title>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .48550/arXiv.2311.03533.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>R. M. J. Byrne</surname>
          </string-name>
          ,
          <article-title>Counterfactuals in explainable artificial intelligence (xai): Evidence from human reasoning</article-title>
          , IJCAI-
          <volume>19</volume>
          (
          <year>2019</year>
          )
          <fpage>6276</fpage>
          -
          <lpage>6282</lpage>
          . doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2019</year>
          /876.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>G.</given-names>
            <surname>Nicora</surname>
          </string-name>
          , et al.,
          <article-title>Bayesian networks in the management of hospital admissions: A comparison between explainable ai and black box ai during the pandemic</article-title>
          ,
          <source>Journal of Imaging</source>
          <volume>10</volume>
          (
          <year>2024</year>
          ). doi:
          <volume>10</volume>
          .3390/jimaging10050117.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>M.</given-names>
            <surname>Danilevsky</surname>
          </string-name>
          , et al.,
          <article-title>A survey of the state of explainable AI for natural language processing</article-title>
          ,
          <year>2020</year>
          . doi:10.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>