<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Trustworthy enough? Evaluation of an AI decision support system for healthcare professionals</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kristýna Sirka Kacafírková</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sara Polak</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Myriam Sillevis Smitt</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shirley A. Elprama</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>An Jacobs</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>imec-SMIT</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vrije Universiteit Brussel</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Exploring end-users' needs and involving them in designing and evaluating an AI system should be a priority. Understanding the system is essential to assess whether to trust it or not. This paper discusses a use case of a decision support system integrated into a platform for healthcare professionals (medical call operators and nurses). The system navigates them by predicting what intervention should be taken when an accident happens to a patient at home. Our use case demonstrates the importance of humancentred evaluation methods and potential struggles with mixed methods as detected by differences between qualitative and quantitative approaches. A subjective scale in combination with group interviews was used to evaluate the trust towards the system. The results showed that while users expressed a relatively high trust in the scale, the qualitative insights indicated uncertainty and the need for better explainability to trust the decision support system. In line with the results, we point out the need for better human-centred evaluation methods, as the current subjective scale needs to be complemented by qualitative methods to ensure rich insights.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;explainable AI (XAI)</kwd>
        <kwd>trustworthy AI</kwd>
        <kwd>DSS in healthcare</kwd>
        <kwd>XAI evaluation</kwd>
        <kwd>explainability needs 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The awareness of the need for explainability2 in AI systems is increasingly rising [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], especially in
sectors such as healthcare, where outcomes recommended by the system have a large impact
affecting humans’ lives. Nonetheless, explanations are often lacking or not adapted to end-users’
needs [
        <xref ref-type="bibr" rid="ref12 ref13 ref2 ref4">2, 4, 12, 13</xref>
        ]. To be able to detect the needs of end-users and enhance the system,
evaluation is crucial. However, the evaluation of explainability in AI systems from a
humancentred perspective is limited [
        <xref ref-type="bibr" rid="ref12 ref14 ref9">9, 12, 14</xref>
        ]. Developers mainly focus on the technical elements of
the explanation, and the user’s feedback is often overlooked [
        <xref ref-type="bibr" rid="ref11 ref9">9, 11</xref>
        ]. Therefore, we focus on
evaluation by end-users and relevant stakeholders regarding the explainability and
trustworthiness of a platform via existing tools. We show a use case of a healthcare platform with
a decision support system (DSS) developed in a Protego3 project. Concretely, we set out to answer
the following research question: What are the limitations of current explainability subjective scales
measuring trust in the DSS system based on AI?
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <sec id="sec-2-1">
        <title>2.1. Evaluating explainability, evaluating trust</title>
        <p>
          When evaluating explainability, it is crucial to acknowledge its link with trust [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Trust is
perceived as one of the main reasons for implementing XAI explanations [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Users seek
explanations primarily as an indicator of whether they can trust the outcome of a system or not
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. However, trust should not be seen as something instantly achievable. It is a long-term process
that the user develops continuously when using the system [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Using explanations should not
aim for over (or under) trusting the AI system but rather an appropriate level of trust [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
Especially in the healthcare context, too much trust can lead to fatal consequences, such as
following an incorrect diagnosis.
        </p>
        <p>
          Available human-centred evaluation methods for trust and explainability are limited [
          <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
          ].
One of the most used evaluation tools in the human-centred domain are subjective scales [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
Simply asking a participant if they trust the system via closed questions, however, often does not
sufficiently show the reasons behind it [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. To demonstrate the difficulties, we combined a
subjective scale with group interviews to show the importance of using both quantitative and
qualitative methods in evaluation.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Protego use case: Decision support system in healthcare</title>
      <sec id="sec-3-1">
        <title>3.1. System description</title>
        <p>A DSS was developed in collaboration with university researchers, industrial partners, and a local
healthcare organisation in Belgium. This system suggests the necessary type of care based on the
context of the alarm. A recommendation is based on health and behavioural data (from sensors
inside the patient’s home), personal information, and data collected by the call operator handling
the alarm. Based on this information, the system suggests a next step of action: 1) call an
ambulance, 2) send a nurse, 3) send an informal carer, 4) or dismiss the alarm. The call operator
then decides what will be the subsequent step. A comprehensive overview of the alarm and its
context is sent to the dispatched caregiver (e.g. a nurse), allowing them to follow up on the alarm
accurately.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data collection</title>
        <p>
          A Proof-of-Concept (PoC) based on the evaluation of mock-up versions was developed and
evaluated in two phases. For the first PoC group interviews, a pre-defined scenario of a diabetes
patient was used to evaluate the PoC with seven stakeholders4: 2 medical call operators (W, 32,
E-4, DW-4; W, 28, E-3, DW-3), 2 home nurses (W, 31, E-3, DW-3; W, 43, E-5, DW-10) and 3 ethical
board members. A refined PoC was then presented, using a pre-defined scenario of a heart
patient, in a final group interview with six stakeholders: one medical call-operator (W, 46, E-8,
DW-15), one home nurse (M, 39, E-10, DW-10) and 4 ethical board members. Participants were
invited to fill out a questionnaire at the end of each interview. For the first iteration, Cahour’s and
Forzy’s [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] trust scale was used with additional questions specific to the platform. The second
iteration also included Hoffman’s trust scale [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which is one of the most used scales for XAI
evaluation to see if there is a significant difference when the scale is adapted to XAI. All
participants signed informed consent. The sessions were organized at the research centre
premises, interviews were audio recorded, and notes were taken. The notes were subsequently
analysed trough a thematic analysis. Questionnaire data were summarized using GoogleSheets,
ChartExpo.
4 Where possible, we included data about: gender (W-woman, M-man, X-other), Age (in years), Experience in a role
(E-number of years), digital working experience (DW – number of years of using a computer in their job)
4. Results: Evaluation and insights from user’s feedback
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>4.1. Subjective scores on trust, predictability, reliability, and decision support</title>
        <p>The questionnaire data indicates that the overall level of trust in the system (for all participants)
is neutral (3) to high (4-5), see Figure 1. In general, home nurses expressed higher trust than call
operators, the ethical board were also less trustful than the professionals. Questions regarding
the user interface showed that the interpretation of percentages of suggested actions was
perceived as less understandable and thus lowered the trust score, which was also expressed
during interviews.</p>
      </sec>
      <sec id="sec-3-4">
        <title>4.2. The need for global explainability and simpler explanations</title>
        <p>A sufficient explanation of “why the percentages were generated” was missing, affecting users’
trust in the outcome. However, as participants said, instead of making the user interface more
transparent with more detailed information, they would be more interested in instructions on
how the system works and when it is updated. Call operators expressed a need to understand the
influence of responding to system questions on the actual recommendation. This might indicate
that there was a high need for a global explanation of the recommender system rather than an
explanation of each outcome (local explanation).</p>
        <p>Participants explained that during the process of action, there is often no time to explore
further what data means or why something was recommended. From a nurse’s point of view, in
an emergency, most of the information collected from the sensors is simply irrelevant and
unnecessary. Immediate action in such a situation is more important. Call operators also
expressed that they would decide on follow-up action according to other factors that can be
sensed only by previous experience, for example, nervosity in voice, difficulty with breathing etc.,
which AI cannot determine. On the other hand, call operators also indicated that personally
reaching a different conclusion than the percentages show would make them doubt their skills
and knowledge, and they would start to lose trust in themselves.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Discussion and conclusion</title>
      <p>The subjective scales indicated rather higher trust and explainability in the system, even though
no typical XAI explanation methods were integrated. However, as we found out during interviews,
participants expressed an unfulfilled need for global explainability and were seeking
explanations of how the system works. More than a need for local explanations, participants were
seeking to understand the logic behind the system and how answering questions in different ways
can affect the outcome.</p>
      <p>
        Furthermore, they tend to trust their experience more than the system. Even though the user
interface offered detailed information via percentages and sensors, these were not useful in the
work context of call operators and nurses, which is in line with Barda et al.’s study [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Too much
information can be overwhelming and not always easy to interpret in a timely manner, which
might also be the case if some XAI explanations had been implemented. This might indicate that
new simpler explanations for users with non-technical backgrounds are needed to be explored.
Nevertheless, it is important to point out that our study used a limited size of sample and was not
rich considering the types of stakeholders. The evaluation is also in the initial phase, more
iterations and capturing the trust over the time are needed. The use of pre-defined scenarios can
also not sufficiently reflect the situation in the real world.
      </p>
      <p>
        Besides, our study also showed that the evaluation of explainability and trust of AI systems
from a human-centred perspective is still limited [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and has room for improvement. For example,
ethical board members pointed out that the scale is too generic and hard to assess while filling in
the questionnaire. Based on the results, we demonstrate that a qualitative approach to end-users’
perception is necessary. Future work could focus on creating and validating scales reflecting
additional dimensions, such as users’ usefulness, understandability, and satisfaction [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that can
affect users’ trust.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work is part of the PROTEGO project, which is an imec.icon research project funded by
imec, Innoviris and Agentschap Innoveren &amp; Ondernemen.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Amann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.:
          <article-title>To explain or not to explain?-Artificial intelligence explainability in clinical decision support systems</article-title>
          .
          <source>PLOS Digital Health. 1</source>
          ,
          <issue>2</issue>
          ,
          <issue>e0000016</issue>
          (
          <year>2022</year>
          ). https://doi.org/10.1371/journal.pdig.
          <volume>0000016</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Barda</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          et al.:
          <article-title>A qualitative research framework for the design of user-centered displays of explanations for machine learning model predictions in healthcare</article-title>
          .
          <source>BMC Med Inform Decis Mak</source>
          .
          <volume>20</volume>
          ,
          <issue>1</issue>
          ,
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1186/s12911-020-01276-x.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Cahour</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forzy</surname>
          </string-name>
          , J.-F.:
          <article-title>Does projection into use improve trust and exploration? An example with a cruise control system</article-title>
          .
          <source>Saf Sci. 47</source>
          ,
          <issue>9</issue>
          ,
          <fpage>1260</fpage>
          -
          <lpage>1270</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rad</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey</article-title>
          . arXiv preprint arXiv:
          <year>2006</year>
          .
          <fpage>11371</fpage>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al.:
          <article-title>Measure Utility, Gain Trust: Practical Advice for XAI Researchers</article-title>
          . In: Proceedings - 2020 IEEE Workshop on TRust and EXpertise in Visual Analytics,
          <string-name>
            <surname>TREX</surname>
          </string-name>
          <year>2020</year>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          Institute of Electrical and Electronics Engineers Inc. (
          <year>2020</year>
          ). https://doi.org/10.1109/TREX51495.
          <year>2020</year>
          .
          <volume>00005</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ferrario</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>How Explainability Contributes to Trust in AI</article-title>
          . In: ACM International Conference Proceeding Series. pp.
          <fpage>1457</fpage>
          -
          <lpage>1466</lpage>
          Association for Computing Machinery (
          <year>2022</year>
          ). https://doi.org/10.1145/3531146.3533202.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Hoffman</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          et al.:
          <article-title>Metrics for Explainable AI: Challenges and Prospects Institute for Human and Machine Cognition</article-title>
          . arXiv preprint. (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Jacovi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          et al.:
          <article-title>Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in AI</article-title>
          .
          <source>In: FAccT 2021 - Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency</source>
          . pp.
          <fpage>624</fpage>
          -
          <lpage>635</lpage>
          Association for Computing Machinery,
          <string-name>
            <surname>Inc</surname>
          </string-name>
          (
          <year>2021</year>
          ). https://doi.org/10.1145/3442188.3445923.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.S.Y.</given-names>
          </string-name>
          et al.:
          <article-title>“Help Me Help the AI”: Understanding How Explainability Can Support Human-AI Interaction</article-title>
          .
          <source>In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          (
          <year>2023</year>
          ). https://doi.org/10.1145/3544548.3581001.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Lopes</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          et al.:
          <article-title>XAI Systems Evaluation: A Review of Human</article-title>
          and
          <string-name>
            <surname>Computer-Centred</surname>
            <given-names>Methods</given-names>
          </string-name>
          , (
          <year>2022</year>
          ). https://doi.org/10.3390/app12199423.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Morley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.:
          <article-title>From What to How: An Initial Review of Publicly Available AI Ethics Tools, Methods and Research to Translate Principles into Practices</article-title>
          .
          <source>Sci Eng Ethics</source>
          .
          <volume>26</volume>
          ,
          <issue>4</issue>
          ,
          <fpage>2141</fpage>
          -
          <lpage>2168</lpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1007/s11948-019-00165-5.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sperrle</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          et al.:
          <string-name>
            <surname>Should We Trust (X)AI</surname>
          </string-name>
          <article-title>? Design Dimensions for Structured Experimental Evaluations</article-title>
          . In: arXiv preprint arXiv:
          <year>2009</year>
          .
          <fpage>06433</fpage>
          . (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Waa</surname>
            ,
            <given-names>J. van der</given-names>
          </string-name>
          et al.:
          <article-title>Interpretable confidence measures for decision support systems</article-title>
          .
          <source>International Journal of Human Computer Studies</source>
          .
          <volume>144</volume>
          , (
          <year>2020</year>
          ). https://doi.org/10.1016/j.ijhcs.
          <year>2020</year>
          .
          <volume>102493</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.:
          <article-title>Evaluating the quality of machine learning explanations: A survey on methods and metrics</article-title>
          .
          <source>Electronics (Switzerland)</source>
          .
          <volume>10</volume>
          ,
          <issue>5</issue>
          ,
          <issue>593</issue>
          (
          <year>2021</year>
          ). https://doi.org/10.3390/electronics10050593.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>