<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>C. Fregosi);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Explanations to Preserve Human Agency</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Caterina Fregosi</string-name>
          <email>caterina.fregosi@unimib.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Natali</string-name>
          <email>chiara.natali@unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Cabitza</string-name>
          <email>federico.cabitza@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Human-AI Interaction, Clinical Decision Support System, Frictional AI</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IRCCS Ospedale Galeazzi-Sant'Ambrogio</institution>
          ,
          <addr-line>Via Cristina Belgioioso 173, Milano, 20157</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Applied Sciences and Arts of Southern Switzerland</institution>
          ,
          <addr-line>Via La Santa 1, Lugano, 6900</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Viale Sarca 336, Milano, 20126</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>What if AI in medicine didn't recommend clinicians what to do, but showed them what to consider? Artificial intelligence-based decision support systems (DSS) are increasingly integrated into clinical workflows to enhance diagnostic accuracy and reduce cognitive efort. However, empirical research shows that such systems may unintentionally foster automation bias, reduce critical engagement, and erode clinicians' sense of agency. Within this context, recent work in Human-AI Interaction has emphasized the role of design features that stimulate deliberation, such as cognitive friction and contrastive reasoning. Yet, little is known about how such interaction protocols afect users' perceived decision agency and diagnostic performance. Here we report a study in progress that evaluates the impact of contrastive explanation formats on clinicians' sense of agency, confidence, and diagnostic accuracy. We compare a Traditional DSS that provides a single recommendation with two “Judicial” protocols-Alternative and Antagonist-that introduce competing diagnoses and justifications to promote evaluative reasoning. The study adopts a mixed within- and between-subjects design involving medical students and clinicians, and includes a novel multidimensional HCI Sense of Agency scale tailored to diagnostic tasks. Data collection is currently ongoing.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Artificial intelligence-based decision support systems (DSS) are increasingly deployed in high-stakes
domains such as medicine, where diagnostic decisions carry significant consequences. These systems
promise to improve accuracy while reducing clinicians’ cognitive workload. Yet, their integration into
clinical practice also raises important concerns. Empirical studies show that AI-assisted decisions can
foster detrimental efects such as overreliance, automation bias, and a diminished sense of responsibility
and critical reasoning [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1, 2, 3, 4</xref>
        ]. Lee et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] conducted extensive surveys among knowledge workers,
revealing that the use of LLMs notably diminishes perceived cognitive efort involved in critical thinking
tasks. Crucially, their findings illustrate that confidence in the tool’s capability correlates inversely
with independent critical engagement. Workers displaying high confidence in AI outputs reported
lower cognitive efort but also reduced independent verification eforts. This observation is critical, as it
suggests that extensive reliance on AI systems, although eficient, risks creating cognitive complacency,
undermining the cultivation and exercise of critical judgement. In clinical contexts, where accountability
and deliberation are central to both professional identity and patient safety, these risks are particularly
pressing.
      </p>
      <p>
        At a functional level, the core challenge is to promote appropriate reliance: the ability to accept
AI-generated advice when it is correct and to reject it when it is not [
        <xref ref-type="bibr" rid="ref6 ref7 ref8 ref9">6, 7, 8, 9</xref>
        ]. Traditional DSSs
typically adopt an “Oracular” format [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], ofering a single, authoritative recommendation, often
accompanied by a persuasive explanation. Although this design may improve short-term performance,
      </p>
      <p>CEUR</p>
      <p>ceur-ws.org
it can also encourage passive acceptance, reduce deliberation, and gradually erode clinicians’ sense of
responsibility [11, 12] and even professional skill, in a phenomenon termed AI-induced deskilling [13].</p>
      <p>Explainable AI (XAI) has emerged as a response to these unintended efects, aiming to make system
outputs more interpretable and actionable [14]. However, growing empirical evidence suggests that
explanations alone do not guarantee improved outcomes; on the contrary, they may mislead users and
degrade decision quality [15, 16]. These findings indicate that efective explainability requires more
than transparency: it requires the alignment between what is presented, how it is interpreted, and the
behaviors it elicits.</p>
      <p>From this perspective, explanations should be understood not merely as outputs, but as components of
an interaction protocol [17]. Current systems often overlook the fact that users bring their own expertise,
mental models, and cognitive styles. Therefore, explanatory strategies must be both intelligible and
behaviorally efective—capable of promoting reflection, mitigating bias, and supporting a sense of agency.
Recent research has proposed the introduction of cognitive friction—deliberate design features that
challenge users’ reasoning to counteract passive reliance and encourage mindful engagement [18, 19, 20].
Within this line of work, the paradigm of Frictional AI encompasses methods that intentionally introduce
cognitive challenges to stimulate critical reflection in human–AI collaboration [ 21, 22, 23].</p>
    </sec>
    <sec id="sec-2">
      <title>2. Beyond Accuracy and Trust</title>
      <p>Most research on human–AI collaboration in clinical decision-making has focused on performance
metrics such as diagnostic accuracy or relational constructs such as trust and reliance [24]. While these
are critical outcomes, they overlook an equally important dimension: whether users remain active and
accountable participants in the decision-making process. This dimension is often described as the sense
of agency—the experiential perception of being the originator and owner of one’s actions and their
consequences [12]. In clinical contexts, agency is not merely a psychological state but a professional
requirement: clinicians must justify their diagnostic reasoning, assume responsibility for their decisions,
and maintain their diagnostic competence over time. If AI systems erode the sense that decisions are
genuinely one’s own, they risk undermining accountability and contributing to professional deskilling
over time [25, 13].</p>
      <p>For this reason, preserving users’ ability to experience decisions as authentically theirs has recently
been highlighted as a key outcome in the design of decision support systems. Several strands of research
have begun to propose alternative design paradigms that aim not only to inform users, but also to
preserve their critical reasoning and experiential sense of authorship. These approaches shift the focus
from explanations as tools of persuasion or transparency, toward interaction protocols that deliberately
foster accountability and user engagement. Hildebrandt [26] introduces the idea of agonistic machine
learning, where systems expose users to disagreement in order to prevent overconformity to algorithmic
advice.</p>
      <p>Miller [11] proposes a paradigm shift from recommendation-driven to hypothesis-driven decision
support that helps users evaluate competing alternatives. These approaches redefine the role of the
system from authoritative oracle to dialogical partner, foregrounding deliberation over persuasion.</p>
      <p>
        A second strand emphasizes the value of deliberate friction. Research on Frictional AI [21, 22, 23]
shows how cognitive efort can be purposefully introduced into the interaction to disrupt unreflective
reliance and stimulate critical engagement. Related work on cognitive forcing functions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] similarly
demonstrates that interventions designed to slow down or challenge the decision process can reduce
automation bias and foster more mindful judgment. While these strategies may temporarily increase
cognitive workload, they aim to preserve long-term competence and accountability in high-stakes
contexts.
      </p>
      <p>Other proposals explicitly focus on surfacing dissent. Haselager et al. [27] describe Reflection Machines ,
systems that create deliberate tension by highlighting counterarguments, prompting users to reconsider
initial intuitions. Reingold et al. [28] propose Dissenting Explanations, where models are trained to
generate adversarial viewpoints rather than converge toward consensus. Similarly, Sarkar argues for
systems that occasionally challenge the user [29]. He presents an intriguing conceptual shift from the
traditional view of AI as an assistant toward viewing AI as a provocateur. In contrast to AI simply
fulfilling user-directed tasks eficiently, a provocateur AI purposefully surfaces counter‑arguments,
fringe cases, or alternative framings, forcing the user to articulate, defend, and possibly revise their
stance. This reframing aligns AI systems with pedagogical objectives, positioning them as tools that
facilitate deeper cognitive engagement rather than merely augmenting productivity. These lines of
work share the assumption that disagreement, when well-structured, can protect against overreliance
and preserve users’ engagement as decision-makers.</p>
      <p>Despite these advances, empirical studies rarely assess whether such designs truly preserve users’
agency. Existing work typically relies on proxies such as trust, reliance, or satisfaction, which do not
capture whether decisions are still experienced as authentically one’s own. To address this gap we
propose the Judicial protocol, which operationalizes evaluative and frictional principles within clinical
DSSs. Unlike conventional DSSs that provide a single authoritative output, Judicial protocols present
two contrastive diagnostic alternatives, each supported by persuasive (yet fallible) justifications. The
goal is to re-engage users’ discriminative capacities by requiring them to actively adjudicate between
competing arguments. We implement this paradigm in two variants: the Alternative Judicial protocol,
in which a single system provides two alternative diagnoses with corresponding justifications, and the
Antagonist Judicial protocol, in which two separate systems each advocate for a diferent diagnosis.
By comparing these designs to the Traditional protocol, which ofers a single recommendation with
explanation, we aim to investigate how contrastive explanations afect not only diagnostic accuracy and
confidence, but also users’ perceived agency, responsibility, and the perceived role of AI in the
decisionmaking process. In parallel, we introduce a novel HCI Sense of Agency scale, tailored to the clinical
decision-making context, which decomposes agency into three dimensions: influence, ownership, and
responsibility. This contribution allows us to move beyond accuracy and trust, and to empirically assess
whether alternative interaction protocols can sustain clinicians’ agency in AI-supported diagnostic
practice.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Research questions</title>
      <p>We aim to address the following research questions:
1. RQ1: Are there significant diferences in users’ perceived sense of agency and responsibility
between the Traditional DSS and the Judicial explanation protocols?
2. RQ2: Are there significant diferences in diagnostic accuracy and user confidence between the</p>
      <p>Traditional DSS and the Judicial explanation format?
3. RQ3: Are there significant diferences in diagnostic accuracy and user confidence between the</p>
      <p>Antagonist and Alternative Judicial conditions?
4. RQ4: Are there significant diferences in the perceived influence and perceived utility of the AI
system between the Traditional DSS and the Judicial explanation formats?
5. RQ5: Are there significant diferences in the perceived utility, influence, sense of agency and
sense of responsibility between the Antagonist and Alternative Judicial protocols?
To further explore potential moderating efects, the sample will be stratified by level of clinical experience,
allowing us to assess whether these variables vary as a function of users’ expertise.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <sec id="sec-4-1">
        <title>4.1. Participants</title>
        <p>To test these hypotheses, a between- and within-subjects experimental design will be implemented (see
Figure 1). Participants, comprising medical students and clinicians from the University of Milan, will be
randomly assigned to one of two groups: Alternative Judicial or Antagonist Judicial. Responses to
the online survey will be collected anonymously.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Procedure</title>
        <p>The experiment was conducted through an online interface built using LimeSurvey1. Each participant
evaluated a total of 10 clinical cases, presented in a fixed and predefined order. In the first phase,
participants completed 5 cases supported by the Traditional DSS, which provided a single diagnostic
recommendation accompanied by an explanation. After each case, participants recorded their diagnostic
choice, rated their confidence in the decision, assessed the perceived utility of the system, and evaluated
the perceived complexity of the case on a 4-point ordinal scale. Upon completion of the five Traditional
cases, they filled in the Sense of Agency scale, which included the three constructs of influence , ownership,
and responsibility.</p>
        <p>In the second phase, participants completed 5 cases supported by one of the Judicial protocols,
depending on group assignment: in the Alternative Judicial condition, a single DSS presented two
alternative diagnoses with corresponding justifications, while in the Antagonist Judicial condition two
distinct DSSs each advocated for a diferent diagnosis with their own explanatory arguments. As in the
ifrst phase, after each case participants provided their diagnostic decision, confidence rating, system
utility rating, and case complexity rating. At the end of this block, they again completed the Sense of
Agency scale.</p>
        <p>Finally, after completing all 10 cases, participants filled out two standardized psychometric
instruments: the short version of the Big Five Inventory (BFI) to measure personality traits, and an adapted
version of the Decision Styles Scale (DSS) tailored to the diagnostic context, focusing on rational and
intuitive decision-making styles. The analyses aim to reveal both main efects and interaction efects
between the DSS format (Traditional vs. Judicial) and the explanation style (Alternative vs. Antagonist)
on key outcome variables.</p>
        <p>All AI recommendations in the study are simulated to ensure consistency across participants and
allow for full control over the diagnostic content. Clinical cases are adapted by an expert clinician from
The New England Journal of Medicine2 and include symptomatology, medical history, and lab results
1https://www.limesurvey.org/it
2https://www.nejm.org/
that together form a realistic diagnostic scenario. AI explanations in the Judicial conditions are crafted
to be persuasive and grounded in the clinical features of each case.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Measures</title>
        <p>The results of this study are evaluated through three straightforward measures—diagnostic accuracy,
self-reported confidence in one’s decision, and self-reported utility of the AI system—alongside a novel
instrument designed to capture participants’ sense of agency.</p>
        <p>Diagnostic accuracy provides the most direct and objective outcome, recorded dichotomously as
either correct (1) or incorrect (0), based on the reference diagnosis reported in the source medical
literature. This measure establishes a clear baseline against which the influence of diferent interaction
protocols can be assessed.</p>
        <p>Complementing this measure, participants also report their confidence in each decision and their
perceived utility of the AI system. Both constructs are measured using 4-point ordinal scales specifically
designed to reduce central tendency bias, thereby ensuring more discriminative responses. Confidence
ofers insight into the degree of certainty with which participants endorsed their diagnostic choices,
while perceived utility reflects the extent to which the system was judged to be supportive in practice.</p>
        <p>Since to our knowledge, there are currently no validated instruments in HCI or clinical
decisionmaking research to directly measure the sense of agency, we developed an ad hoc scale for this study.</p>
        <p>Building on prior work in psychology, the scale operationalizes agency across three complementary
constructs: Influence (the extent to which users felt steered or constrained by the system), Ownership
(the degree to which users experienced the decision as authentically their own), and Responsibility
(the extent to which users perceived themselves as accountable for the outcomes).</p>
        <p>The scale comprises 22 items, presented on a 4-point ordinal scale (1 = strongly disagree, 4 = strongly
agree), with an additional Not applicable option to account for cases where participants felt an item
did not reflect their experience. Negatively worded items (indicated as INV) were included to control
for acquiescence bias. To avoid priming efects, items from the three constructs were presented in a
randomized order rather than grouped by construct.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Limitations</title>
      <p>This study has several limitations that should be acknowledged. First, the study design may be subject
to learning and anchoring efects, as all participants completed the Traditional DSS block prior to the
Judicial block. This fixed ordering could artificially inflate or suppress performance and perceived
agency in the second block. Future iterations should incorporate counterbalancing or explicitly model
trial index to disentangle genuine protocol efects from order-related influences.</p>
      <p>Second, the generalizability of findings is constrained by the use of a single-institution convenience
sample consisting of medical students and clinicians drawn from one context. Moreover, the decision
support system was simulated rather than integrated into a real-world clinical workflow, which may
limit ecological validity and the applicability of results to actual practice environments.</p>
      <p>Third, the repeated administration of the agency scale after each block may introduce scale reactivity.
By repeatedly prompting participants to reflect on their sense of influence, ownership, and responsibility,
the measure itself may shape the construct under study. Future work should examine alternate forms
of the instrument, reduce measurement frequency, or introduce spacing strategies to minimize such
reactivity.</p>
      <p>Finally, the study’s reliance on binary accuracy scoring simplifies diagnostic reasoning to a
dichotomous correct/incorrect outcome. This approach may overlook partial correctness, reasonable diferential
diagnoses, or clinically meaningful prioritization of possibilities. Future research could adopt graded
scoring schemes or use expert panel adjudication to capture the nuance of diagnostic quality.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The growing discourse on frictional AI has emphasised the need for interaction protocols that move
beyond accuracy and trust as the only measures of success. Clinical decision support systems must
not only inform but also sustain clinicians’ capacity to deliberate, justify, and ultimately own their
decisions. Yet, empirical tools to assess such beyond-utility outcomes remain scarce. This study takes a
step in that direction by proposing a dual contribution: first, the introduction of judicial AI protocols
that operationalise contrastive and agonistic principles through Alternative and Antagonist designs;
and second, the development of a multidimensional Sense of Agency scale that captures clinicians’
perceived influence, ownership, and responsibility in diagnostic tasks.</p>
      <p>Together, these contributions set the ground for a systematic evaluation of whether deliberately
frictional interaction protocols can preserve professional agency without compromising diagnostic
performance. By reframing decision support from oracular recommendation to judicial adjudication,
our work foregrounds the clinician as an active arbiter rather than a passive recipient of algorithmic
advice. The proposed scale, in turn, ofers a means to empirically test this claim and to anchor design
choices in robust measures of agency.</p>
      <p>Looking ahead, the framework outlined here aspires to shift both research and design practice: from
interventions judged solely by short-term accuracy gains, towards systems that also safeguard the
longer-term values of accountability, responsibility, and professional competence. In doing so, we hope
to contribute to a broader reorientation of Human–AI Interaction—one in which diagnostic AI is not
only evaluated for what it adds, but also for what it preserves.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>C. Fregosi and F. Cabitza acknowledge funding support provided by the Italian project PRIN PNRR
2022 InXAID - Interaction with eXplainable Artificial Intelligence in (medical) Decision making. CUP:
H53D23008090001 funded by the European Union - Next Generation EU.</p>
      <p>C. Natali acknowledges the financial support provided by the Federal Commission for Scholarships for
Foreign Students in the form of the Swiss Government Excellence Scholarship (ESKAS No. 2024.0002)
for the academic year 2024-25.</p>
    </sec>
    <sec id="sec-8">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT-5 to improve readability (grammar,
style, and clarity). The tool was not used to generate ideas or substantive content, and all text was
reviewed and edited by the authors.
[11] T. Miller, Explainable ai is dead, long live explainable ai! hypothesis-driven decision support
using evaluative ai, in: Proceedings of the 2023 ACM conference on fairness, accountability, and
transparency, 2023, pp. 333–342.
[12] J. W. Moore, What is the sense of agency and why does it matter?, Frontiers in psychology 7
(2016) 1272.
[13] C. Natali, L. Marconi, L. D. Dias Duran, F. Cabitza, Ai-induced deskilling in medicine: A
mixedmethod review and research agenda for healthcare and beyond, Artificial Intelligence Review 58
(2025) 1–40.
[14] L. Longo, M. Brcic, F. Cabitza, J. Choi, R. Confalonieri, J. Del Ser, R. Guidotti, Y. Hayashi, F. Herrera,
A. Holzinger, et al., Explainable artificial intelligence (xai) 2.0: A manifesto of open challenges and
interdisciplinary research directions, Information Fusion (2024) 102301.
[15] G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, D. Weld, Does the whole exceed
its parts? the efect of ai explanations on complementary team performance, in: Proceedings of
the 2021 CHI conference on human factors in computing systems, 2021, pp. 1–16.
[16] F. Cabitza, C. Fregosi, A. Campagner, C. Natali, Explanations considered harmful: The impact of
misleading explanations on accuracy in hybrid human-ai decision making, in: World Conference
on Explainable Artificial Intelligence, Springer, 2024, pp. 255–269.
[17] F. Cabitza, L. Famiglini, C. Fregosi, S. Pe, E. Parimbelli, G. A. La Maida, E. Gallazzi, From oracular
to judicial: Enhancing clinical decision making through contrasting explanations and a novel
interaction protocol, in: Proceedings of the 30th International Conference on Intelligent User
Interfaces, 2025, pp. 745–754.
[18] A. Cooper, The inmates are running the asylum, Springer, 1999.
[19] A. L. Cox, S. J. Gould, M. E. Cecchinato, I. Iacovides, I. Renfree, Design frictions for mindful
interactions: The case for microboundaries, in: Proceedings of the 2016 CHI conference extended
abstracts on human factors in computing systems, 2016, pp. 1389–1397.
[20] Z. Chen, R. Schmidt, Exploring a behavioral model of “positive friction” in human-ai interaction,
in: International Conference on Human-Computer Interaction, Springer, 2024, pp. 3–22.
[21] F. Cabitza, A. Campagner, D. Ciucci, A. Seveso, Programmed ineficiencies in dss-supported human
decision making, in: Modeling Decisions for Artificial Intelligence: 16th International Conference,
MDAI 2019, Milan, Italy, September 4–6, 2019, Proceedings 16, Springer, 2019, pp. 201–212.
[22] C. Natali, et al., Per aspera ad astra, or flourishing via friction: Stimulating cognitive activation
by design through frictional decision support systems, in: CEUR workshop proceedings, volume
3481, CEUR-WS, 2023, pp. 15–19.
[23] F. Cabitza, C. Natali, L. Famiglini, A. Campagner, V. Caccavella, E. Gallazzi, Never tell me the odds:
Investigating pro-hoc explanations in medical decision making, Artificial intelligence in medicine
150 (2024) 102819.
[24] C. Natali, A. Campagner, F. Cabitza, Answering the call to go beyond accuracy: An online tool for
the multidimensional assessment of decision support systems., in: BIOSTEC (2), 2024, pp. 219–229.
[25] Y. S. J. Aquino, W. A. Rogers, A. Braunack-Mayer, H. Frazer, K. T. Win, N. Houssami, C. Degeling,
C. Semsarian, S. M. Carter, Utopia versus dystopia: professional perspectives on the impact
of healthcare artificial intelligence on clinical roles and skills, International Journal of Medical
Informatics 169 (2023) 104903.
[26] M. Hildebrandt, Privacy as protection of the incomputable self: From agnostic to agonistic machine
learning, Theoretical Inquiries in Law 20 (2019) 83–121.
[27] P. Haselager, H. Schrafenberger, S. Thill, S. Fischer, P. Lanillos, S. Van De Groes, M. Van Hoof,
Reflection machines: Supporting efective human oversight over medical decision support systems,
Cambridge Quarterly of Healthcare Ethics 33 (2024) 380–389.
[28] O. Reingold, J. H. Shen, A. Talati, Dissenting explanations: Leveraging disagreement to reduce
model overreliance, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38,
2024, pp. 21537–21544.
[29] A. Sarkar, Ai should challenge, not obey, Communications of the ACM 67 (2024) 18–21.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kliegr</surname>
          </string-name>
          , Š. Bahník,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fürnkranz</surname>
          </string-name>
          ,
          <article-title>A review of possible efects of cognitive biases on interpretation of rule-based machine learning models</article-title>
          ,
          <source>Artificial Intelligence</source>
          <volume>295</volume>
          (
          <year>2021</year>
          )
          <fpage>103458</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Buçinca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Malaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Z.</given-names>
            <surname>Gajos</surname>
          </string-name>
          ,
          <article-title>To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making, Proceedings of the ACM on Human-computer Interaction 5 (</article-title>
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cabitza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Campagner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Angius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Natali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reverberi</surname>
          </string-name>
          ,
          <article-title>Ai shall have no dominion: on how to measure technology dominance in ai-supported human decision-making</article-title>
          ,
          <source>in: Proceedings of the 2023 CHI conference on human factors in computing systems</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vered</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Livni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. D. L.</given-names>
            <surname>Howe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Sonenberg,</surname>
          </string-name>
          <article-title>The efects of explanations on automation bias</article-title>
          ,
          <source>Artificial Intelligence</source>
          <volume>322</volume>
          (
          <year>2023</year>
          )
          <fpage>103952</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.-P. H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Tankelevitch</surname>
          </string-name>
          , I. Drosos,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rintel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Banks</surname>
          </string-name>
          , N. Wilson,
          <article-title>The impact of generative ai on critical thinking: Self-reported reductions in cognitive efort and confidence efects from a survey of knowledge workers (</article-title>
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>See</surname>
          </string-name>
          ,
          <article-title>Trust in automation: Designing for appropriate reliance</article-title>
          ,
          <source>Human factors 46</source>
          (
          <year>2004</year>
          )
          <fpage>50</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schemmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kuehl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Benz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bartos</surname>
          </string-name>
          , G. Satzger,
          <article-title>Appropriate reliance on ai advice: Conceptualization and the efect of explanations</article-title>
          ,
          <source>in: Proceedings of the 28th International Conference on Intelligent User Interfaces</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>410</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Schmitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wambsganß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Söllner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Janson</surname>
          </string-name>
          ,
          <article-title>Towards a trust reliance paradox? exploring the gap between perceived trust in and reliance on algorithmic advice</article-title>
          ,
          <source>in: Proceedings of the International Conference on Information Systems (ICIS)</source>
          <year>2021</year>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Hartline</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hullman</surname>
          </string-name>
          ,
          <article-title>A decision theoretic framework for measuring ai reliance</article-title>
          ,
          <source>in: The 2024 ACM Conference on Fairness, Accountability, and Transparency</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>221</fpage>
          -
          <lpage>236</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Masarie</surname>
          </string-name>
          <string-name>
            <surname>Jr</surname>
          </string-name>
          ,
          <article-title>The demise of the “greek oracle” model for medical diagnostic systems</article-title>
          ,
          <source>Methods of information in medicine 29</source>
          (
          <year>1990</year>
          )
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>