<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>HiTZ@Antidote: Argumentation-driven Explainable Artificial Intelligence for Digital Medicine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rodrigo Agerri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iñigo Alonso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aitziber Atutxa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ander Berrondo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ainara Estarrona</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iker Garcia-Ferrero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iakes Goenaga</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Koldo Gojenola</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maite Oronoz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Perez-Tejedor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>German Rigau</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anar Yeginbergenova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HiTZ Center - Ixa, University of the Basque Country UPV/EHU</institution>
          ,
          <addr-line>Donostia-San Sebastián</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Providing high quality explanations for AI predictions based on machine learning is a challenging and complex task. To work well it requires, among other factors: selecting a proper level of generality/specificity of the explanation; considering assumptions about the familiarity of the explanation beneficiary with the AI task under consideration; referring to specific elements that have contributed to the decision; making use of additional knowledge (e.g. expert evidence) which might not be part of the prediction process; and providing evidence supporting negative hypothesis. Finally, the system needs to formulate the explanation in a clearly interpretable, and possibly convincing, way. Given these considerations, ANTIDOTE fosters an integrated vision of explainable AI, where low-level characteristics of the deep learning process are combined with higher level schemes proper of the human argumentation capacity. ANTIDOTE will exploit cross-disciplinary competences in deep learning and argumentation to support a broader and innovative view of explainable AI, where the need for high-quality explanations for clinical cases deliberation is critical. As a first result of the project, we publish the Antidote CasiMedicos dataset to facilitate research on explainable AI in general, and argumentation in the medical domain in particular.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Explainable AI</kwd>
        <kwd>Digital Medicine</kwd>
        <kwd>Question Answering</kwd>
        <kwd>Argumentation</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>under consideration, (iii) referring to specific elements
that have contributed to the decision, (iv) making use of
ANTIDOTE1 is a European CHIST-ERA project where additional knowledge (e.g. metadata) which might not be
each partner is funded by their national Science Agencies. part of the prediction process, (v) selecting appropriate
As the Spanish partner in the Consortium is the HiTZ examples and, (vi) providing evidence supporting
negaCenter - Ixa, from the University of the Basque Coun- tive hypotheses. Finally, the system needs to formulate
try UPV/EHU, the project was funded by the Proyectos the explanation in a clearly interpretable, and possibly
de Colaboración Internacional (PCI 2020) program of the convincing, way.</p>
      <p>Spanish Ministry of Science and Innovation. The other Taking into account these considerations, ANTIDOTE
European partners are the following: Université Côte fosters an integrated vision of Explainable AI (XAI),
d’Azur (UCA) from France and coordinators of the in- where the low-level characteristics of the deep learning
ternational consortium, Fondazione Bruno Kessler (FBK) process are combined with higher level schemes proper
from Italy, KU Leuven/Computer Science, in Belgium and of human argumentation. Following this, the ANTIDOTE
Universidade Nova de Lisboa (NOVA) in Portugal. integrated vision is supported by three considerations.</p>
      <p>The aim of ANTIDOTE is to exploit cross-disciplinary First, in neural architectures the correlation between
incompetences in three areas, namely, deep learning, ar- ternal states of the network (e.g., weights assumed by
gumentation and interactivity, to support a broader and single nodes) and the justification of the network
classifiinnovative view of explainable AI. cation outcome is not well studied. Second, high quality</p>
      <p>Providing high quality explanations for AI predictions explanations are crucially based on argumentation
mechbased on machine learning is a challenging and complex anisms (e.g., provide supporting examples and rejected
task. To work well it requires, among other aspects: (i) alternatives). Finally, in real settings, providing
explanaselecting a proper level of generality/specificity of the tions is inherently an interactive process involving the
explanation, (ii) considering assumptions about the fa- system and the user.
miliarity of the explanation beneficiary with the AI task Thus, ANTIDOTE will exploit cross-disciplinary
competences in three areas, namely, deep learning,
argumentation and interactivity, to support a broader and
innovative view of explainable AI. There are several research
challenges that ANTIDOTE will address to advance the
state-of-the-art in explainable AI.</p>
      <p>The first challenge is to take advantage of the huge
SEPLN-PD 2023: Annual Conference of the Spanish Association for
Natural Language Processing 2023: Projects and System
Demonstrations
$ rodrigo.agerri@ehu.eus (R. Agerri)</p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)</p>
      <p>1https://univ-cotedazur.eu/antidote
body of past research on argumentation to complement
state-of-the-art approaches on explainability. In addition,
the recent resurgence of AI highlights the idea that
lowlevel system behavior not only needs to be interpretable
(e.g., showing those elements that most contributed to the
system decision), but that also needs to be joined by high
level human argumentation schemes. The second
challenge is to automatically learn explanatory
argumentation schemas in Natural Language (NL) and to efectively
combine evidence-based decision making with high level
explanations. The third challenge for ANTIDOTE is that
a task-specific prediction model and a general
argumentation model need to be combined to produce explanatory
argumentations.</p>
      <p>While neural networks for medical diagnosis have
become exceedingly accurate in many areas, their ability
to explain how they achieve their outcome remains
problematic. Herein lies the main novelty of the ANTIDOTE
project: it focuses on elaborating argumentative
explanations to diagnosis predictions in order to assist student
clinicians to learn making informed decisions.</p>
      <p>The explanatory argumentative scenario envisaged by
ANTIDOTE will involve a student clinician, who will
need to hypothesize about the clinical case of a patient
and will have to provide argumentative explanations
about them. The focus of the experimental setting is
set on the capacity of the ANTIDOTE Explanatory AI
system to provide correct predictions and consistent
arguments, without forgetting also the linguistic quality of
the dialogues (e.g., naturalness of the utterances, etc.).</p>
      <p>In our scenario depicted in Figure 1, the clinician
queries the ANTIDOTE XAI for explanations (arguments)
on its diagnosis of the clinical case. The ANTIDOTE XAI
provides hypotheses (diferential diagnosis) about the
clinical case, as well as arguments to support its
prediction and arguments discarding alternative predictions.</p>
      <p>The student clinician has the possibility to take the
initiative to ask additional questions and clarifications. The
goal of the explanatory argumentation in a diferential
diagnosis is to validate the correctness of the
diagnosis and the ANTIDOTE XAI capacity to argue in favour
of the correct hypothesis and to counter-argue against
alternative hypotheses.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>In this section we review the most relevant previous work focusing on argumentation and explainable AI for the medical domain.</title>
        <sec id="sec-2-1-1">
          <title>2.1. Argumentation mining and generation</title>
          <p>Argumentation mining is a research area that moves
between natural language processing, argumentation
theory and information retrieval. The aim of argumentation
mining is to automatically detect the argumentation of
a document and its structure. This implies the
detection of all the arguments involved in the argumentation
process, their individual or local structure (rhetorical or
argumentative relationships between their propositions),
and the interactions between them, namely, the global
argumentation structure.</p>
          <p>Argumentation mining in Natural Language
Processing has been applied to various domains such as
persuasive essays, legal documents, political debates and
social media data [1]. For instance, Stab and Gurevych
[2] built an annotated dataset of persuasive essays with
corresponding argument components and relations.
Using this corpus, Eger et al. [3] developed an end-to-end
neural method for argument structure identification.
Furthermore, Nguyen and Litman [4] also applied an
endto-end method to parse argument structure and used the
argument structure features to improve automated
persuasive essay scoring. Other approaches studied
contextdependent claim detection by collecting annotations for
Wikipedia articles [5]. Using this corpus, the task of
automatically identifying the corresponding pieces of
evidence given a claim has also been investigated [6].</p>
          <p>Argumentation generation remains a research area
in which there is still a long way to go. Recent work
has made progress towards this goal through the
automated generation of argumentative text [7, 8, 9, 10].
Thus, Alshomary et al. [11] proposed a Bayesian
argument generation system to generate arguments given the
corresponding argumentation strategies. Sato et al. [9]
presented a sentence-retrieval-based end-to-end
argument generation system that can participate in English
debating games.</p>
          <p>Adadi and Berranda [18] presented an extensive
literature review, collecting and analyzing 381 diferent
scientific papers between 2004 and 2018. They arranged
all of the scientific work in the field of explainable AI
along four main axes and stressed the need for more
formalism to be introduced in the field of XAI and for more
interaction between humans and machines.</p>
          <p>In a more recent study [19] introduced a diferent
type of arrangement that initially distinguishes
transparent and post-hoc methods and subsequently created
sub-categories.</p>
          <p>Taking into account argumentation principles,
ANTIDOTE will explain machine decisions based on four
modes of explanations to be auditable by humans: (i)
analytic statements in NL that describe the elements and
context that support a choice, (ii) visualizations that
highlight portions of the raw data that support a choice, (iii)
cases that invoke specific examples, and (iv) rejections
of alternative choices that argue against less preferred
answers based on analytics, cases, and data.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2Source from DARPA XAI program: https://www.darpa.mil/</title>
        <p>program/explainable-artificial-intelligence</p>
      </sec>
      <sec id="sec-2-3">
        <title>There have also been some works exploring counter</title>
        <p>argument generation to select the main talking points to
generate a counter-argument [12]. In this line of research,
Hidey and McKeown [13] proposed a neural model that
edited the original claim semantically to produce a claim
with an opposing stance. They also incorporated external
knowledge into the encoder-decoder architecture show- 3. Methology and Work Plan
ing that their model generated arguments that were more
likely to be on topic. The main scientific challenge for the project is the
com</p>
        <p>Finally, an autonomous debating system (Project De- bination of three models depicted in Figure 1: (1) The
bater) able to engage in competitive debates with humans Prediction Model has to predict appropriate International
was developed. The system consisted of a pipeline of four Classification of Diseases (ICD) codes given a clinical
main modules: argument mining, an argument knowl- case; (2) the Argumentative Model selects proper
arguedge base, argument rebuttal, and debate construction ments (i.e., entity and relations) to support or attack a
[14]. given topic. It may use both information included in the
clinical cases used by the prediction model and additional
2.2. Explainable AI sources of knowledge; (3) the Interaction Model provides
argumentative explanations about a certain prediction.</p>
        <p>Explainable artificial intelligence (XAI) aims to address An integrated approach is proposed to both predict the
the needs of users wanting to understand how a pro- outcome of a clinical course of action and justify a
medigram’s artificial intelligence works and how to evaluate cal diagnosis by a language model. A starting point will
the results obtained. Otherwise, there is no basis for real be using current large language models [20] to generate
confidence in the work of the AI system, as illustrated appropriate explanations guided by the activated view on
by Figure 22 The transparency ofered by explainable a textual snippet that contributed to the decision, namely,
AI is therefore essential for the acceptance of artificial the argument for the decision.
intelligence.</p>
        <p>There has been a surge of interest in explainable artifi- 3.1. Work Plan
cial intelligence (XAI) in recent years. This has produced
a myriad of algorithmic and mathematical methods to The Work Plan is structured in six Work Packages of
explain the inner workings of machine learning models which three are focused on the scientific contributions
[15]. However, despite their mathematical rigor, these of the project.
works sufer from a lack of usability and practical
interpretability for real users. Although the concepts of
interpretability and explainability are hard to rigorously
define, multiple attempts have been made towards that
goal [16, 17].</p>
      </sec>
      <sec id="sec-2-4">
        <title>WP2: Methodology and Design (Leader: FBK). Partic</title>
        <p>ipants: UCA, UPV/EHU, KU, NOVA. The
purpose of WP2 is to define, adapt and integrate the
modules, resources, data structures, data formats
and module APIs of the ANTIDOTE architecture.
This includes designing the experiments, datasets,
standard protocols, information flow and main
architecture of ANTIDOTE.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Ongoing Work</title>
      <sec id="sec-3-1">
        <title>WP3: Machine Learning (ML) for predicting clinical</title>
        <p>outcomes (Leader: KU). Participants: UPV/EHU,
FBK, UCA, NOVA. WP3 targets (1) the develop- There are a number of tasks currently being undertaken
ment of a multitask learning model to jointly pre- within the project. In this section we provide details of
dict and justify a medical diagnosis by a deep the most central ones with respect to the objectives and
learning model; (2) surfacing and making explicit motivation provided in the introduction.
the underlying aspects (identification of the most
relevant/informative terms, identification of re- 4.1. ANTIDOTE Datasets
lations among terms) driving neural network
decisions during the diagnosis prediction process; In order to carry out the tasks related to the main use-case
(3) retrieve external information to support the presented in Figure 1, we need to identify, collect and
explanation. annotate the most suitable corpus with which to train
WP4: Explanatory arguments in natural language diferent models. In this regard, we have identified two
(Leader: UCA). Participants: UPV/EHU, FBK, possible data sources that will help us meet our objectives
UCA, NOVA. WP4 relies on the textual arguments and that will constitute an important contribution of the
that form the basis for the decisions generated in ANTIDOTE project: SAEI and CasiMedicos.
WP3. WP4 targets (1) the definition and
analysis of explanatory argumentative patterns to be The SAEI Corpus is a collection of diferential
diagused to construct natural language explanatory nosis in Spanish collected by La Sociedad Andaluza de
arguments of predictions; (2) the creation of a re- Enfermedades Infecciosas (The Andalusian Society of
Insource of annotated natural language explanatory fectious Diseases)3. This society is a non-profit
associaarguments; (3) the development of explanatory tion formed almost entirely by physicians specializing
arguments in natural language by mining and in Internal Medicine with special dedication to the
mancollecting them from trusted textual resources in agement of infectious diseases, whose general purpose
the medical domain. is the promotion and development of this medical
disciWP5: Evaluation (use cases in healthcare) (Leader: pline (training, care and research). We have selected the
UPV/EHU). Participants: KU, UCA, FBK, NOVA. books that are of interest to us in order to carry out our
WP5 aims to (1) evaluate the efectiveness and objectives, namely, those that include clinical cases of
quality of the prediction and the plausible alterna- infectious diseases for residents that are available, for the
tives (2) the quality of the generated explanatory years 2011, 2015, 2016, 2017 and 2020. Among all these
arguments regarding the supporting evidence books we have extracted cleaned and pre-processed a
found in the clinical case in favor of the prediction total of 244 clinical cases with diferential diagnosis.
and the positive or negative evidence found to
discard other plausible alternatives, (3) the intrinsic
quality of the generated arguments.</p>
      </sec>
      <sec id="sec-3-2">
        <title>CasiMedicos is a community and collaborative med</title>
        <p>ical project run by volunteer medical doctors4. Among
all the information created and made publicly available
3.2. Evaluation by this collaborative project, we have identified as an
adequate data source the MIR exams commented by
volThe generation of arguments will be quantitatively eval- untary medical doctors with the aim of providing answers
uated by computing metrics used in text generation to and explanations5 to the MIR exams annually published
measure their overlap with ground truth arguments [21]. by the Spanish Ministry of Health. In this data source
Moreover, the argumentative model will be evaluated we have extracted and pre-processed 622 commented
following the criteria of coherence, simplicity, and gen- questions from the MIR exams held between the years
erality [22]: explanations with structural simplicity, co- 2005, 2014, 2016, 2018, 2019, 2020, 2021 and 2022. The
herence, or minimality are preferred. With respect to cleaned corpus, named the Antidote Casimedicos dataset,
argument mining, standard metrics such as F1 and accu- is publicly available to encourage research on explainable
racy will be used. AI in the medical domain in general, and argumentation</p>
        <p>Generation of explanatory arguments will be also qual- in particular6.
itative evaluated by medical students. Given the objec- Unlike popular Question Anwering (QA) datasets for
tives and context of the project, ANTIDOTE will be based English based on medical exams [24], both SAEI and
on previous work by Johnson [23], whereby the argu- CasiMedicos include not only the explanations for the
ments will be evaluated for their (informal) inferential
structure in terms of acceptability, relevance, and sufi- 43hhttttppss::////wwwwww..scaaesiim.oregdicos.com/
ciency of reasons provided, as well as their answerability 5https://www.casimedicos.com/mir-2-0/
to human agents’ doubts and objections. 6https://github.com/ixa-ehu/antidote-casimedicos
corrent answer (diagnosis or treatment), but also explana- approaches to Question Answering techniques in the
tory arguments written by medical doctors explaining medical domain.
why the rest of the possible answers are incorrect.</p>
        <p>After pre-processing, these datasets have been
translated from Spanish to English with the objective of start- 5. Concluding Remarks
ing various annotation tasks at various levels of
complexity: (i) linking the explanatory sequences with
respect to each possible answer; (ii) labeling of hierarchical
argumentative structures; (iii) discourse markers. The
resulting corpus will be the first corpus (multilingual or
otherwise) with this type of annotations for the medical
domain. Whenever ready, the corpus will be distributed
under a free license to promote further research and to
ensure reproducibility or results.</p>
      </sec>
      <sec id="sec-3-3">
        <title>In this paper we provide a description of the ANTIDOTE</title>
        <p>project, mostly focusing on identifying and generating
high-quality argumentative explanations for AI
predictions in the medical domain. So far, ongoing work has
been focused on dataset collection and annotation and
novel experimental work on Question Answering and
Crosslingual Argument Mining. This work has leveraged
multilingual encoder and decoder large language models
[28, 20] for both extractive and generative
experimentation.</p>
        <p>Still, providing high-quality explanations for AI
predictions based on machine learning is a challenging and
complex task [24]. To work well, it requires, among other
factors, making use of additional knowledge (e.g.
medical evidence) which might not be part of the prediction
process, and providing evidence supporting negative
hypotheses. With these issues in mind, ANTIDOTE aims to
address the challenge of providing an integrated vision
of explainable AI, where low-level characteristics of the
deep learning process are combined with higher level
schemes proper of the human argumentation capacity. In
order to do so, ANTIDOTE will be focused on a number
of deep learning tasks for the medical domain, where
the need for high quality explanations for clinical cases
deliberation is critical.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>4.2. Question Answering in the Medical</title>
      </sec>
      <sec id="sec-4-2">
        <title>Domain</title>
        <sec id="sec-4-2-1">
          <title>While there are several QA datasets for English based</title>
          <p>on medical exams [24], none of the previously published
works contain two features which are unique of both
SAEI and CasiMedicos: (i) the presence of explanations
for both correct and incorrect answers; (ii) an
argumentative structure arguing and counter-arguing about the
possible answers. These features make it possible to
define new Question Answering tasks, both from an
extractive and generative point of view. In extractive QA,
the objective would consist of identifying, in a given
context, the explanation to the correct answer. In terms of
generative QA, it will also allow us to leverage large
language models [20] to learn generating the explanatory
arguments with respect to both correct and incorrect
possible answers.
4.3. Crosslingual Knowledge Transfer We thank the CasiMedicos Proyecto MIR 2.0 for
their permission to share their data for research
purThe only corpus annotated with argumentative structure poses. ANTIDOTE (PCI2020-120717-2) is a project
currently available for the medical domain is the AbstRCT funded by MCIN/AEI/10.13039/501100011033 and by
dataset, which consists of English clinical trials [25]. In European Union NextGenerationEU/PRTR. Rodrigo
order to investigate the diferent strategies of transfer- Agerri currently holds the RYC-2017-23647 fellowship
ring knowledge from English to other languages, espe- (MCIN/AEI/10.13039/501100011033 and by ESF
Investcially those applying model- and data-transfer techniques ing in your future). Iker García-Ferrero is supported
previously discussed for other application domains [26], by a doctoral grant from the Basque Government
ongoing work is focused on adapting such knowledge (PRE_2021_2_0219) and Anar Yeginbergenova
acknowltransfer techniques for argumentation in the medical do- edges the PhD contract from the UPV/EHU (PIF 22/159).
main. As a result, we are undertaking novel experimental
work on argument mining in Spanish for the medical
domain [27]. This also involves the generation of the first References
Spanish dataset annotated with argumentative structures
for the medical domain. Finally, the plan is to apply the [1] M. Dusmanu, E. Cabrio, S. Villata, Argument
mindeveloped technique to other languages of interest for ing on twitter: Arguments, facts and sources, in:
the ANTIDOTE project (French and Italian). Proceedings of the 2017 Conference on Empirical</p>
          <p>In this line of research, and taking as starting point Methods in Natural Language Processing, 2017, pp.
the ongoing work mentioned in the previous section, 2317–2322.
we plan to investigate also crosslingual and multilingual
[2] C. Stab, I. Gurevych, Parsing argumentation struc- Nature 591 (2021) 379–384.</p>
          <p>tures in persuasive essays, Computational Linguis- [15] O. Biran, C. Cotton, Explanation and justification in
tics 43 (2017) 619–659. machine learning: A survey, in: IJCAI-17 workshop
[3] S. Eger, J. Daxenberger, I. Gurevych, Neural end- on explainable AI (XAI), volume 8, 2017, pp. 8–13.
to-end learning for computational argumentation [16] Z. C. Lipton, The Mythos of Model Interpretability:
mining, arXiv preprint arXiv:1704.06104 (2017). In machine learning, the concept of interpretability
[4] H. Nguyen, D. Litman, Argument mining for im- is both important and slippery., Queue 16 (2018)
proving the automated scoring of persuasive essays, 31–57.
in: Proceedings of the AAAI Conference on Artifi- [17] F. Doshi-Velez, B. Kim, Towards a rigorous science
cial Intelligence, volume 32, 2018. of interpretable machine learning, arXiv preprint
[5] R. Levy, Y. Bilu, D. Hershcovich, E. Aharoni, arXiv:1702.08608 (2017).</p>
          <p>N. Slonim, Context dependent claim detection, [18] A. Adadi, M. Berrada, Peeking inside the black-box:
in: Proceedings of COLING 2014, the 25th Inter- a survey on explainable artificial intelligence (XAI),
national Conference on Computational Linguistics: IEEE access 6 (2018) 52138–52160.</p>
          <p>Technical Papers, 2014, pp. 1489–1500. [19] A. B. Arrieta, N. Díaz-Rodríguez, J. Del Ser, A.
Ben[6] R. Rinott, L. Dankin, C. Alzate, M. M. Khapra, netot, S. Tabik, A. Barbado, S. García, S. Gil-López,
E. Aharoni, N. Slonim, Show me your evidence-an D. Molina, R. Benjamins, et al., Explainable
Arautomatic method for context dependent evidence tificial Intelligence (XAI): Concepts, taxonomies,
detection, in: Proceedings of the 2015 conference opportunities and challenges toward responsible
on empirical methods in natural language process- AI, Information fusion 58 (2020) 82–115.
ing, 2015, pp. 440–450. [20] L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou,
[7] R. Bar-Haim, L. Eden, R. Friedman, Y. Kantor, D. La- A. Siddhant, A. Barua, C. Rafel, mT5: A massively
hav, N. Slonim, From arguments to key points: To- multilingual pre-trained text-to-text transformer,
wards automatic argument summarization, arXiv in: NAACL, 2021.</p>
          <p>preprint arXiv:2005.01619 (2020). [21] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger,
[8] X. Hua, L. Wang, Neural argument generation aug- Y. Artzi, BERTScore: Evaluating Text Generation
mented with externally retrieved evidence, arXiv with BERT, in: International Conference on
Learnpreprint arXiv:1805.10254 (2018). ing Representations (ICLR), 2020.
[9] M. Sato, K. Yanai, T. Miyoshi, T. Yanase, [22] T. Miller, Explanation in artificial intelligence:
InM. Iwayama, Q. Sun, Y. Niwa, End-to-end argument sights from the social sciences, Artificial
intelligeneration system in debating, in: Proceedings of gence 267 (2019) 1–38.</p>
          <p>ACL-IJCNLP 2015 System Demonstrations, 2015, [23] R. H. Johnson, Manifest rationality: A pragmatic
pp. 109–114. theory of argument, Routledge, 2012.
[10] M. Alshomary, S. Syed, A. Dhar, M. Potthast, [24] K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei,
H. Wachsmuth, Argument Undermining: Counter- H. W. Chung, N. Scales, A. Tanwani, H.
ColeArgument Generation by Attacking Weak Premises, Lewis, S. Pfohl, P. Payne, M. Seneviratne, P. Gamble,
arXiv preprint arXiv:2105.11752 (2021). C. Kelly, N. Scharli, A. Chowdhery, P. Mansfield,
[11] M. Alshomary, S. Syed, M. Potthast, H. Wachsmuth, B. A. y Arcas, D. Webster, G. S. Corrado, Y. Matias,
Target inference in argument conclusion genera- K. Chou, J. Gottweis, N. Tomasev, Y. Liu, A.
Ration, in: Proceedings of the 58th Annual Meeting jkomar, J. Barral, C. Semturs, A. Karthikesalingam,
of the Association for Computational Linguistics, V. Natarajan, Large language models encode clinical
2020, pp. 4334–4345. knowledge, in: arXiv 2212.13138, 2022.
[12] X. Hua, Z. Hu, L. Wang, Argument generation with [25] T. Mayer, S. Marro, E. Cabrio, S. Villata, Enhancing
retrieval, planning, and realization, arXiv preprint Evidence-Based Medicine with Natural Language
arXiv:1906.03717 (2019). Argumentative Analysis of Clinical Trials, Artificial
[13] C. Hidey, K. McKeown, Fixed that for you: Gen- Intelligence in Medicine (2021) 102098.
erating contrastive claims with semantic edits, in: [26] I. García-Ferrero, R. Agerri, G. Rigau, Model and
Proceedings of the 2019 Conference of the North data transfer for cross-lingual sequence labelling in
American Chapter of the Association for Computa- zero-resource settings, in: Findings of the
Associational Linguistics: Human Language Technologies, tion for Computational Linguistics: EMNLP 2022,
Volume 1 (Long and Short Papers), 2019, pp. 1756– 2022.</p>
          <p>1767. [27] A. Yeginbergenova, R. Agerri, Cross-lingual
ar[14] N. Slonim, Y. Bilu, C. Alzate, R. Bar-Haim, B. Bogin, gument mining in the medical domain, in: arXiv
F. Bonin, L. Choshen, E. Cohen-Karlik, L. Dankin, 2301.10527, 2023.</p>
          <p>L. Edelstein, et al., An autonomous debating system, [28] J. Devlin, M. Chang, K. Lee, K. Toutanova, BERT:
pre-training of deep bidirectional transformers for
language understanding, in: NAACL-HLT, 2019,
pp. 4171–4186.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>