<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>These authors contributed equally.
$ giulia.vilone@tudublin.ie (G. Vilone); luca.longo@tudublin.ie (L. Longo)
 https://giuliavilone.github.io/ (G. Vilone); lucalongo.eu/about (L. Longo)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>networks⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giulia Vilone</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Longo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, Technological University Dublin</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Explaining the logic of a data-driven Machine Learning (ML) model can be seen as a defeasible reasoning process that is likely non-monotonic. This means a conclusion linked to a set of premises can be withdrawn when new information becomes available. Argumentation Theory (AT) formalises reasoning with a defeasible knowledge base. Abstract Argumentation Frameworks (AAF) organise conflicting arguments in a dialogical structure, allowing formal semantics to resolve conflicts. This study proposes an XAI method for automatically forming an AAF-based representation, using weighted attacks to model conflictual information. The concept of inconsistency budget is employed to eliminate the weakest attacks. Findings showed that the variation of the inconsistency budget could afect, albeit limited, the evaluation metrics computed over the resulting rulesets.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Explainable artificial intelligence</kwd>
        <kwd>Argumentation</kwd>
        <kwd>Non-monotonic reasoning</kwd>
        <kwd>Automatic attack extraction</kwd>
        <kwd>Weighted argumentation frameworks</kwd>
        <kwd>Inconsistency budget</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Numerous eXplainable AI (XAI) methods generate explanations of ML models in diferent
formats (numerical, rules, textual, visual or mixed) [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Rule-based explanations are considered
naturally transparent and intelligible because they are a structured, compact, and intuitive
format for reporting information [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Rule-based approaches usually consist of rulesets clarifying
the relationships between the inputs of a model and its outputs. However, these approaches
neither verify if rules are consistent with the background knowledge nor handle potential
inconsistencies among rules [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. Understanding the inferential process of a model can be
considered a non-monotonic reasoning process that allows the withdrawal of some conclusions,
carried out by some rules, in light of new information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Similarly, a model’s predictions
may be discarded when in conflict with the existing knowledge. This decision should follow a
process grounded on logic and evidence. Argumentation studies how conflicting arguments,
usually formalised with a first-order logical language, can be presented, supported or discarded
in a defeasible reasoning process and investigates formal approaches to evaluate the validity of
their conclusions [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. AT provides the basis for implementing these processes computationally
[
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], usually based on the notion of ‘arguments’ and ‘attacks’, often treated in an abstract
way [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Some scholars proposed methods for automatically mining arguments and attacks
from neural networks [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. However, more work needs to be done to assign weights to
arguments or attacks extracted from a trained model in an automatic way [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This study
focuses on automatically forming an argumentation framework consisting of rules and attacks
among them, and it uses the concept of inconsistency budget [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to determine a threshold
for the strength of such attacks, and it investigates the impact of its variation to the resulting
argumentation framework.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        AT can generate efective explanations by translating a model’s inferences in an argumentation
process that shows, step by step, how it concludes sets of conflicting arguments [
        <xref ref-type="bibr" rid="ref12 ref15">12, 15, 16, 17</xref>
        ].
Formal non-monotonic logic studies formal frameworks to capture and represent defeasible
inferences. A defeasible concept consists of a set of pieces of information, called arguments,
that can be invalidated by adding new information [18]. Defeasible argumentation provides
a sound formalisation for reasoning from a defeasible knowledge base [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. This process
frequently requires the recursive analysis of conflicting arguments in a dialectical setting to
determine which arguments should be accepted or discarded [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Abstract Argumentation
Theory (AAT) is the dominant paradigm, whereby arguments are abstractly considered in a
dialogical structure. The AAT-based frameworks share a defeasible knowledge base made of
arguments, a set of attacks to model conflicts between two arguments, and a formal semantic for
conflict resolution that implements non-monotonicity in practice and assigns a dialectical status
(accepted or rejected) to the arguments [16, 19]. [
        <xref ref-type="bibr" rid="ref14">20, 14</xref>
        ] assigned numeric weights to attacks,
thus introducing Weighted Argumentation Frameworks (WAF). Weights must be positive, real
values that assess the strength of the attack or, equivalently, measure the inconsistency between
two arguments. The notion of inconsistency budget indicates how much inconsistency must be
tolerated. Given an inconsistency budget  , all the attacks whose sum of weights is less than or
equal to  can be disregarded. However, there are no indications of determining the value of
the inconsistency budget to form a WAF with the optimal set of attacks and arguments.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Design</title>
      <p>The experiment conducted as part of this research consists of training and explaining a model
using a WAF. It contains six phases, as described below.</p>
      <p>Phase 1: Dataset preparation. The first step was to choose a set of five training datasets
containing multi-dimensional data handcrafted by domain experts and a categorical labelled
target variable. The experiment was conducted on the Adult, Avila, Bank, Credit Card Default
and Letter Recognition public datasets downloaded from the UCI Machine Learning Repository1.</p>
      <p>Phase 2: Model training. A feed-forward neural network with two fully-connected hidden
layers was trained on each dataset. The networks’ hyper-parameters (optimiser, activation
function, dropout rate, number of hidden neurons, and batch size) were tuned with a grid search
to reach the highest prediction accuracy; the training process was early-stopped to prevent
overfitting.</p>
      <p>
        Phase 3: Automatic formation of a knowledge base. The ML models and the datasets
were fed into a rule-extraction method, presented in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] that generates a set of  −  
rules using a two-step algorithm. Each rule corresponds to a ‘defeasible’ argument in the
resulting WAF. The weighted attacks were automatically extracted from the generated rules by
following the process proposed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Generally, attacks are binary relations between two
conflicting arguments and can be of diferent kinds [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This study considers only the following
two types: 1) rebutting, and 2) undercutting attacks [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Phase 4: Conflict evaluation. The weight of each attack measures the degree of
inconsistency between pairs of arguments. This inconsistency is the diference in the number of
instances supporting one of the two conflictual rules and belonging to their overlapping area [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
The concept of ‘inconsistency budget’ was used to determine how much inconsistency must be
tolerated [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In this study, the inconsistency budget varied between 10% and 90%.
      </p>
      <p>
        Phase 5: Dialectical status and accrual of arguments. Given conflicting arguments, their
acceptance status must be assigned. The ranking-base categoriser semantic [21, 22, 23] assigns
a rank value to each argument by considering the number of its attacks. If there are multiple
arguments with the highest rank, they are grouped into sets according to the conclusion they
support [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The most credible set is the one with the highest cardinality. In the case of ties, no
conclusion can be reached. This semantics does not consider the notion of weights of attacks.
      </p>
      <p>
        Phase 6: Explainability Objective evaluation. Eight metrics were chosen to objectively
and quantitatively measure the degree of explainability of the generated rule sets, per the
evaluation approach presented in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Objectivity is reached by excluding any human intervention in
the evaluation process. Number of rules and average rule length assess the syntactic simplicity of
the rules and must be minimised [24]. Fraction of output classes and fraction of overlap quantify
the rules’ clarity and coherence. The former should be as low as possible to avoid conflicts,
whereas the latter must be maximised to guarantee that all the target classes are considered.
A ruleset must also be complete, correct, faithful to the model’s predictions, and robust to be a
valid representation of a model’s inferential process [25, 26, 27, 28, 24].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results and conclusions</title>
      <p>The variation in the value of the inconsistency budget has, generally speaking, a limited impact
on the values of the metrics as the results remained almost unaltered throughout the five datasets
(see Fig. 1). Completeness and fraction of classes remained constant at 100%, whereas the values
of the other metrics for some datasets vary when the inconsistency budget goes above 50%.
The inconsistency budget is the cause of these variations in the metrics as they correspond to a
1https://archive.ics.uci.edu/ml/index.php
sharp increase in the number of eliminated attacks (see Fig. 2). This supports the assumption
that the number of attacks afects the robustness, correctness, fidelity and, sometimes, other
metrics calculated over the predictions made by a WAF, even if this efect is limited in some
instances. However, the limitations of this study in terms of the number and variety of models
and datasets prevent reaching definitive conclusions. Further studies with datasets containing
additional types of input data, such as texts and images, and ML models based on deeper neural
networks or other learning architectures would tell if the few variations in the metric values
are exceptions or if the inconsistency budget truly has an impact. Future work will extend this
research study by using other semantics considering the attacks’ weights.
[16] S. Modgil, F. Toni, F. Bex, I. Bratko, C. I. Chesnevar, W. Dvořák, M. A. Falappa, X. Fan, S. A.</p>
      <p>Gaggl, A. J. García, et al., The added value of argumentation, in: Agreement technologies,
Springer, 2013, pp. 357–403.
[17] A. Vassiliades, N. Bassiliades, T. Patkos, Argumentation and explainable artificial
intelligence: a survey, The Knowledge Engineering Review 36 (2021) e5.
[18] L. Longo, L. Rizzo, P. Dondio, Examining the modelling capabilities of defeasible
argumentation and non-monotonic fuzzy reasoning, Knowledge-Based Systems 211 (2021)
106514.
[19] S. A. Gómez, C. I. Chesnevar, Integrating defeasible argumentation with fuzzy art neural
networks for pattern classification, Journal of Computer Science &amp; Technology 4 (2004)
45–51.
[20] P. E. Dunne, A. Hunter, P. McBurney, S. Parsons, M. J. Wooldridge, Inconsistency tolerance
in weighted argument systems., in: AAMAS (2), 2009, pp. 851–858.
[21] L. Amgoud, J. Ben-Naim, Ranking-based semantics for argumentation frameworks, in:
International Conference on Scalable Uncertainty Management, Springer, 2013, pp. 134–
147.
[22] L. Amgoud, J. Ben-Naim, D. Doder, S. Vesic, Ranking arguments with
compensationbased semantics, in: Fifteenth International Conference on the Principles of Knowledge
Representation and Reasoning, 2016, pp. 12–212.
[23] P. Besnard, A. Hunter, A logic-based theory of deductive arguments, Artificial Intelligence
128 (2001) 203–235.
[24] H. Lakkaraju, S. H. Bach, J. Leskovec, Interpretable decision sets: A joint framework
for description and prediction, in: Proceedings of the 22nd ACM SIGKDD international
conference on knowledge discovery and data mining, ACM, San Francisco, California,
USA, 2016, pp. 1675–1684.
[25] G. Bologna, Y. Hayashi, A comparison study on rule extraction from neural network
ensembles, boosted shallow trees, and svms, Applied Computational Intelligence and Soft
Computing 2018 (2018). doi:10.1155/2018/4084850.
[26] C. Ferri, J. Hernández-Orallo, M. J. Ramírez-Quintana, From ensemble methods to
comprehensible models, in: International Conference on Discovery Science, Springer, Lübeck,
Germany, 2002, pp. 165–177. doi:10.1007/3-540-36182-0\_16.
[27] A. A. Freitas, Are we really discovering interesting knowledge from data, Expert Update
(the BCS-SGAI magazine) 9 (2006) 41–47.
[28] A. Ignatiev, Towards trustable explainable AI, in: Proceedings of the Twenty-Ninth
International Joint Conference on Artificial Intelligence, IJCAI, Yokohama, Japan, 2020, pp.
5154–5158. doi:10.24963/ijcai.2020/726.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruggieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Turini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giannotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pedreschi</surname>
          </string-name>
          ,
          <article-title>A survey of methods for explaining black box models, ACM computing surveys (CSUR) 51 (</article-title>
          <year>2018</year>
          )
          <volume>93</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>93</lpage>
          :
          <fpage>42</fpage>
          . doi:
          <volume>10</volume>
          .1145/3236009.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vilone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <article-title>Classification of explainable artificial intelligence methods through their output formats</article-title>
          ,
          <source>Machine Learning and Knowledge Extraction</source>
          <volume>3</volume>
          (
          <year>2021</year>
          )
          <fpage>615</fpage>
          -
          <lpage>661</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H. K.</given-names>
            <surname>Dam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ghose</surname>
          </string-name>
          ,
          <article-title>Explainable software analytics</article-title>
          ,
          <source>in: Proceedings of the 40th International Conference on Software Engineering: New Ideas and Emerging Results</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , Gothenburg, Sweden,
          <year>2018</year>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F. K.</given-names>
            <surname>Došilović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brčić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hlupić</surname>
          </string-name>
          ,
          <article-title>Explainable artificial intelligence: A survey</article-title>
          ,
          <source>in: 41st International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>0210</fpage>
          -
          <lpage>0215</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z. C.</given-names>
            <surname>Lipton</surname>
          </string-name>
          ,
          <article-title>The mythos of model interpretability</article-title>
          ,
          <source>Commun. ACM</source>
          <volume>61</volume>
          (
          <year>2018</year>
          )
          <fpage>36</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vilone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <article-title>A novel human-centred evaluation approach and an argument-based method for explainable artificial intelligence</article-title>
          ,
          <source>in: IFIP International Conference on Artificial Intelligence Applications and Innovations</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>447</fpage>
          -
          <lpage>460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bryant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Krause</surname>
          </string-name>
          ,
          <article-title>A review of current defeasible reasoning implementations</article-title>
          ,
          <source>The Knowledge Engineering Review</source>
          <volume>23</volume>
          (
          <year>2008</year>
          )
          <fpage>227</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <article-title>Argumentation for knowledge representation, conflict resolution, defeasible inference and its integration with machine learning</article-title>
          ,
          <source>in: Machine Learning for Health Informatics</source>
          , Springer,
          <year>2016</year>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>208</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L.</given-names>
            <surname>Rizzo</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Longo,</surname>
          </string-name>
          <article-title>An empirical evaluation of the inferential capacity of defeasible argumentation, non-monotonic fuzzy reasoning and expert systems</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>147</volume>
          (
          <year>2020</year>
          )
          <fpage>113220</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>P. M. Dung</surname>
          </string-name>
          ,
          <article-title>On the acceptability of arguments and its fundamental role in nonmonotonic reasoning, logic programming and n-person games</article-title>
          ,
          <source>Artificial intelligence 77</source>
          (
          <year>1995</year>
          )
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>O.</given-names>
            <surname>Cocarascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Cyras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Explanatory predictions with artificial neural networks and argumentation</article-title>
          ,
          <source>in: Proceedings of the 2nd Workshop on Explainable Artificial Intelligence (XAI</source>
          <year>2018</year>
          ),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>O.</given-names>
            <surname>Cocarascu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Toni</surname>
          </string-name>
          ,
          <article-title>Argumentation for machine learning: A survey.</article-title>
          ,
          <source>in: COMMA</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>230</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vilone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <article-title>A global model-agnostic xai method for the automatic formation of an abstract argumentation framework and its objective evaluation</article-title>
          ,
          <source>in: 1st International Workshop on Argumentation for eXplainable AI co-located with 9th International Conference on Computational Models of Argument (COMMA</source>
          <year>2022</year>
          ),
          <source>CEUR Workshop Proceedings</source>
          ,
          <year>2022</year>
          , p.
          <fpage>2119</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Dunne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>McBurney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Parsons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wooldridge</surname>
          </string-name>
          ,
          <article-title>Weighted argument systems: Basic definitions, algorithms, and complexity results</article-title>
          ,
          <source>Artificial Intelligence</source>
          <volume>175</volume>
          (
          <year>2011</year>
          )
          <fpage>457</fpage>
          -
          <lpage>486</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Gómez</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. I. Chesnevar</surname>
          </string-name>
          ,
          <article-title>Integrating defeasible argumentation and machine learning techniques: A preliminary report</article-title>
          , in: In Procs. V Workshop of Researchers in Comp. Science,
          <year>2003</year>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>