<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Discovering the Rationale of Decisions: Experiments on Aligning Learning and Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cor Steging</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silja Renooij</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bart Verheij</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bernoulli Institute of Mathematics</institution>
          ,
          <addr-line>Computer Science and Arti cial Intelligence</addr-line>
          ,
          <institution>University of Groningen</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information and Computing Sciences, Utrecht University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In AI and law, systems that are designed for decision support should be explainable when pursuing justice. In order for these systems to be fair and responsible, they should make correct decisions and make them using a sound and transparent rationale. In this paper, we introduce a knowledge-driven method for model-agnostic rationale evaluation using dedicated test cases, similar to unit-testing in professional software development. We apply this new method in a set of machine learning experiments aimed at extracting known knowledge structures from arti cial datasets from ctional and non- ctional legal settings. We show that our method allows us to analyze the rationale of black box machine learning systems by assessing which rationale elements are learned or not. Furthermore, we show that the rationale can be adjusted using tailor-made training data based on the results of the rationale evaluation.</p>
      </abstract>
      <kwd-group>
        <kwd>Responsible AI</kwd>
        <kwd>Explainable AI</kwd>
        <kwd>Machine Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In AI and Law, explainability is a key requirement in system design, due to
the need for the justi cation of decisions. For machine-supported decisions, this
is nowadays encoded in the GDPR's right to explanation. Four types of
explanations are distinguished [Miller, 2019] and have been applied in AI and
Law [At
        <xref ref-type="bibr" rid="ref3 ref4">kinson et al., 2020</xref>
        b]: Contrastive explanations show why a decision is
made and others are not. Examples include HYPO's counterexamples and
hypothetical situations [Rissland and Ashl
        <xref ref-type="bibr" rid="ref22">ey, 1987</xref>
        , Ashley, 1990] and argument
diagrams [Verheij, 2003a]. In selective explanations, the focus is on the most
salient elements needed, for instance by the use of the critical questions of
argumentation schemes [At
        <xref ref-type="bibr" rid="ref3 ref4">kinson et al., 2020</xref>
        a, Verheij, 2003b]. Probabilistic
explanations are grounded in statistical correlations, and are less applicable in
law with its focus on speci c circumstances. An example is the explanation of
evidential Bayesian networks [Vlek et al., 2016] in terms of scenarios and the
evidence for and against them. Lastly, social explanations emphasise the transfer
of knowledge between individuals, as in models of the dialogue between parties
and in courts, specifying shared and unshared commi
        <xref ref-type="bibr" rid="ref5">tments Hage et al. [1993</xref>
        ],
Gordon [1995], At
        <xref ref-type="bibr" rid="ref3 ref4">kinson et al. [2020</xref>
        a].
      </p>
      <p>
        This requirement of explainability is problematic for the application of
central machine learning techniques in law. Neural networks, for example, are known
to perform well, but behave like a black box algorithm. Hence, explanation
techniques have been developed to `open the black box'
        <xref ref-type="bibr" rid="ref14 ref21 ref27">(cf. LIME [Ribeiro et al.,
2016], SHAP [Lundberg and Lee, 2017])</xref>
        . Even in the domain of vision (where
the successes of neural networks are especially signi cant), the necessity of such
methods is underpinned by studies regarding adversarial attacks that show that
slight perturbations of images, invisible to the human observer, can radically
change the outcome of a class
        <xref ref-type="bibr" rid="ref7">i er [Goodfellow et al., 2015</xref>
        ].
      </p>
      <p>
        In this paper, we expand upon the method introdu
        <xref ref-type="bibr" rid="ref23">ced in [Steging et al.,
2021</xref>
        ], where we investigate black box machine learning methods with a focus
on proper explainability, and not only in terms of accuracy as in the standard
machine learning protocol. We are in particular interested in the discovery of
the rationale underlying decisions, where the rationale is the knowledge
structure that can justify a decision, such as the rule applied. We aim to measure
the quality of rationale discovery, with an eye on the possibility of improving
rationale discovery.
      </p>
      <p>To measure and possibly improve rationale discovery, we create dedicated
test datasets, on which a machine learning system can only perform well if it
has learned a particular component of the knowledge structure that de ned
the data. The idea is similar to how unit testing works in professional software
development: we de ne a set of cases, targeting a speci c component, in which
we know what the answer should be, and compare that to the output that the
system gives.</p>
      <p>
        To be able to focus on what is methodologically feasible, we do not use
natural language corpora
        <xref ref-type="bibr" rid="ref12 ref15 ref16 ref17 ref2 ref25 ref30 ref6 ref9">(as for instance in argument mining [Mochales Palau
and Moens, 2009, Wyner et al., 2010], conceptual retrieval [Grabmair et al.,
2015] or case prediction [Ashley, 2019, Medvedeva et al., 2019, Bruninghaus and
Ashley, 2003])</xref>
        . Instead we work with datasets of arti cial decisions with known
underlying generating rationale.
      </p>
      <p>
        Our work builds on a study investigating whether neural networks are able
to tackle open
        <xref ref-type="bibr" rid="ref5">texture problems [Bench-Capon, 1993</xref>
        ]. The study used a ctional
legal domain
        <xref ref-type="bibr" rid="ref18 ref29">(also investigated in [Wardeh et al., 2009, Mozina et al., 2005])</xref>
        , in
which the eligibility for a welfare bene t for elderly citizens is determined based
on six conditions. Arti cial datasets were generated specifying personal
information of elderly citizens with their eligibility for the welfare bene t. Multilayer
perceptrons were trained and tested on these datasets, and managed to perform
with high accuracy scores (above 98%). It was shown that the neural networks
were unable to properly learn two of the six conditions. By making adjustments
to the training dataset, the neural networks were able to learn conditions more
adequately, while maintaining similar accuracy scores. But also after adjustment,
the conditions that de ned the data were not learned fully correctly. Other
earlier discussions of neural networks in law are [Philipps an
        <xref ref-type="bibr" rid="ref11">d Sartor, 1999</xref>
        , Hunter,
1999, Stranieri et a
        <xref ref-type="bibr" rid="ref20">l., 1999</xref>
        ].
      </p>
      <p>
        In the upcoming sections,
        <xref ref-type="bibr" rid="ref5">the study by Bench-Capon [1993</xref>
        ] will rst be
replicated as closely as possible, using modern, widely-used neural network methods,
in order to a rm whe
        <xref ref-type="bibr" rid="ref5">ther the claims made in 1993</xref>
        still hold today. Then a
simpli ed version of the welfare bene t domain is examined to see how well the
networks are able to extract a simpli ed rationale. Lastly, we study a real legal
setting, namely Dutch tort law. That domain uses only Boolean variables, but
allows for exceptions to underlying rules.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Domains and Datasets</title>
      <p>For each of the three domains considered in this paper, this section describes
the underlying knowledge structure using logic, from which we will generate
datasets to train a series of neural networks. These networks will subsequently
be analysed using a method we propose for assessing the quality of their rational
discovery. To this end we need two types of datasets for the purpose of testing.
The rst are standard test sets sampled from the complete domain to evaluate
the accuracy of the networks. The second type is a dedicated test set designed
to target a speci c aspect of the domain knowledge. This section describes all
datasets we use.
2.1</p>
      <sec id="sec-2-1">
        <title>Domains</title>
        <p>
          Welfare bene t domain This ctional domain in
          <xref ref-type="bibr" rid="ref5">troduced in [Bench-Capon,
1993</xref>
          ] concerns the eligibility of a person for a welfare bene t to cover the
expenses for visiting their spouse in the hospital, and can be formalised as follows:
Eligible(x) () C1(x) ^ C2(x) ^ C3(x) ^ C4(x) ^ C5(x) ^ C6(x)
C1(x) () (Gender(x) = f emale ^ Age(x) 60)_
        </p>
        <p>(Gender(x) = male ^ Age(x) 65)
C2(x) () jCon1(x); Con2(x); Con3(x); Con4(x); Con5(x)j 4
C3(x) () Spouse(x)
C4(x) () :Absent(x)
C5(x) () :Resources(x) 3000
C6(x) () (T ype(x) = in ^ Distance(x) &lt; 50)_</p>
        <p>(T ype(x) = out ^ Distance(x) 50)
That is, a person is eligible i he/she is of pensionable age (60 for a woman,
65 for a man), paid four out of the last ve contributions Coni, is the patient's
spouse, is not absent from the UK, has capital resources not amounting to more
than £3,000, and lives at a distance of less than 50 miles from the hospital if the
relative is an in-patient, or beyond that for an out-patient.</p>
        <p>The six independent conditions for eligibility are de ned in terms of 12
variables, which are the features of the generated datasets. These features and their
possible values are shown in Table 1. In addition to these 12 features, the datasets
will contain 52 noise features unrelated to eligibility, just as in the original
experiment, giving a total of 64 features plus an eligibility label for each instance.
All datasets are valid in the sense that the given eligibility labels follow from
evaluating the 6 conditions above.</p>
        <p>
          Simpli ed domain Experiments with di erent models for the welfare domain
all concluded that it was not possible to extract all six conditions for eligibili
          <xref ref-type="bibr" rid="ref5">ty
[Bench-Capon, 1993</xref>
          , Wardeh et al., 2009, Mozina et al., 2005]. The complexity
of the original problem with 6 di erent conditions and 64 features, complicate
a proper analysis of the networks' rationale, since each condition and feature
could potentially in uence it. To facilitate this analysis, we simpli ed the original
problem in two ways. First, the 52 noise variables, which did not seem to a ect
the performance of
          <xref ref-type="bibr" rid="ref5">the networks [Bench-Capon, 1993</xref>
          ], are removed. Secondly,
we de ne eligibility solely by the age-gender (C1) and patient-distance (C6)
conditions that were examined in the original experiment to justify the rationale
of the network:
        </p>
        <p>Eligible(x) ()</p>
        <p>C1(x) ^ C6(x)
Eligibility is thus determined through a combination of a XOR-like function (C6)
and a nuanced threshold function (C1).</p>
        <p>Tort law domain Our third domain concerns Dutch tort law: articles 6:162
and 6:163 of the Dutch civil code that describe when a wrongful act is
committed and resulting damages must be repaired. This `duty to repair' (dut) can be
formalised as follows:
dut(x) () c1(x) ^ c2(x) ^ c3(x) ^ c4(x) ^ c5(x)
c1(x) () cau(x)
c2(x) () ico(x) _ ila(x) _ ift (x)
c3(x) () vun(x) _ (vst (x) ^ :jus(x)) _ (vrt (x) ^ :jus(x))
c4(x) () dmg (x)
c5(x) () :(vst (x) ^ :prp(x))
where the elementary propositions are provided alongside an argumentative
model of the law in Figure 1 [Verheij, 2017], and conditions c2 and c3 capture
the legal notions of unlawfulness (unl ) and imputability (imp), respectively.</p>
        <p>Compared to the ctional welfare domain, the Dutch tort law domain is
captured in 5 conditions for duty to repair (dut ), based upon 10 Boolean features.
Each condition is a disjunction of one or more features, possibly with exceptions.
The feature capturing a violation of a statutory duty (vst ) is present in both
condition c3 and c5, rendering these dependent.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Datasets</title>
        <p>
          For each experiment, we generate datasets of di erent types, for di erent
purposes3 For most types of datasets, the generating process is at least partly
stochastic and repeated for every repetition of an experiment. Using the same
3 The Jupyter notebooks used for data generation can be found in a Github repository:
https://github.com/CorSteging/DiscoveringTheRationaleOfDecisions
type of dataset, for example in training and testing a neural network, does
therefore not mean using the exact same dataset. Table 2 shows an overview of the
domains and their datasets, and illustrates their di erences and similarities.
Welfare bene t datasets Within this domain, four types of datasets are
generated, following the original s
          <xref ref-type="bibr" rid="ref5">tudy [Bench-Capon, 1993</xref>
          ] as closely as possible:
type A, type B, Age-Gender and Patient-Distance datasets. Each dataset
contains the 12 features as de ned in Table 1, as well as 52 noise features with
integer values ranging from 0 to 100. The original study used training sets with
2,400 instances, which is quite small by todays standards [At
          <xref ref-type="bibr" rid="ref3 ref4">kinson et al., 2020</xref>
          b].
To make sure conclusions are not the result of using too little data, we will also
include training sets with more data (50,000 instances).
        </p>
        <p>Type A datasets are generated with either 2,400 instances or 50,000 instances.
Exactly half of the instances are eligible, creating a balanced label distribution,
as is common practice in machine learning problems. For the eligible instances,
feature values are generated (randomly where possible) such that they satisfy
the conditions C1 C6. For each condition, 61 th of the ineligible instances is
designed to fail on that speci c condition; where possible the values of the features
involved are generated randomly such that the condition fails. All remaining
features in these instances are generated randomly across their full range of values
(see Table 1); as a result, it is possible for ineligible instances to fail on multiple
conditions, and some conditions will fail more often than others.</p>
        <p>In the original 1993 study it was argued that it was too easy to achieve high
accuracy scores with networks trained and tested on type A datasets, which
contained an average of 4.1 conditions that were not satis ed for ineligible cases.
Using only 4 out of 6 conditions was shown to be su cient for classifying 98.95%
of the instances correctly.</p>
        <p>Type B datasets were subsequently introduced to make the problem more
challenging. These datasets di er from type A datasets only in that ineligible
instances fail on exactly one condition, rather than at least one condition; the
other ve conditions are always satis ed. Type B datasets again contain either
2,400 instances or 50,000 instances.</p>
        <p>
          The original study investigated whether \an acceptable rationale can be
uncovered by an examina
          <xref ref-type="bibr" rid="ref5">tion of the net" [Bench-Capon, 1993</xref>
          ]. To this end a set
of test cases was constructed in which all conditions except one were guaranteed
to be satis ed. From these it was concluded that the age-gender condition (C1)
and the patient-distance condition (C6) are not learned by network. The original
paper does not specify how many test cases were constructed, nor exactly how
they were constructed. We will generate dedicated datasets that are tailor made
to evaluate whether the trained networks have actually uncovered these same
two conditions.
        </p>
        <p>The Age-Gender datasets are generated by sampling the age and gender
features across their full range of values, this time considering only multiples
of 5 for age. The values for the other features are generated such that every
condition is satis ed except for the age-gender condition (C1). As a result, the
eligibility of an instance in these datasets is solely determined by whether or not
condition C1 is satis ed. The Age-Gender sets contain 40,000 instances, with
every possible combination of values for age and gender occurring a 1000 times.
This gives a slightly unbalanced label distribution with 42.5% of the instances
being eligible, and 57.5% ineligible. Because the dataset is only used to test the
networks, rather than to train the network, this is not an issue.</p>
        <p>The Patient-Distance datasets are similarly generated by sampling the
distance and patient type features across their full range of values, this time
considering only multiples of 5 for distance. The eligibility of an instance in these
datasets is thus determined by whether or not condition C6 is satis ed. The
Patient-Distance datasets also contain 40,000 instances, with every possible
combination of values for patient type and distance occurring a 1000 times. In these
datasets, exactly 50% of the instances is eligible.</p>
        <p>Simpli ed datasets For the simpli ed welfare domain, the same type of datasets
are generated as above, with the same properties except that all noise features
and 8 of the 12 actual features are excluded. For type B datasets, this means
that ineligible instances fail on either C1 or C6, but not on both, while in type</p>
        <p>A datasets the ineligible instances can fail on both conditions. Moreover, in the
Age-Gender dataset the patient-distance condition C6 is always satis ed, and
in the Patient-Distance dataset the age-gender condition C1 is always satis ed.
Type A and type B datasets again contain 50,000 instances each. The
AgeGender dataset now contains only two features and 4,242 instances, that is, one
unique instance for every possible combination of age and gender. Likewise, the
Patient-Distance dataset contains 3,234 unique instances.</p>
        <p>Tort law datasets With 10 Boolean features there are 210 = 1024 possible
unique cases that can be generated from the argumentation structure of the
tort law domain in Figure 1. Each case has a corresponding outcome for dut,
indicating whether or not there is a duty to repair someone's damages. We will
again consider four types of datasets.</p>
        <p>The unique dataset contains these 1024 unique instances for the 10 features
plus the label. In this dataset, there are 912 instances where dut is false and 112
instances where dut is true (11%).</p>
        <p>The regular type datasets are generated such that dut is true in exactly half of
the instances. The sets are regular in the sense that balanced label distributions
are common in machine learning problems. These regular datasets are generated
by sampling uniformly from the subset of cases from the unique dataset, such
that each possible case is represented equally within the 50/50 label distribution.
In practice, only a subset of the possible cases is typically available and presented
to a network, upon which the network will have to learn to generalize to all
possible cases. In addition to generating regular type datasets with 5,000 cases,
we therefore also generate smaller regular type datasets with only 500 instances;
the latter contains 35.35% of the unique instances.</p>
        <p>In the tort law domain we focus on the notions of unlawfulness (c2) and
imputability (c3) to assess whether the networks are able to discover conditions
in the data. For each of the two conditions, we again create a dedicated dataset.</p>
        <p>The Unlawfulness dataset is the subset of the unique dataset in which the
features for the unlawfulness condition c2 can take on any of their values, while
the other features have values that are guaranteed to satisfy the remaining
conditions.Whether or not there is a duty to repair is therefore solely determined
by whether or not condition c2 is satis ed. All combinations of values of the
other features are considered. The Unlawfulness dataset therefore consists of
168 unique instances, of which 66.66% have a positive dut value.</p>
        <p>The Imputability dataset is a similar subset of the unique dataset, but now
the features for the imputability condition (c3) can take on any value, provided
that the value of vst is such that condition c5 is satis ed. The value of dut(x)
now completely depends on whether or not condition c3 evaluates to true. Due
to the interdependency of conditions c3 and c5, the Imputability dataset only
has 128 unique instances, with 87.5% of them having a positive dut value.</p>
        <p>Simpli ed
welfare
bene t
Tort law</p>
        <p>Regular (5,000 instances)
Regular (500 instances)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental setup and results</title>
      <p>
        In this section we describe and motivate the experiments we performed and
report on their results.
In order to demonstrate our method for assessing and improving rationale
discovery of models learned from data, we rst need such models. Though our
method is model agnostic, we choose to use neural ne
        <xref ref-type="bibr" rid="ref5">tworks, like in
[BenchCapon, 1993</xref>
        ]. We assume that assessing and improving rationale discovery is
relevant only for models that are considered to be a good match with the data
they were learned from. Our rst step, after training the above mentioned
neural networks, is therefore to evaluate their performance on typical test sets in
terms of the standard accuracy measure. Subsequently we will evaluate the
performance of the networks on the dedicated, knowledge-driven test sets that were
speci cally designed for assessing the networks' quality of rationale discovery.
Neural network architectures In
        <xref ref-type="bibr" rid="ref5">the original experiments in 1993</xref>
        , three
multilayer perceptrons were used with one, two and three hidden layers,
respec
        <xref ref-type="bibr" rid="ref5">tively [Bench-Capon, 1993</xref>
        ]. These networks were created using the Aspirin
softwa
        <xref ref-type="bibr" rid="ref13">re [Leighton and Wieland, 1994</xref>
        ], but the exact details regarding the
networks and its parameters (e.g. the learning rate, activation function, gradient
descent method) were left out of the original publication. The networks all had
64 input nodes (one for each feature in the datasets), a varying number of nodes
in the hidden layers, and one output node that determines the eligibility.
      </p>
      <p>
        In this paper we will use a similar set-up and network architecture for all
three domains. The output is always a single node, representing either eligibility
or duty to repair, depending on the domain under consideration. The number of
input nodes corresponds to the number of features and is therefore dependent on
the domain (see Table 2). More speci cally, the welfare bene t domain will have
64 input nodes, the simpli ed domain will have 4, and the tort law domain will
have 10 input nodes. The node con guration (i.e. number of nodes per layer) of
each network is as follows, where input represents the number of nodes in the
input layer:
{ One hidden layer network: input -12-1
{ Two hidden layer network: input -24-6-1
{ Three hidden layer network: input -24-10-3-1
In
        <xref ref-type="bibr" rid="ref5">the replication of the 1993</xref>
        experiment, the MLPClassi er of the scikit-learn
package is used [Pedregosa et al., 2011]. The networks use the sigmoid function as
their activation function, which was the most common activation function when
the original study was done. The networks use the Adam stochastic
gradientbased opt
        <xref ref-type="bibr" rid="ref7">imizer [Kingma and Ba, 2015</xref>
        ], with a constant learning rate of 0.001.
A total of 50,000 training iterations are used with a batch size of 50. Recall
that the focus of this study is not on creating the best possible classi er, but to
demonstrate our method of assessing rationale discovery.
      </p>
      <p>Training and performance testing The three types of neural networks will
be trained and tested on a combination of di erent datasets, from each of the
three domains. A complete overview of the datasets used in the experiments is
shown in Table 3. This table shows the datasets that the networks will train
on, and the datasets that the networks will be tested with. For each domain,
every combination of training dataset and testing dataset is evaluated in terms
of the accuracy of the resulting network on the test data. Because some of the
datasets are stochastic (each generated dataset is slightly di erent), the whole
process of data generation, training and testing is repeated 50 times. The mean
classi cation accuracies along with their standard deviations will be reported.</p>
      <p>To assess the rationale discovery capabilities of all the trained networks, we
study their performance on the dedicated test sets for the age-gender,
patientdistance, unlawfulness and imputability conditions. Performance will be
measured both quantitatively, using standard accuracy, and qualitatively by a more
detailed comparison of actual and expected outcomes.
3.2</p>
      <sec id="sec-3-1">
        <title>Results</title>
        <p>We will rst report the accuracy scores for all combinations of training and
testing datasets in the di erent domains and subsequently focus on rationale
discovery. Results will be discussed in detail in Section 4.
Accuracy Tables 5 { 8 show the mean classi cation accuracies over 50 runs,
together with their standard deviations, for the di erent combinations of training
and testing sets in the three domains. These tables include the quantitatively
measured performance on the various dedicated test sets. Tables 5 and 6 present
the results for the replication experiment with the original sizes of the Type
A and Type B datasets and for the replication experiment with more data,
respectively. The accuracies as reported in the orginal paper are provided in
Table 4 [Bench-Capon, 1993]). Results for the simpli ed welfare bene t domain
are shown in Table 7, and for the tort law domain in Table 8.
Rationale discovery Each dedicated test dataset is designed to measure how
well a model has learned a speci c condition from the domain. Since performance
on these test sets in terms of accuracy is comparable for the di erent neural
network architectures used, we present the results for the qualitative evaluation of
their rationale discovery capabilities for what is in theory the most sophisticated
one: the models with 3 hidden layers.</p>
        <p>In the welfare bene t domain, the Age-Gender datasets are used to measure
how well condition C1 is learned. In addition to measuring accuracy on these
dedicated datasets, we can plot the actual output of the neural network, which
should be 1 (eligible) for an individual of pensionable age and 0 otherwise, against
age, for both values of gender. Such plots, showing the mean output of the
networks over 50 runs, are shown in Figure 2 for both the welfare bene t domain
and its simpli ed version, for networks trained on each of the training sets under
consideration.</p>
        <p>Similarly, the Patient-Distance datasets are used to evaluate how well
condition C6 is learned, which can be assessed by plotting the networks' output
against the distance to hospital, for both in-patients and out-patients. The plots
showing the mean network output over 50 runs for both the welfare bene t
do(A) Replication
(B) More data
(C) Simpli ed
(D) Replication
(E) More data
(F) Simpli ed
Fig. 3: For all training sets from (simpli ed) welfare domain: mean network
output vs distance on Patient-Distance test set when trained on type A training
sets (A-C) and on type B training sets (D-F).
main and its simpli ed version, and for networks trained on each of the training
sets under consideration, are shown in Figure 3.</p>
        <p>In the tort law domain, we can similarly evaluate how well conditions c2
(unlawfullness) and c3 (imputability) are learned. For these conditions, the network
should output 1 in cases of the Unlawfulness dataset where the case is
unlawful (c2), or in the Imputability dataset where the case can be imputated to a
person (c3); otherwise the output should be 0. Since the tort law domain only
contains Boolean features, the outputs of the networks are presented in tables
rather than plots. The mean output over 50 runs for the two training sets on the
Unlawfulness and Imputability datasets is presented in Table 9.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>In this section we discuss in detail the results we found and the conclusions we can
draw from them. We separately focus on standard classi cation accuracy and on
rational discovery capabilities. We conclude by introducing the approach we took
as a general knowledge-driven method for model-agnostic rationale evaluation.
Standard accuracy is measured to see whether the learned models are able to
solve the classi cation problem, regardless of whether or not they discovered the
rationale underlying the data.</p>
      <p>Welfare bene t The accuracies obtained in the replication experiment (Table
5) di er from those in the original study (Table 4), but show similar trends.
Originally, networks trained on a type A training set performed well on type
A test sets (around 99%), but much worse on the type B test set (around
7076%). When trained on a type B training set, the accuracies on test set A in
the original study stayed the same, with accuracies on test set B increasing to
around 98%. In the replication experiment, training on a type B test set slightly
decreases accuracies on type A test sets, while the accuracy on type B test
sets increased less substantially than in the original experiment. In both cases,
changing the distribution of the training data from type A to type B served
to increase performance on test sets of the latter type while hardly a ecting
performance on test sets of the former type. Since type B datasets exploited
some knowledge of the domain (i.e. bene t is typically denied due to failure on
only a single condition), this suggests that overall performance can be improved
using tailor made training sets.</p>
      <p>Using more data, we nd higher accuracies (Table 6). This is not surprising,
as more training data generally leads to a better performance. Still, for networks
trained on type A data sets, accuracies on a type B test sets (below 85%) are
much lower than on type A test sets (around 99.8%). Training the networks
on type B training sets with more data shows signi cantly better results than
with fewer datapoints. Accuracies on type A test sets then are still around 99%,
whereas the accuracies on type B test sets are around 98%. The above
observations therefore still hold, even with more data and modern machine learning
methods.</p>
      <p>Simpli ed In the simpli ed domain, high accuracies are found across all datasets,
averaging out at around 99% (Table 7). Accuracies on type B test sets are only
slightly lower than on type A test sets, unlike in the other welfare domain
experiments. The networks do seem to perform slightly better on both types of
test sets when trained on a type B training set as compared to a type A training
set. However, this di erence is much more nuanced than in the regular welfare
bene t dataset. This can be explained by the fact that cases in type A data sets
can now fail on at most 2 conditions, rather than the original 6, which is only 1
more than the single failed condition in the type B datasets.</p>
      <p>Tort law In the Tort law domain we nd accuracies of 100% or near 100% for
networks trained on all instances (see Table 8). When presented with all unique
instances, the networks with one and two hidden layers are able to perfectly
predict the outcome from the Dutch tort law, and the network with three hidden
layers can create a very close approximation.</p>
      <p>Presenting a neural network with all available cases is in practice often
infeasible. If it is possible, then a simple lookup table rather than a neural network
would most likely su ce. For this reason, we also trained the networks on a
subset of only around 35% of the unique instances (see Table 8). As expected,
the accuracies of the networks on the general test sets drop, but only slightly
(to 98-99%). Even on the unique test set, accuracies remain around 96%. This
suggests that it is possible to approximate tort law with a small subset of the
unique cases.
4.2</p>
      <sec id="sec-4-1">
        <title>Rationale Discovery</title>
        <p>Looking at the performance of the networks on the dedicated test sets partially
exposes the rationale captured by the network. We designed these test sets such
that each one targets a single condition from the domain. In addition to
considering the accuracy on these dedicated test sets, we qualitatively evaluate the
rational discovery capabilities of the networks by comparing their outputs with
the actual outputs we would ideally expect for the di erent domains.
Welfare bene t In the welfare bene t domain, the Age-Gender dataset is
used to measure how well condition C1 is learned, that is, whether the networks
output 1 if the individual is of pensionable age (male and over the age of 65 or
female and over the age of 60), and output 0 otherwise. Plotting the age of the
individuals from the Age-Gender dataset against the output of the network, for
each gender, should ideally result in the graph on the left side of Figure 4. Here
the output of the network spikes instantly from 0 to 1 at the age of 60 for women
and 65 for men.</p>
        <p>Similarly, for cases from the Patient-Distance dataset the networks should
only output 1 (eligible) if the relative is an in-patient and the distance to the
hospital is less than 50 miles, or if the relative is an out-patient and the distance
to the hospital is further than 50 miles (condition C6). Plotting the distance
against the output of the network for both types of patients would ideally result
in the graph shown on the right in Figure 4.</p>
        <p>
          In our replication experiment the output graphs show a similar pa
          <xref ref-type="bibr" rid="ref5">ttern as in
the original 1993</xref>
          experiment. In the latter (not shown), the networks trained on
a type A training set do not show the expected pattern, for neither condition;
in fact for the Patient-Distance test cases always a 1 is returned. Training on
a type B training set improved the results, but the turning point at which the
networks output 1 is o . For the Age-Gender dataset it occurs at 45 for women,
rather than 60, and at 50 for men, instead of 65. For the Patient-Distance dataset
the turning point was too gradual, and takes place at 40, rather than 50 miles.
In our replication experiment the outputs of the networks trained on a type B
Fig. 4: An idealistic expectation of the outputs of a network on the Age-Gender
dataset versus the age for both genders (left) and on the Patient-Distance dataset
versus the distance for both patient types (right).
dataset (Figures 2(D) and 3(D)) more closely resemble the ideal outputs than the
outputs of the networks trained on a type A dataset (Figures 2(A) and 3(A)), for
both conditions. For the patient-distance condition (C6), the turning point does
occur at 50 after training on a type B dataset, unlike in the original experiment.
Training on a type B dataset seems to have a signi cant impact on the way the
rationale of the networks is formed, as networks are able to internalize condition
C1 and C6 better when trained on a type B dataset. This is furthermore re ected
in the accuracies on the Age-Gender and patient distance dataset as shown in
Table 5, which increase by roughly 30% when training on a type B dataset. These
accuracies on the Age-Gender and Patient-Distance datasets were not present
in the original study.
        </p>
        <p>Upon repeating the replication experiment with more data, this indeed
increases the performance of the networks signi cantly, but we still nd
performance for networks trained on type B datasets to be better than that of networks
trained on type A datasets. Output patterns more closely resemble the ideal ones
after training on type B datasets (see Figure 2(B) versus (E) and Figure 3(B)
versus (E)) and accuracies also increase (see Table 6). Interestingly, the turning
point for condition C1 does occur at the right place when training on more type
B training data: at 60 for females and 65 for males (see Figure 2(E)).
Simpli ed The simpli ed domain consists of only the two conditions C1 and
C6, without any other conditions or noise variables. In this less complex version
of the domain, overall performance is much higher, and conditions C1 and C6 are
learned quite successfully. Figures 2(C) and (F), and 3(C) and (F), respectively,
are very close to the ideal output graphs, with turning points in the correct
places. This is also re ected in near perfect accuracy scores on Age-Gender and
Patient-Distance datasets in Table 7. As argued before, the di erence between
type A and type B datasets is much smaller than in the original domain, hence
the results found for these to datasets are now quite similar.</p>
        <p>Tort law Recall that in the tort law domain, on the Imputability dataset,
networks should output 1 if the case can be imputated to the person, and 0
otherwise; on the Unlawfulness dataset, the networks should output 1 if the
case is unlawful, and 0 otherwise. Table 9 shows well the networks were able to
internalize the notions of unlawfulness and imputability. When trained on all
instances, the mean output of the networks is 0 if a case is not unlawful, and
1 if it is, which is exactly what it should do. Networks trained on all instances
attain a perfect score on the Imputability dataset as well. This can also be seen
in Table 8, where the networks score 100% accuracy on the Unlawfulness and
Imputability datasets after training on all instances.</p>
        <p>With less data, however, accuracies drop to around 92-95% for the
Unlawfulness dataset and 91-94% for the Imputability dataset. This accuracy may
still seem high, but we should take into account the label distributions
(66.6733.33% and 87.5-12.5%, respectively). Table 9 shows that networks still perform
perfectly on cases in which the unlawfulness and imputability conditions
evaluate to true. When the conditions are false, however, mistakes are made. The
average output of networks on the Unlawfulness dataset increases to 0.018, which
should be 0, meaning that it classi es some lawful cases as unlawful. In the
Imputability dataset, the mean output increased more drastically to 0.875 when
imputability is false. Meaning that in 87.5% of the instances in which the case
cannot be imputed to a person, the network incorrectly decided that it should.
This means that despite its high accuracy on the general test set, the networks
largely ignored the concept of imputability.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>A Method for Rationale Evaluation</title>
        <p>Although our experiments and discussion focused on speci c example domains
and neural networks, our approach for rationale evaluation can be seen as a
general method independent of the machine learning algorithm applied. This paper
therefore proposes a knowledge-driven method for model-agnostic rationale
evaluation, consisting of three distinct steps:
1. Measure the accuracy of a trained system, and proceed if the accuracy is
su ciently high;
2. Design dedicated test sets for rationale evaluation targeting selected
rationale elements based on expert knowledge of the domain;
3. Evaluate the rationale through the performance of the trained system on
these dedicated test sets.</p>
        <p>The rst step is based on the assumption that e orts for assessing and possibly
improving the rationale discovery capabilities of a learned model are only taken
if the general performance of the model is already considered good enough. Here
we assume performance is measured using accuracy, but other measures can be
employed as well and the threshold of what is considered good enough may vary
per domain and application.</p>
        <p>The second step in our method depends on domain knowledge. Hence the
method e ectively is a quantitative human-in-the-loop solution for rationale
evaluation.</p>
        <p>In the third step, performance is again evaluated, by now not only considering
accuracy but also examining model output and expected output in terms of the
dedicated test sets. Our examples have shown that the latter depend on the type
of features involved.</p>
        <p>Subsequently, the information gained by using this rationale evaluation method
can be used to improve the rationale of the system by adjusting the training data
accordingly, imposing sound rationale discovery.</p>
        <p>The method does not currently specify how the dedicated test sets are
constructed. We aim to further operationalize the rationale evaluation method by
using information about the knowledge in the domain, and the distribution of
examples, for instance building on Bayesian networks.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        The work in this paper was inspired by Bench-Capon's 1993 paper that
investigated whether neural networks are able to tackle open texture problems.
The conclusions were that neural networks can perform very well on such
problems in terms of accuracy, even if some conditions from the domain are no
        <xref ref-type="bibr" rid="ref5">t
learned [Bench-Capon, 1993</xref>
        ].
      </p>
      <p>
        In this paper we rst replicated the original experiments as closely as possible
to verify that we can reproduce
        <xref ref-type="bibr" rid="ref5">the results from the 1993</xref>
        paper. In addition,
we repeated the experiments with larger training datasets to ensure that the
original conclusions about conditions that were not learned are not due to a lack
of data. The idea of constructing test cases to test speci c conditions inspired us
to propose a method for assessing rationale discovery capabilities by designing
dedicated test datasets and to evaluate performance on these knowledge-driven
test sets, combining quantitative and qualitative evaluation elements in a hybrid
way. Type B datasets served to complicate the problem in the original study, but
also demonstrate that training can be improved using knowledge-driven tailor
made training sets.
      </p>
      <p>
        We investigated three legal domains, in which neural networks were trained
on labelled cases and tasked with predicting unlabelled cases. We started o
with an arti cial domain from the literature, followed by a simpli ed adaptation
of that domain. Lastly, we investigated a real life domain as well. The results
indicate that the network are able to achieve high accuracies in each of the
three domains. The networks are therefore able to make the right decisions in
most cases, with accuracies averaging around 99% on type A or regular test
sets. This is how machine learning problems are usually evaluated. Using our
approach of rationale evaluation, however, we show that the networks do not
necessarily learn the conditions, despite their high accuracy scores. Performance
on the dedicated test sets, type B, Age-Gender and Patient-Distance dataset
show that the networks are unable to learn the conditions C1 and C6. This was
suggested in the original experimen
        <xref ref-type="bibr" rid="ref5">t [Bench-Capon, 1993</xref>
        ] and it holds true in
the replication study with modern, commonly used machine learning techniques
and more data. By adjusting the distribution of the training data based on
expert domain knowledge (training on a type B dataset) these accuracies increase.
Simplifying the domain shows that systems are able to learn the conditions C1
and C6, though still not perfectly. Even in the real life tort law domain, with a
non- ctional knowledge structure and di erent characteristics, a similar pattern
can be observed. The networks failed to learn the independent condition that
de nes imputability, despite its high accuracies on the general test set.
      </p>
      <p>This study therefore rea rms the conclusions from previous work, while
simultaneously introducing a model-agnostic method for assessing rationale
discovery capabilities of machine learned black box models, using dedicated test
datasets designed with expert knowledge of the domain. In future research, we
aim to further detail and extend our method such that by employing it, the
soundness of the rationale becomes tangible, and its quality can be asserted.
Ultimately, based on this evaluation, the training data of the black-box systems
can be altered to improve their rationale. Further expanding upon this design
method will bring us closer to AI that is both explainable and responsible.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>K. D. Ashley</surname>
          </string-name>
          .
          <article-title>Modeling Legal Arguments: Reasoning with Cases and Hypotheticals</article-title>
          . The MIT Press, Cambridge (Massachusetts),
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>K. D. Ashley</surname>
          </string-name>
          .
          <article-title>A brief history of the changing roles of case prediction in ai and law</article-title>
          . Law in Context,
          <volume>36</volume>
          (
          <issue>1</issue>
          ):
          <volume>93</volume>
          {
          <fpage>112</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bench-Capon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bex</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Gordon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Prakken</surname>
          </string-name>
          , G. Sartor, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>In memoriam douglas n. walton: the in uence of doug walton on ai and law</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          , pages
          <volume>1</volume>
          {
          <fpage>46</fpage>
          ,
          <year>2020a</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>K.</given-names>
            <surname>Atkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bench-Capon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Bollegala</surname>
          </string-name>
          .
          <article-title>Explanation in ai and law: Past, present and future</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>289</volume>
          :
          <fpage>103387</fpage>
          ,
          <year>2020b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Bench-Capon</surname>
          </string-name>
          .
          <article-title>Neural networks and open texture</article-title>
          .
          <source>In Proceedings of the 4th International Conference on Arti cial Intelligence and Law</source>
          ,
          <source>ICAIL '93</source>
          , pages
          <fpage>292</fpage>
          {
          <fpage>297</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York,
          <year>1993</year>
          . ISBN 0-89791-606-9.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Bru</surname>
          </string-name>
          <article-title>ninghaus and</article-title>
          <string-name>
            <surname>K. D. Ashley</surname>
          </string-name>
          .
          <article-title>Predicting outcomes of case based legal arguments</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Arti cial Intelligence and Law (ICAIL</source>
          <year>2003</year>
          ), pages
          <fpage>233</fpage>
          {
          <fpage>242</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York (New York),
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          .
          <article-title>Explaining and harnessing adversarial examples</article-title>
          .
          <source>In Proceedings of International Conference on Learning Representations</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Gordon</surname>
          </string-name>
          .
          <source>The Pleadings Game: An Arti cial Intelligence Model of Procedural Justice</source>
          . Kluwer, Dordrecht,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Grabmair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Ashley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sureshkumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Nyberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Walker</surname>
          </string-name>
          .
          <article-title>Introducing LUIMA: an experiment in legal conceptual retrieval of vaccine injury decisions using a uima type system and tools</article-title>
          .
          <source>In Proceedings of the 15th International Conference on Arti cial Intelligence and Law</source>
          , pages
          <volume>69</volume>
          {
          <fpage>78</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York (New York),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Hage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Leenes</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Lodder</surname>
          </string-name>
          .
          <article-title>Hard cases: a procedural approach</article-title>
          .
          <source>Arti cial intelligence and law</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <volume>113</volume>
          {
          <fpage>167</fpage>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Hunter</surname>
          </string-name>
          .
          <article-title>Out of their minds: Legal theory in neural networks</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          ,
          <volume>7</volume>
          (
          <issue>2</issue>
          ):
          <volume>129</volume>
          {
          <fpage>151</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In Proceedings of 3rd International Conference on Learning Representations</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Leighton</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Wieland</surname>
          </string-name>
          .
          <source>The Aspirin/Migraines Software Package</source>
          , pages
          <volume>209</volume>
          {
          <fpage>227</fpage>
          . Springer, New York, Boston, MA,
          <year>1994</year>
          . ISBN 978-1-
          <fpage>4615</fpage>
          - 2736-7.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Lundberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>A uni ed approach to interpreting model predictions</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pages
          <fpage>4765</fpage>
          {
          <fpage>4774</fpage>
          . Curran Associates, Inc.,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Medvedeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vols</surname>
          </string-name>
          , and
          <string-name>
            <surname>M. Wieling.</surname>
          </string-name>
          <article-title>Using machine learning to predict decisions of the european court of human rights</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          , pages
          <volume>1</volume>
          {
          <fpage>30</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Explanation in arti cial intelligence: Insights from the social sciences</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>267</volume>
          :1{
          <fpage>38</fpage>
          ,
          <year>2019</year>
          . ISSN 0004-3702.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>R. Mochales</given-names>
            <surname>Palau</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Moens</surname>
          </string-name>
          .
          <article-title>Argumentation mining: the detection, classi cation and structure of arguments in text</article-title>
          .
          <source>In Proceedings of the 12th International Conference on Arti cial Intelligence and Law (ICAIL</source>
          <year>2009</year>
          ), pages
          <fpage>98</fpage>
          {
          <fpage>107</fpage>
          . ACM Press, New York (New York),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Mozina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zabkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bench-Capon</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Bratko. Argument</surname>
          </string-name>
          <article-title>based machine learning applied to law</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          ,
          <volume>13</volume>
          (
          <issue>1</issue>
          ):
          <volume>53</volume>
          {
          <fpage>73</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <article-title>Scikit-learn: machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          :
          <fpage>2825</fpage>
          {
          <fpage>2830</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Philipps</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Sartor.</surname>
          </string-name>
          <article-title>Introduction: from legal theories to neural networks and fuzzy reasoning</article-title>
          .
          <source>Arti cial Intelligence and law</source>
          ,
          <volume>7</volume>
          (
          <issue>2</issue>
          ):
          <volume>115</volume>
          {
          <fpage>128</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>M. T. Ribeiro</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Singh</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Guestrin</surname>
          </string-name>
          .
          <article-title>"why should I trust you?": Explaining the predictions of any classi er</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , San Francisco, CA, USA, pages
          <volume>1135</volume>
          {
          <fpage>1144</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>E. L.</given-names>
            <surname>Rissland</surname>
          </string-name>
          and
          <string-name>
            <given-names>K. D.</given-names>
            <surname>Ashley</surname>
          </string-name>
          .
          <article-title>A case-based system for trade secrets law</article-title>
          .
          <source>In Proceedings of the 1st International Conference on Arti cial Intelligence and Law</source>
          ,
          <source>ICAIL '87</source>
          , pages
          <fpage>60</fpage>
          {
          <fpage>66</fpage>
          , New York, NY, USA,
          <year>1987</year>
          . ACM. ISBN 0897912306.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Steging</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Renooij</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>Discovering the rationale of decisions: Towards a method for aligning learning and reasoning (accepted)</article-title>
          .
          <source>In Proceedings of the 18th International Conference on Arti cial Intelligence and Law</source>
          , ICAIL '
          <fpage>21</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Stranieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zeleznikow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gawler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Lewis</surname>
          </string-name>
          .
          <article-title>A hybrid rule{neural approach for the automation of legal reasoning in the discretionary domain of family law in australia</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          ,
          <volume>7</volume>
          (
          <issue>2</issue>
          -3):
          <volume>153</volume>
          {
          <fpage>183</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>Arti cial argument assistants for defeasible argumentation</article-title>
          .
          <source>Arti cial Intelligence</source>
          ,
          <volume>150</volume>
          (
          <issue>1</issue>
          {2):
          <volume>291</volume>
          {
          <fpage>324</fpage>
          ,
          <year>2003a</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>Dialectical argumentation with argumentation schemes: An approach to legal logic</article-title>
          .
          <source>Arti cial intelligence and Law</source>
          ,
          <volume>11</volume>
          (
          <issue>2-3</issue>
          ):
          <volume>167</volume>
          {
          <fpage>195</fpage>
          ,
          <year>2003b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>Formalizing arguments, rules and cases</article-title>
          .
          <source>In Proceedings of the 16th International Conference on Arti cial Intelligence and Law</source>
          ,
          <source>ICAIL '17</source>
          , pages
          <fpage>199</fpage>
          {
          <fpage>208</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York,
          <year>2017</year>
          . ISBN 978-1-
          <fpage>4503</fpage>
          -4891-1.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Vlek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Prakken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Renooij</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Verheij</surname>
          </string-name>
          .
          <article-title>A method for explaining bayesian networks for legal evidence with scenarios</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          ,
          <volume>24</volume>
          (
          <issue>3</issue>
          ):
          <volume>285</volume>
          {
          <fpage>324</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Wardeh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Bench-Capon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Coenen</surname>
          </string-name>
          .
          <article-title>Padua: a protocol for argumentation dialogue using association rules</article-title>
          .
          <source>Arti cial Intelligence and Law</source>
          ,
          <volume>17</volume>
          (
          <issue>3</issue>
          ):
          <volume>183</volume>
          {
          <fpage>215</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Wyner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mochales-Palau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Moens</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Milward</surname>
          </string-name>
          .
          <article-title>Approaches to text mining arguments from legal cases</article-title>
          .
          <source>In Semantic Processing of Legal Texts</source>
          , pages
          <volume>60</volume>
          {
          <fpage>79</fpage>
          . Springer, Berlin,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>