<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Machine Coaching with Proxy Coaches</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vassilis Markos</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marios Thoma</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Loizos Michael</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CYENS Centre of Excellence</institution>
          ,
          <addr-line>Nicosia</addr-line>
          ,
          <country country="CY">Cyprus</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Open University of Cyprus</institution>
          ,
          <addr-line>Nicosia</addr-line>
          ,
          <country country="CY">Cyprus</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We evaluate the Machine Coaching paradigm, a human-in-the-loop machine learning methodology, according to which a human coach and a machine engage in an iterative bidirectional exchange of explanations, towards improving the machine's ability to reach conclusions and justify them in a way that is acceptable to the human coach. To support the systematic empirical investigation of the eficacy and eficiency of Machine Coaching, we adopt proxy (algorithmic) coaches in the stead of human ones.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;machine coaching</kwd>
        <kwd>proxy coaching</kwd>
        <kwd>explainable AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Learning in humans proceeds in many and diverse ways. Under what could be called the
autodidactic paradigm [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ], the human learner utilizes whatever information is available. A
human supervisor, if present, may complete / enrich that information, so that the human learner
faces a more benign / informative environment for learning.
      </p>
      <p>
        Another, significant, part of human learning takes place under what could be called the
coaching-based paradigm [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], where a human supervisor, or coach, shares with the human
learner not only what the case is in a certain state of the environment, but chiefly why that
case is. This happens whenever parents warn their children to “not run while holding scissors,
because they will hurt themselves”, whenever teachers explain to their students to “conclude
that two triangles are similar because they have congruent angles”, when chess instructors
teach their pupils to “place pieces on squares from which they cannot be easily deflected”, and
when managers direct their assistants to “book a hotel close to the meeting venue for same-day
trips”.
      </p>
      <p>
        Such pieces of advice from the coach do not typically come unprompted, but as a reaction
to a wrong or wrongly-justified decision by the learner. A coach ofering advice is, in efect,
proactively completing / enriching missing information in states of the environment that the
learner might encounter, by providing conditions on states under which a certain decision should
be reached. Compared to the autodidactic paradigm, the substantial amount of information
communicated under the coaching-based paradigm is conducive to more eficient, albeit
coachspecific, learning, while the reactive nature of advising entails only marginally extra efort
from the coach, as humans are known to be competent at identifying counter-arguments when
challenged [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Machine Coaching1 [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] was proposed as an argumentation-based learning-theoretic
framework under which the coaching-based paradigm of learning can be studied. It formalizes the
computational resources available to the coach and the learner, the protocol and language of
their interaction, and the expected quality guarantees on what is learned. Under that framework,
one considers a human coach, with access to a target policy, interacting with a machine learner,
and formally establishes that if the human coach ofers appropriate, in a defined-sense, advice,
then the machine learner can utilize that advice to learn a hypothesis policy that approximates
the target policy with the expected quality guarantees.
      </p>
      <p>Undertaking a study with human participants in the role of the coach is, clearly, the ultimate
way to empirically validate the Machine Coaching framework. Yet there are three major
challenges to overcome in such an envisioned study.</p>
      <p>
        The first challenge stems from the lack of fluency of average humans in formal logic , making
it hard for them to exchange explanations in the native (logic-based) language of Machine
Coaching. This issue could be alleviated by adopting a natural language interface for exchanging
explanations, which would be automatically translated to/from the native language. Although
there is some work in this direction — even using the Machine Coaching framework itself at a
meta-level [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] — the problem is suficiently complex, and currently lacks a robust solution.
      </p>
      <p>The second challenge stems from the inaccessibility of the target policy that humans have in
mind when coaching. Without a target policy, one cannot empirically evaluate the quality of the
learned hypothesis policy. If, in addition, no policy agrees with all the advice ofered by the
human coach, then the empirical study would confound the validation of the Machine Coaching
framework with the determination of the expressivity of the framework’s native language, and
with the compatibility of the interaction with human cognition.</p>
      <p>The third challenge stems from the physical and mental limitations of humans in prolonged
interactions. A systematic evaluation of the Machine Coaching framework would require
substantial rounds of interactions between the coach and the machine, across diverse target
policies. It is unlikely that humans would maintain eficient and consistent behavior over time,
either due to fatigue, an unconscious adaptation to the empirical setting, or a suppression of
their natural reactions when coaching with an externally-given target policy.</p>
      <p>
        To eschew these challenges, we resort to the use of a proxy (algorithmic) coach that can
communicate in the native language of Machine Coaching, cope with a given explicit target
policy, and provide advice eficiently and consistently. These features need to be balanced
against the desire to stay close to human coaches, who have only approximate access to their
target policy — conceivably as a compilation, in some intensional or extensional form, of their
relevant life experiences — and are unable to divulge the target policy on cue, but can still react
if some part of it is “challenged” by an external position and associated argument [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Though
the proxy coaches introduced below do not claim to serve as a comprehensive simulation of
human ones, they aim to capture the interaction aspects that are important for a sound in vitro
assessment of Machine Coaching. Consequently, we consider our work as a stepping stone
between the theoretically proven eficiency of Machine Coaching and an empirical assessment
1Resources available at: https://cognition.ouc.ac.cy/prudens.
of that through a study involving actual human participants.
      </p>
      <p>Accordingly, we consider proxy coaches with access to an exemplar set of data, labeled with
the predictions of a fixed (but hidden to the proxy coaches) target policy. The key technical
question, then, is how to extract appropriate pieces of advice from the exemplar set. In the sequel
we seek to answer this question by demonstrating how to develop proxy coaches with both
an intensional and an extensional compilation of the exemplar data, and continue to evaluate
them empirically. We shall remark at this point that our focus in this paper is mostly shifted
towards investigating the various alternatives of proxy coaches and their qualitative diferences.
Consequently, while we do present and discuss quantitative results for all chosen proxy coaches,
they mostly serve as a means to stress the efects each proxy coaching protocol has on the
coaching process.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <p>
        The process whereby an agent learns under the supervision of a more experienced tutor / coach,
who ofers advice as a reaction to the learner’s decisions or actions, appears as a suggestion
for the development of AI systems in John McCarthy’s seminal work on the “Advice Taker”
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Recent work [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] has sought to formalize this process, termed Machine Coaching, in
learning-theoretic terms, as a variant of the Probably Approximately Correct (PAC) model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        The eXplainable Interactive Learning (XIL) framework [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] also considers a human supervisor
in the role of a learning “coach”, providing advice either in the form of a (more) correct label,
or a correct explanation. Unlike Machine Coaching, which adopts learning and reasoning
semantics compatible with the language of formal argumentation, XIL assumes black-box access
to an active learning algorithm [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and a local post-hoc explainer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Correspondingly, the
explanations provided by the “coach” in XIL cannot be directly and elaboration-tolerantly [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
embedded into the learned model, but are rather used to create additional labeled data for the
further training of an opaque learned model.
      </p>
      <p>
        The idea of online ingestion of human advice in Machine Learning processes is present in
other works as well. A way to incorporate user advice into Support Vector Machines is proposed
in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], using the advice to reduce the number of data points that need to be labeled. In [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], an
interactive framework is proposed that allows human users to label selected data points for the
purpose of training a document classifier. Similarly, in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a learning paradigm is proposed
that systematically utilizes user-provided explanations to reduce learning complexity, stressing
that data richness (i.e., including explanations along with labels) over data volume (i.e., having
access to many labeled data without explanations) leads, in certain cases, to equally good or
better performance.
      </p>
      <p>
        Beyond speeding up learning, more recent works examine how human knowledge may
benefit the entire learning process, by being an integral part of it. Coactive Learning, proposed
in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], is an interactive Machine Learning paradigm where a learner receives sub-optimal user
advice that is, inherently, only “slightly better” than the learner’s current prediction. However,
it is shown that several existing supervised algorithms can be altered to accommodate this type
of human-machine interaction.
      </p>
      <p>Although our work herein focuses on the empirical evaluation of Machine Coaching, the
technique of proxy coaches that we follow would also seem to be applicable, to varying extents,
to the frameworks above. The rest of the works that we briefly review below are not tied to
coaching per se, but relate to our chosen implementation of the proxy coaches.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], the authors attempt to ofer post-hoc explanations for a black-box model by training
a random forest, extracting rules from it, and using them to construct an argumentation theory.
One of our considered proxy coach types also adopts the view that a random forest can be the
source of arguments, but in our case these arguments are not the end product itself, but are
rather the pieces of advice that are given from the proxy coach to the ultimate learner.
      </p>
      <p>
        On the other hand, we also consider a proxy coach type that does not compile the available
data into an explicit learned model, but uses them only implicitly. This approach relates to the
works in [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ], which investigate a paradigm of implicit learning, where one can reason and
respond to queries by directly consulting the data. Whereas the emphasis of those works is on
answering given queries, our emphasis is on constructing appropriate pieces of advice for the
proxy coach to ofer to the ultimate learner.
      </p>
      <p>
        To facilitate the choice of appropriate advice in this implicit learning setting, we use natural
selection over an evolutionary process. The formalization of evolution that we adopt is that in
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], where evolution is cast as a learning problem that seeks to approximate a hidden target
function. The evolutionary process in our case attempts to identify the next appropriate piece of
advice, and to integrate it into the next generation of organisms, with each organism encoding a
collection of rules aimed to approximate the entire target policy. Thus, our approach resembles
a Pittsburgh-style Learning Classifier System [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. From Direct to Proxy Coaches</title>
      <p>Following the Machine Coaching framework, a (direct) coach and a learner hold, respectively,
a fixed target policy and a revisable hypothesis policy, each represented in the form of a
partially-ordered set of “if-then” rules. Upon perceiving a new context, the learner returns the
predictions of the current hypothesis policy, and an explanation in the form of hypothesis
policy rules that support those predictions. The coach responds by ofering advice in the form
of target policy rules that support why the learner’s predictions were incorrect, incomplete, or
improperly-justified according to the coach. The learner revises the hypothesis policy, and the
process repeats.</p>
      <p>While arguments within Machine Coaching are generally considered to be arbitrary trees
of rules, for the sequel, we restrict our attention to shallow propositional policies, whose
rules have a special atom (or its negation) as their head, so that rules are not chained during
reasoning — and, thus, each argument comprises of a single rule. Relatedly, we restrict our
attention to complete contexts, which specify truth-values for all remaining atoms. Without
loss of generality, we hence let the special atom be output.</p>
      <p>Thus, if output were the ability to fly, and given a context {penguin, bird}, the learner
could predict output by ofering the following explanation: “ bird implies output”, and
the coach could react by ofering the following piece of advice: “ penguin implies -output”,
which would be integrated with higher priority in the revised hypothesis policy.</p>
      <p>Even with these restrictions, policies can still vary along other dimensions, which afect the
Stack
Stack</p>
      <sec id="sec-3-1">
        <title>R1 :: a implies output</title>
        <sec id="sec-3-1-1">
          <title>R2 :: a, b implies -output</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>R3 :: a, c implies -output</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>R4 :: x implies output</title>
        <sec id="sec-3-2-1">
          <title>R5 :: x, y implies -output</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>R6 :: x, y, z implies output</title>
          <p>policy size = 6 rules
rule width = 11 atoms / 6 rules
flip count = 4 batch transitions
batch size = 6 rules / 5 batches
stack depth = 5 batches / 2 stacks
conceptual and computational complexity of learning (through Machine Coaching, or otherwise)
and reasoning. Policies can vary in terms of their policy size, which counts the number of rules
in the policy, their rule width, which counts the average number of atoms in rule bodies, their
lfip count , which counts the flips / transitions between batches of consecutive rules with the
same head polarity (positive or negative), the batch size, which counts the average number of
rules across batches, and the stack depth, which counts the average number of consecutive
batches such that rules in a batch have logically more specific conditions than the rules in the
preceding batch (cf. Figure 1).</p>
          <p>Proxy coaches have only indirect knowledge of some target policy , by being given access
to an exemplar set  of contexts labeled according to . In broad terms, then, a proxy coach
̂︀
faces a context  ∈  from some coaching set , receives the prediction / explanation ℎ() of
the learner’s current hypothesis policy ℎ on , and then uses ̂︀ to generate a piece of advice
 (, ℎ(); ̂︀) to be ofered to the learner. An epoch concludes whenever a piece of advice is
ofered, noting that the proxy coach need not ofer advice for each  ∈ .</p>
          <p>The precise manner on how  (, ℎ(); ) is generated depends on the proxy coach type, but
̂︀
it generally seeks to consider the anticipated efect that each candidate piece of advice would
have on the learner’s hypothesis policy, in terms of: (i) improving its coverage, by reducing the
number of contexts on which it abstains; and (ii) improving its accuracy, by reducing the ratio
of wrong over correct predictions it makes.
3.1. Intensional Proxy Coaches
The first class of proxy coaches that we consider encode intensionally their knowledge of the
target policy, by using the exemplar set ̂︀ once, before the coaching phase, to supervise the
training of a white-box model . The ante-hoc explainability of  makes it, then, natural to
seek to “extract” a piece of advice  (, ℎ(); ) from the explanation of (), given a context
̂︀
 ∈ . The exemplar set, therefore, takes the role of a training set. Coaching proceeds, then, as in
Algorithm 1.</p>
          <p>Our chosen white-box models are those of decision trees and random forests, from which we
have identified three natural ways of extracting advice as part of Step 6 of Algorithm 1.</p>
          <p>Our first approach considers a white-box model  based on a single decision tree. Given the
current context  ∈ , it identifies the active path of the tree, from which it computes, as usual,
Algorithm 1 Intensional Proxy Coaching
1: input: exemplar set ; coaching set .</p>
          <p>̂︀
2: Train white-box model , using  as training instances.</p>
          <p>̂︀
3: for the next context  ∈  do
4: Ask for learner’s prediction / explanation ℎ().
5: if predictions of () and ℎ() difer then
6: Get advice from explanations of () and ℎ().
7: Give advice to revise current hypothesis policy ℎ.
8: end if
9: end for
· · ·
the prediction of (). The full path itself acts as the explanation of (), and is returned
as a piece of advice. By construction, the advice so generated is “cautious”, in that a full path
in decision trees tends to have high accuracy on the exemplar set, as otherwise the decision
tree learning algorithm would have chosen to expand the path further. For the same reason,
however, the advice is also “specific”, in that it applies only on a few contexts because of the
path length. Overall, then, this type of a proxy coach chooses to provide advice based on the
expectation that the hypothesis policy will improve its coverage only marginally, but it will be
highly efective in improving and maintaining its accuracy.</p>
          <p>Our second approach is a variant of the first, where instead of considering the full active
path of the decision tree as the (unique candidate for a) piece of advice, it considers all of
its prefixes as candidates. Among them, it chooses the minimal one that is simultaneously
more specific than the explanation of ℎ(), and such that the prediction of the decision tree
would have remained the same had the active path been pruned to be that particular prefix.
By construction, the advice so generated is more “loose”, in that a prefix of a path in decision
Algorithm 2 Extensional Proxy Coaching
1: input: exemplar set ; coaching set .</p>
          <p>̂︀
2: Initialize the current hypothesis policy ℎ to be empty.
3: for the next context  ∈  do
4: Create mutations of types (M0), (M+), (M↓).
5: Compute fitness score  of each mutation ℎ .
6: Select “good” beneficial or neutral mutation ℎ.
7: Set the current hypothesis policy to be ℎ.</p>
          <p>8: end for
trees tends to have lower accuracy on the exemplar set, which is why the decision tree learning
algorithm had chosen to expand the prefix further. For the same reason, however, the advice is
also more “general”, in that it applies on more contexts because of the shorter prefix length.
Overall, then, this type of a proxy coach chooses to provide advice based on the expectation
that the hypothesis policy will improve its coverage measurably, but at the cost of being less
efective in improving and maintaining its accuracy.</p>
          <p>Our third approach considers a white-box model  based on a random forest. Given the
current context  ∈ , it computes, as usual, the prediction of (). It then proceeds to identify
the active path from each individual tree whose prediction matches that of (), and considers
those full paths as candidates. There are many strategies for choosing one of the candidates, but
for concreteness, our particular approach chooses the one that is most “specific” (breaking ties
randomly), in that it is activated by as few contexts as possible in the exemplar set. Given that
paths in trees in random forests are typically short and hence “general”, the approach tends
to favor the generation of a “maximally cautious” advice among “typically loose” candidates.
Overall, then, this type of a proxy coach chooses to provide advice based on the expectation of
striking a balance between the improvement of the coverage and accuracy of the hypothesis
policy.
3.2. Extensional Proxy Coaches
The second class of proxy coaches that we consider retain extensionally their knowledge of
the target policy, without compiling it into another form. Rather, they use the exemplar set ̂︀
repeatedly, during the coaching phase, to decide whether a candidate piece of advice should
become  (, ℎ(); ), given a context  ∈ . The exemplar set, therefore, takes the role of a
̂︀
validation set. Coaching proceeds, then, as in Algorithm 2.</p>
          <p>This extensional perspective suggests a trial-and-error view of coaching, where the coach
tests candidate pieces of advice to measure their efect. Since a systematic testing of all possible
pieces of advice is infeasible, the challenge for the coach is to select the set of candidates, and to
identify how to test the efect of each candidate on the hypothesis policy.</p>
          <p>Our approach adopts an evolutionary mechanism, where, roughly, the process of mutation
corresponds to the generation of the candidate pieces of advice, and the process of natural
selection corresponds to their testing towards selecting  (, ℎ(); ). We discuss below in more
̂︀
detail the nuances that come from the fact that the same evolutionary mechanism needs to
simulate both the proxy coach and the learner.</p>
          <p>At the start of a generation, the population comprises the current hypothesis policy ℎ. Given
the current context  ∈ , certain candidate pieces of advice are considered, and each is
“provisionally” ofered as advice to a diferent copy of ℎ, giving rise to its ofspring. Specifically:
(M0) an ofspring is created by ofering no advice; (M+) an ofspring is created by ofering
the advice “ implies output”, if ℎ() predicts -output or abstains; (M+) an ofspring is
created by ofering the advice “  implies -output”, if ℎ() predicts output or abstains;
(M↓) an ofspring is created, for each , by ofering the advice “ body−  implies head”, where
“body implies head” is the latest advice that was integrated in ℎ, and body−  is body minus
its -th literal.</p>
          <p>These aforementioned pieces of advice are said to be ofered “provisionally” in the sense that
the proxy coach may, at a subsequent generation, seek to refine a previously given piece of
advice by means of type (M↓) mutations. A piece of advice is “conclusively” given at the end of
a streak of (zero or more) type (M0) or (M↓) mutations that follow a type (M+) mutation. This is
the point at which an epoch concludes.</p>
          <p>At each generation, exactly one of the ofspring is chosen to survive to initialize the next
generation. To support this selection process, each ofspring ℎ is evaluated in terms of its
improvement in accuracy and coverage relative to its parent against each exemplar  ∈ ̂︀. The
change from ℎ() to ℎ() is considered: positive, if a wrong prediction changes to a correct
one or an abstention, or an abstention changes to a correct prediction; negative, if a correct
prediction changes to a wrong one or an abstention, or an abstention changes to a wrong
prediction. By giving a +1 or − 1 fitness point to ofspring ℎ for each respective positive or
negative change across the exemplar set , we end up with an intuitive metric  that aggregates
̂︀
the improvement efect of the advice given to ofspring ℎ on both its coverage and its accuracy.</p>
          <p>The ofspring are, then, grouped based on the relation of their relative fitness to a fixed
threshold parameter . An ofspring ℎ is detrimental, neutral, beneficial if its relative
iftness  belongs in (−∞ , − ), in [− , +], in (+, +∞), respectively. Among the beneficial
ones, if available, the ofspring ℎ is selected to survive with probability / ∑︀ , where the
exponent  is a fixed non-linearity parameter. Otherwise, a neutral ofspring (whose existence
is guaranteed by a type (M0) mutation) is selected uniformly at random.</p>
          <p>The proposed approach searches greedily for advice that is as “cautious” and as “general”
as possible, while starting the greedy search from a fully “cautious” and “specific” seed advice
selected by a type (M+) mutation. Overall, then, this type of a proxy coach chooses to provide
advice based on the expectation of striking a balance between the improvement of the coverage
and of the accuracy of the hypothesis policy.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Empirical Investigation</title>
      <p>We present below the key aspects and results of our empirical investigation. Additional details
are given in the appendices. All related materials may be found at https://github.com/VMarkos/
proxy-coaching-argml-2022.
4.1. Empirical Setting
We first fixed a set  of 20 atoms (other than output) for use in contexts and rule bodies.
We then constructed 10 groups with 20 policies each, with each pair of groups corresponding
to high versus low values in one of the variability dimensions discussed in Section 3, giving
rise to a set  of 132 distinct target policies. For each target policy  ∈  , we constructed an
associated exemplar set ̂︀ by uniformly at random sampling 1000 contexts from 2, labeling
each context according to , filtering out any contexts on which  abstained, and keeping 70%
of the labeled contexts. The other 30% was used to populate an evaluation set ℰ . An additional
500 (unlabeled) contexts were sampled to populate a coaching set .</p>
      <p>
        We ran an experiment for each pairing of a target policy  ∈  with a proxy coach type
discussed in Section 3: (E1) a decision tree with “full path” advice, (E2) a decision tree with “min
prefix” advice, (E3) a random forest, and (E4) an evolutionary mechanism. Decision trees in
experiments (E1), (E2), and (E3) were trained on the exemplar set ̂︀ using ID3 [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Random
forests in experiments (E3) comprised 20 trees, each trained on a 10% random fraction of ̂︀.
The evolutionary mechanism in experiment (E4) used a threshold parameter  = 0, and a
non-linearity parameter  = 2.
      </p>
      <p>The current hypothesis policy was examined at the end of each epoch in each experiment.
Performance values recorded how many of the predictions of the current hypothesis policy on
the evaluation set ℰ were correct against , wrong against , or abstentions, and the accuracy
of the definite predictions (i.e., the number of correct predictions over the number of correct
or wrong predictions). Relative Size values recorded the size of the current hypothesis policy
(relative to the size of ), and the size of the epoch (relative to the size of ), the size of each
given piece of advice (relative to the size of ).</p>
      <p>
        The above values for each proxy coach type were aggregated across  , and were plotted
against the epochs. To account for diferent numbers of epochs across experiments, the set
of epochs of each experiment was uniformly distributed over the interval [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], using
splineinterpolation to fill-in the values between the points that corresponded to the epochs in that
experiment. In particular, a value of 0 corresponds to the pre-coaching state of afairs, where
the current hypothesis policy is empty, and a value of 1 corresponds to the post-coaching state
of afairs, where the current hypothesis policy is the one after the integration of the final piece
of advice.
4.2. Results and Analysis
A first general observation is that experiments agree qualitatively on the eficacy of Machine
Coaching. As more pieces of advice are given, the performance of the current hypothesis policy
improves in terms of increased correct predictions and decreased abstentions. The size of the
current hypothesis policy grows linearly with the given pieces of advice, due to the uniform
distribution of epochs on the  axis. The sizes of the epochs and the given pieces of advice
remain steady or increase, confirming the intuition that advice is given less frequently and
becomes more specific as coaching progresses.
      </p>
      <p>Experiments (E1) and (E2)
The first set of experiments cleanly demonstrates the efect of ofering “cautious” but “specific”
advice. As expected, wrong predictions remain efectively zero throughout the coaching phase,
while abstentions are very gradually replaced with correct predictions, maintaining an efectively
perfect accuracy.</p>
      <p>Analogously, the second set of experiments demonstrates the efect of ofering more “loose”
and “general” advice. Correct predictions increase slightly faster compared to experiments (E1),
while abstentions decrease at an even faster pace, causing the introduction of wrong predictions,
and a sub-perfect accuracy. Interestingly, most wrong predictions are corrected over time,
leading to suficiently high accuracy.</p>
      <p>In both sets of experiments, advice becomes more specific over time and is given less
frequently, with the epoch size increasing significantly during the end of the coaching phase.
These considerations, along with the continual improvement of performance, suggest that the
given advice is indeed beneficial. The quality of the given advice is further supported by the
size of the hypothesis policy remaining distinctly smaller than the size of the target policy,
indicating that the given advice is able to compress parts of the target policy. The more “loose”
and “general” advice in experiments (E2) seems to correspond to a less frequent need for advice,
and a more concise hypothesis policy, at the expense of occasionally making wrong predictions,
but also efectively never abstaining.</p>
      <p>Experiments (E3) and (E4)
The third and fourth sets of experiments demonstrate the efect of ofering advice chosen
greedily through bounded local searching, and without consulting the explanations of the
current hypothesis policy on its predictions. Accordingly, performance is measurably worse
than in the first two sets of experiments, with fewer correct predictions, more wrong predictions,
higher number of abstentions, and lower accuracy.</p>
      <p>Although the last two sets of experiments demonstrate a more aggressive take on decreasing
abstentions and increasing correct predictions earlier on, this seems to come at the expense of
introducing a considerable number of wrong predictions that persist across epochs, indicating a
lower quality of the given advice. This indication is further corroborated by the persistent size
of the advice, which fails to become more “cautious” and “specific” over time, and likely leads
to the introduction of nearly equally-many new wrong predictions in the place of any of the
existing wrong predictions it corrects.</p>
      <p>Relatedly, the epoch size does not increase as rapidly as in the first two sets of experiments,
leading to a larger hypothesis policy size, to the extent, in fact, that the hypothesis policy ends
up being larger than the target policy, showing an inability for compression, and induction, in
the given advice.</p>
      <p>Comparatively, the size of the given piece of advice in experiments (E4) is multiple times
larger than that in experiments (E3). This can be directly attributed to the initialization of the
search used in each case, with the former constructing advice by starting from a fully “cautious”
and “specific” seed advice, and the latter selecting a piece of advice from typically “loose” and
“general” candidate pieces of advice.</p>
      <p>Machine Coaching Eficiency
The primary metric on eficiency for Machine Coaching is the number of epochs required to
get to a certain degree of performance. Figure 4 shows the same information as Figure 3, but
with aggregation happening over each particular epoch. Since diferent target polices lead to
diferent numbers of resulting epochs, the aggregation happens over a diminishing set of target
policies which eventually becomes unrepresentative of the initial set, and should not form the
basis for conclusions.</p>
      <p>Looking, therefore, at the parts of the plots that include at least 50% of the entire set of target
policies being considered, we observe qualitatively that a few (around 10–20) pieces of advice
seem to sufice to get relatively high performance.</p>
      <p>Although reporting absolute computation times is not informative, we do remark that
experiments (E4) are the most time-intensive ones, as a result of the repeated testing against the
entire exemplar set, for each mutation in each generation.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>
        We have put forward proxy (algorithmic) coaches as a means to provide empirical evidence that
complements prior theoretical work on the eficacy and eficiency of Machine Coaching. In
ongoing work, we continue to investigate coaching-based learning further with proxy coaches
in more expressive settings (e.g., relational policies, richer reasoning semantics with rule
chaining, and, relatedly, without assuming complete contexts), and contrast its performance
and scalability against autodidactic learning algorithms that map exemplar data directly into a
form of prioritized rules [
        <xref ref-type="bibr" rid="ref23 ref24 ref25 ref26">23, 24, 25, 26</xref>
        ].
      </p>
      <p>Ultimately, our plan is to undertake empirical studies with humans in the role of direct
coaches, by identifying meaningful solutions to the challenges discussed in Section 1, perhaps
by pairing human coaches with proxy (algorithmic) coaches. Whether human coaches will ofer
advice in a manner analogous to one of the proxy coaches considered herein, whether they will
remain consistent across multiple interactions, and whether they will find the coaching protocol
to be cognitively light, are all questions that we wish to investigate. In turn, we expect that
answers to these questions will ofer guidelines towards improving the cognitive compatibility
of the language, semantics, and protocol of Machine Coaching.</p>
      <p>Acknowledgements This work was supported by funding from the European Regional
Development Fund and the Government of the Republic of Cyprus through the Research and
Innovation Foundation under grant agreement no. INTEGRATED/0918/0032, from the EU’s
Horizon 2020 Research and Innovation Programme under grant agreements no. 739578 and no.
823783, and from the Government of the Republic of Cyprus through the Deputy Ministry of
Research, Innovation, and Digital Policy.</p>
    </sec>
    <sec id="sec-6">
      <title>A. Empirical Setting (Details)</title>
      <p>We synthetically generated a target policy by starting from a set of atoms , an average batch
size , a flip count ℓ, and a number of stacks , and then partitioning ℓ + 1 into  positive
integers 1, . . . ,  ∈ N, and producing for each one of them a stack of depth ,  = 1, . . . , . In
order to generate a stack, given , , and the special atom output, we first generated a batch
containing a single rule, referred to as the stack’s root, with body from  and head output,
and then we iteratively generated batches of average size  and interchangeably conflicting
heads. Although this process does not explicitly determine the constructed target policy’s size
or the average body size of its rules, we manipulated these attributes indirectly, by varying
across policies the average batch size, flip count, number of stacks, and each stack’s root body
size . Namely, for the purposes of our experiments,  and  both ranged from 1 to 31 with a
step of 2, while ℓ ranged from 1 to 31 with a step of 3, and  ranged from 1 to ℓ with a step of 2.
For each assignment of values to , , ℓ, , we generated 10 diferent target policies, to account
(to some extent) for the various possible partitions of ℓ + 1 to  positive integers. Lastly, in all
above cases,  contained 20 atoms. All related metrics that were manipulated during the data
generation stage are illustrated in an example policy in Figure 1.</p>
      <p>We measured the predictive ability of each hypothesis policy against each target policy in  ,
and, in the case of intensional proxy coaches, against the corresponding learned model. We refer
to these metrics as learning performance and learning conformance, respectively, with each of the
two recording correct predictions, wrong predictions, abstentions, and accuracy, as discussed in
Section 4.1. The fitness metric used in the evolutionary mechanism of the extensional proxy
coach was as described in Section 3.2, and is summarized in Table 1.</p>
      <p>t 
n
re 
a
P  

0
1
1</p>
      <p>Ofspring</p>
    </sec>
    <sec id="sec-7">
      <title>B. Performance Results (Details)</title>
      <p>Table 2 contains all results regarding the average final performance using all four proxy coaches,
while in Tables 3 and 4 we present performance on evaluation set aggregated within each of
the several groups of target policies we have produced. For evaluation purposes, we have also
labeled the coaching set and kept its labeled part, so in what follows we report results on all
three datasets (exemplar, evaluation and coaching).</p>
      <p>Most target policy attributes do not seem to significantly afect performance when it comes
to single-tree approaches with the exception of policies containing few flips or large rule
batches. Regarding flips, this behavior is somewhat surprising since policies containing few</p>
      <p>Exemplar</p>
      <p>Final performance with respect to all used sets w.r.t. both the target policy (Accuracy) and the Proxy
Coach (Conformance) (c: “correct”, w: “wrong”, a: “abstain”,  : mean, Q1: 1st quantile, Q2: median, Q3:
lfips are structurally simpler, which was expected to benefit coaching and requires, thus, further
investigation. On the contrary, poorer performance on policies containing large rule batches
was expected since larger batches impose, in general, a richer and more complex policy structure
overall. As far as forests are concerned, with the exception of relatively large policies, where the
corresponding hypotheses seem to be suficiently accurate, correct predictions seem to remain
unafected by variations in target policy’s attributes. This unexpected eficiency on larger target
policies is also subject to further investigation.</p>
      <p>In Figure 5, we present two sets of single policies, each chosen from one of the groups defined
in Section 4.1. Namely, on the left column (Figure 5a) each policy is the group’s median with
respect to the policy attribute varied within that group, while on the right column (Figure 5b),
each policy is the group’s median with respect to the total number of epochs required during
coaching. In both columns, the four stripes correspond to experiments E1-E4, respectively,
while, within each stripe, the first row corresponds to high-value groups and the second row to
low-value ones.</p>
    </sec>
    <sec id="sec-8">
      <title>C. Conformance Results (Details)</title>
      <p>So far, we have discussed results regarding learning eficiency against the target theory
(performance). We have also computed, whenever applicable, learning eficiency with respect to proxy
coaches (conformance). Namely, Conformance values recorded how many of the predictions
of the current hypothesis policy on the evaluation set ℰ were correct against the proxy coach,
wrong against the proxy coach, or abstentions, and the accuracy of the definite predictions
(i.e., the number of correct predictions over the number of correct or wrong predictions). As in
Section 4.1, Relative Size values recorded the size of the current hypothesis policy (relative to</p>
      <p>Exemplar
the size of ), and the size of the epoch (relative to the size of ), the size of each piece of advice
given (relative to the size of ).</p>
      <p>All results regarding the above metrics are presented in Figure 6 (not applicable for E4).
Overall, the trends observed for the three diferent advice protocols are similar to those presented
in Figure 3. Nevertheless, one may observe that there are also some minor diferences. To begin
with, regarding the full path protocol, wrong predictions, in terms of conformance to the proxy
coach’s ones, are zero. This is due to the advice-giving protocol being highly conservative on
the returned advice which, as discussed in Section 3.1, is quite narrow in terms of coverage but
highly accurate at the same time.</p>
      <p>Another interesting fact is that, regardless of whether eficiency is measured against the
target policy or the proxy coach, coaching under the forest protocol is persistently less eficient
than the other two. This, again, could be interpreted as another hint for the ineficiency of
the underlying advice mining mechanism when it comes to forests or just as a side efect of a
forest’s relative opacity compared to single trees.</p>
      <p>0
0
0
0
2
5
Epochs
2
Epochs
Epochs
10
Epochs
10</p>
      <p>20
Epochs
10
Epochs
10
Epochs
Epochs
10
Epochs
0.0
1.0
e0.75
c
n
a
rm0.50
fro
Pe0.25
0.0
1.0
e0.75
c
n
a
rm0.50
fro
Pe0.25
(a) Median policies per group,
w.r.t. each policy
(b) Median policies per group
w.r.t.</p>
      <p>the total
attribute.</p>
      <p>number of epochs.
w.r.t. each group’s manipulated policy
attribute (left column) and the total number of epochs within each group (right column). Within each
column, the four stripes correspond to E1, E2, E3 and E4 respectively. Within each stripe, groups vary
from left to right as follows: policy size, f lip count, stack count, batches and width.</p>
      <p>Also, the top row of
each stripe corresponds to</p>
      <p>high-value groups, while the bottom to low-value ones. Color coding is the
same as in</p>
      <p>Figure
10
Epochs
20
Epochs
10
Epochs
Epochs
e
z
i</p>
      <p>S
1.0 ive
t
a
l
e
0.5 R
0.0</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <source>Autodidactic Learning and Reasoning, Ph.d. thesis</source>
          , School of Engineering and Applied Sciences, Harvard University, USA,
          <year>2008</year>
          . URL: https://dl.acm.org/doi/abs/10.5555/ 1467943.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <source>Partial Observability and Learnability, Artificial Intelligence</source>
          <volume>174</volume>
          (
          <year>2010</year>
          )
          <fpage>639</fpage>
          -
          <lpage>669</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.artint.
          <year>2010</year>
          .
          <volume>03</volume>
          .004.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <source>The Advice Taker 2.0, in: Proceedings of the 13th International Symposium on Commonsense Reasoning</source>
          , volume
          <year>2052</year>
          , London, U.K,
          <year>2017</year>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-2052/#paper13.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          , Machine Coaching,
          <source>in: IJCAI 2019 Workshop on Explainable Artificial Intelligence</source>
          , Macau, China,
          <year>2019</year>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>86</lpage>
          . URL: https://cognition.ouc.ac.cy/loizos/papers/ Michael_2019_
          <article-title>MachineCoaching</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mercier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sperber</surname>
          </string-name>
          , Why Do Humans Reason?
          <article-title>Arguments for an Argumentative Theory</article-title>
          .,
          <source>Behavioral and Brain Sciences</source>
          <volume>34</volume>
          (
          <year>2011</year>
          )
          <fpage>57</fpage>
          -
          <lpage>74</lpage>
          . doi:
          <volume>10</volume>
          .1017/S0140525X10000968.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ioannou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <article-title>Knowledge-Based Translation of Natural Language into Symbolic Form</article-title>
          ,
          <source>in: Proceedings of the 7th Linguistic and Cognitive</source>
          Approaches To Dialog Agents Workshop - LaCATODA
          <year>2021</year>
          , Montreal, Canada,
          <year>2021</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>32</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2935</volume>
          /#paper3.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>J. McCarthy</surname>
          </string-name>
          ,
          <article-title>Programs with Common Sense</article-title>
          ,
          <source>in: Proceedings of the Teddington Conference on the Mechanization of Thought Processes</source>
          , London, U.K,
          <year>1959</year>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>91</lpage>
          . URL: http: //jmc.stanford.
          <source>edu/articles/mcc59/mcc59.pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Valiant</surname>
          </string-name>
          ,
          <article-title>A Theory of the Learnable</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>27</volume>
          (
          <year>1984</year>
          )
          <fpage>1134</fpage>
          -
          <lpage>1142</lpage>
          . doi:
          <volume>10</volume>
          .1145/
          <year>1968</year>
          .
          <year>1972</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Teso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kersting</surname>
          </string-name>
          ,
          <string-name>
            <surname>Explanatory</surname>
          </string-name>
          <article-title>Interactive Machine Learning</article-title>
          ,
          <source>in: Proceedings of the 2019 AAAI/ACM Conference on AI</source>
          ,
          <string-name>
            <surname>Ethics</surname>
          </string-name>
          , and Society, AIES '19,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , pp.
          <fpage>239</fpage>
          -
          <lpage>245</lpage>
          . doi:
          <volume>10</volume>
          .1145/3306618.3314293.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Settles</surname>
          </string-name>
          ,
          <article-title>Active Learning Literature Survey</article-title>
          ,
          <source>Technical Report</source>
          , University of WisconsinMadison Department of Computer Sciences,
          <year>2009</year>
          . URL: https://minds.wisconsin.edu/ handle/1793/60660.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <article-title>"Why Should I Trust You?": Explaining the Predictions of Any Classifier</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , KDD '16,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          . doi:
          <volume>10</volume>
          .1145/2939672.2939778.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>J. McCarthy</surname>
          </string-name>
          , Elaboration Tolerance, in: Common Sense 98, London, U.K,
          <year>1998</year>
          . URL: http://www-formal.stanford.edu/jmc/elaboration.html.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Raghavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Allan</surname>
          </string-name>
          ,
          <article-title>An Interactive Algorithm for Asking and Incorporating Feature Feedback into Support Vector Machines</article-title>
          ,
          <source>in: Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '07,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2007</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>86</lpage>
          . doi:
          <volume>10</volume>
          . 1145/1277741.1277758.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>B.</given-names>
            <surname>Settles</surname>
          </string-name>
          , Closing the Loop:
          <article-title>Fast, Interactive Semi-Supervised Annotation With Queries on Features and Instances</article-title>
          ,
          <source>in: Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , Association for Computational Linguistics, Edinburgh, Scotland, UK.,
          <year>2011</year>
          , pp.
          <fpage>1467</fpage>
          -
          <lpage>1478</lpage>
          . URL: https://aclanthology.org/D11-1136.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Zaidan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eisner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Piatko</surname>
          </string-name>
          ,
          <article-title>Using “Annotator Rationales” to Improve Machine Learning for Text Categorization</article-title>
          ,
          <source>in: Human Language Technologies</source>
          <year>2007</year>
          :
          <article-title>The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference, Association for Computational Linguistics</article-title>
          , Rochester, New York,
          <year>2007</year>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>267</lpage>
          . URL: https://aclanthology.org/N07-1033.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Shivaswamy</surname>
          </string-name>
          , T. Joachims, Coactive Learning,
          <source>Journal of Artificial Intelligence Research</source>
          <volume>53</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          . doi:
          <volume>10</volume>
          .1613/jair.4539.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Prentzas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nicolaides</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kyriacou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kakas</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Pattichis,</surname>
          </string-name>
          <article-title>Integrating Machine Learning with Symbolic Reasoning to Build an Explainable AI Model for Stroke Prediction</article-title>
          ,
          <source>in: 2019 IEEE 19th International Conference on Bioinformatics and Bioengineering (BIBE)</source>
          , Athens, Greece,
          <year>2019</year>
          , pp.
          <fpage>817</fpage>
          -
          <lpage>821</lpage>
          . doi:
          <volume>10</volume>
          .1109/BIBE.
          <year>2019</year>
          .
          <volume>00152</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>R.</given-names>
            <surname>Khardon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roth</surname>
          </string-name>
          , Learning to Reason,
          <source>Journal of the ACM</source>
          <volume>44</volume>
          (
          <year>1997</year>
          )
          <fpage>697</fpage>
          -
          <lpage>725</lpage>
          . doi:
          <volume>10</volume>
          . 1145/265910.265918.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Juba</surname>
          </string-name>
          ,
          <article-title>Implicit Learning of Common Sense for Reasoning</article-title>
          ,
          <source>in: Proceedings of the 23rd International Joint Conference on Artificial Intelligence</source>
          , AAAI Press, Beijing, China,
          <year>2013</year>
          , pp.
          <fpage>939</fpage>
          -
          <lpage>946</lpage>
          . URL: https://www.ijcai.org/Proceedings/13/Papers/144.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Valiant</surname>
          </string-name>
          , Evolvability,
          <source>Journal of the ACM</source>
          <volume>56</volume>
          (
          <year>2009</year>
          )
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . doi:
          <volume>10</volume>
          .1145/1462153. 1462156.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Urbanowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <article-title>Learning Classifier Systems: A Complete Introduction, Review, and Roadmap</article-title>
          ,
          <source>Journal of Artificial Evolution and Applications</source>
          <year>2009</year>
          (
          <year>2009</year>
          )
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2009</year>
          /736398.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Quinlan</surname>
          </string-name>
          ,
          <article-title>Induction of Decision Trees, Machine Learning 1 (</article-title>
          <year>1986</year>
          )
          <fpage>81</fpage>
          -
          <lpage>106</lpage>
          . doi:
          <volume>10</volume>
          . 1023/A:
          <fpage>1022643204877</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Rivest</surname>
          </string-name>
          ,
          <article-title>Learning Decision Lists, Machine Learning 2 (</article-title>
          <year>1987</year>
          )
          <fpage>229</fpage>
          -
          <lpage>246</lpage>
          . doi:
          <volume>10</volume>
          .1023/A:
          <fpage>1022607331053</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dimopoulos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kakas</surname>
          </string-name>
          ,
          <article-title>Learning Non-Monotonic Logic Programs: Learning Exceptions</article-title>
          , in: N.
          <string-name>
            <surname>Lavrac</surname>
          </string-name>
          , S. Wrobel (Eds.),
          <source>Machine Learning: ECML-95</source>
          , volume
          <volume>912</volume>
          of Lecture Notes in Computer Science, Springer, Berlin, Heidelberg,
          <year>1995</year>
          , pp.
          <fpage>122</fpage>
          -
          <lpage>137</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 3-540-59286-5_
          <fpage>53</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          , Causal Learnability,
          <source>in: Proceedings of the Twenty-Second International Joint Conference on Artificial Intelligence</source>
          , AAAI Press, Barcelona, Spain,
          <year>2011</year>
          , pp.
          <fpage>1014</fpage>
          -
          <lpage>1020</lpage>
          . doi:
          <volume>10</volume>
          .5591/978-1-
          <fpage>57735</fpage>
          -516-8/
          <fpage>IJCAI11</fpage>
          -174.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>L.</given-names>
            <surname>Michael</surname>
          </string-name>
          ,
          <article-title>Cognitive Reasoning and Learning Mechanisms</article-title>
          ,
          <source>in: Proceedings of the 4th International Workshop on Artificial Intelligence and Cognition</source>
          , volume
          <volume>1895</volume>
          <source>of CEUR Workshop Proceedings</source>
          , CEUR, New York City, NY,
          <year>2016</year>
          , pp.
          <fpage>2</fpage>
          -
          <lpage>23</lpage>
          . URL: http://ceur-ws.
          <source>org/</source>
          Vol-1895/#AIC16_
          <fpage>paper1</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>