<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Children's Exploration, Aha! Moments and Explanations in Model Building for Self-Regulated Problem-Solving</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vicky Charisi</string-name>
          <email>vasiliki.charisi@ec.europa.eu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Natalia Díaz-Rodríguez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barbara Mawhin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis Merino</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DaSCI Andalusian Institute in Data Science and Computational Intelligence</institution>
          ,
          <addr-line>Univers. of Granada</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Human Factors Department, EBT-Salient Aero Foundation</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Joint Research Centre, European Commission</institution>
          ,
          <addr-line>Seville</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Service Robotics Laboratory, University Pablo de Olavide</institution>
          ,
          <addr-line>Seville</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In certain problem-solving tasks that require Human-AI interactions, a mutual understanding of the reasoning behind the performed actions can benefit both humans and artificial agents. However, identifying and predicting the cognitive strategies involved in such a hybrid setting, especially in novel, self-regulated exploratory tasks, is a challenging endeavour. Our aim is to identify behavioural properties relevant to young children's cognitive strategies that are present in problem-solving, with an emphasis on the Aha! moment as an intermediate step between exploratory actions, that typically relate to the development of tacit knowledge, and the generation of explanations that requires explicit knowledge. We use data from existing, previously published, behavioural studies with children 5 to 7 years old to explore these mechanisms in two selfregulated problem-solving tasks. In addition, we reflect on our observations of an Artificial Agent (Q-learning algorithm) that learns to solve the same task. Our findings indicate that while in current reinforcement learning practice, detecting the moment of the cognitive transformation of the problem representation normally translates into observing convergence curves of the objective functions being optimized, in young children this involves more complex behavioural properties, such as verbal metacognition. These behavioural processes can be used as a proxy for the identification of the Aha! moment. Finally, we propose a conceptual map which integrates the observed behaviours that are used to detect, communicate and corroborate learning both in humans and machines and we discuss the association of children's exploratory behaviours, the Aha! moments and ultimately their explanation generation.</p>
      </abstract>
      <kwd-group>
        <kwd>Explainability</kwd>
        <kwd>Child development</kwd>
        <kwd>Human intelligence</kwd>
        <kwd>Problem-solving</kwd>
        <kwd>Behavioural indicators</kwd>
        <kwd>Explainable AI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        For efective hybrid environments where humans
collaborate with Artificial Intelligence (AI) systems to make a
decision, a mutual understanding of the reasoning behind
certain actions or recommendations can be of catalytic
importance.
tual understanding and trust development [1], and can be
ing models can be explained towards a customized and
diverse set of audiences [
        <xref ref-type="bibr" rid="ref27">2</xref>
        ], debugged, and audited. For
edge should become explicit, which often includes the
cognitive process known as the Aha! moment or
insight. We adopt the definition of the Aha! moment in
problem solving as a sudden transformation of the
problem representation [3, 4]; this difers from the solution
EBeM’22: IJCAI-ECAI Workshop on AI Evaluation Beyond Metrics,
0000-0001-7677-027X (V. Charisi∗); 0000-0003-3362-9326
for AI approaches often requires focusing on infants or
young children in the context of structured or
unstructured activities [
        <xref ref-type="bibr" rid="ref39">5, 6, 7</xref>
        ]. Self-regulated play, for example,
that allows children to perform exploratory actions and
come up with insights and discoveries in problems they
generated has previously been correlated with the
development of their implicit knowledge and their gradual reorganization of component acts and modularization.
understanding of the surrounding physical world [8]. Although Bruner’s examples came from infants in the
However, what cognitive process are mobilized for the ifrst year of life, his ideas have been applied to the
acquisitransformation of tacit into explicit knowledge in young tion of more complex skills beyond infancy. Additionally,
children? And what behavioural properties can be used he argued that play is the best way to promote
developas a proxy for the identification of those processes? ment as it can occur with any physical material or with
      </p>
      <p>
        Based on a series of behavioural studies with children imagination, alone or with others and can take place
5 to 7 years of age, we identify behavioural properties rel- in various settings [10]. The connection of play with
evant to cognitive processes that are present in problem- the development of fundamental cognitive processes and
solving tasks, with an emphasis on identifying the Aha! human learning has been well-established [
        <xref ref-type="bibr" rid="ref25">11, 12, 13</xref>
        ].
moments, as an intermediate step between exploratory Self-directed and intrinsically motivated goal generation
actions and the generation of explanations (see Fig. 1), and problem-solving are among children’s cognitive tools
aiming to inform current and future approaches on ex- that afect their overall development [
        <xref ref-type="bibr" rid="ref39">7, 14</xref>
        ]. In free play,
plainable AI (XAI). children set novel goals, discover unexpected
information, and invent problems they would not otherwise
encounter. In this context, children apply exploratory
pro2. Relevant Work cesses that allow them to progressively reduce
uncertainty about their environment [14].
2.1. Problem-Solving in Young Children In this context, a problem is defined as a situation in
To understand the fundamentals of problem-solving as which a solver needs to change a given state to a desired
a cognitive process, developmental psychologists have one but there are obstacles. There are diferent types of
extensively explored the involved faculties and the ways problems such as the routine problem vs. the non-routine
they interact with each other. To this end, classic and con- problem. The first one refers to a situation in which the
temporary work has examined various tasks that were solver knows a solution method whereas the second is
used depending on the child’s age and areas of interest. when the solver has to create a solution method. There is
Bruner, for example, laid out a plan for the development also the well-defined problem where the state, goal and
of skilled action [9]. First there is intention, then an as- set of operators are clearly defined. It is opposed to the
sembling of “constituent acts”. They initially occur out ill-defined problem where the elements are not clearly
of order but later become properly sequenced to reach defined.
the goal. Bruner emphasised the role of exploratory be- The problem solving process occurs when a person has
haviour and play prior to achieving skilled action. Flex- to invent a way to solve it following two main stages: the
ibility and higher order acts become possible through problem representation and the problem solution. The
solvers need to comprehend the problem and create a
model of the problem situation. Then, they have to build The process named scafolding is described as a process
a solution by using processes of planning, executing and that enables a child or novice to solve a task or achieve a
they have to monitor it using awareness and control. It goal that would be beyond his unassisted eforts [ 20]. To
implies cognitive and metacognitive processes. Problem achieve more complex tasks (like problem-solving), it is
solving is always domain-specific but the thinking by necessary to combine simpler skills in order to achieve
analogy strategy seems to be almost always successful. a higher level of competence. This promotes cognitive
Thinking of a related problem already known and even growth. The shared space of an activity involving
colbetter, already solved, helps for success. An application laboration mechanisms between peers is also at great
of this is the heuristics which allow a solver to go faster importance whether it is a human or an artificial agent
to an acceptable solution even if it is not perfectly ac- [21, 22].
curate. Considering the bounded rationality of humans,
heuristics allows us to make judgements, choices and 2.3. Insight in Problem-Solving
adapt our behaviours eficiently. This is closely related
to the concepts of “social learning” and “adaptation” in Most commonly, this phenomenon is called the “Aha!”
human development. experience describing the moment when a person gets
      </p>
      <p>
        In order to solve a problem, two mental representations the solution to a problem that up to this point had left
are needed: one of the current state and one of the goal her puzzled. In cognitive science it is referred to as
instate. As it is goal-oriented and contextualized, a plan sight problem solving and it is accompanied by a feeling
detailing the solution step by step is required. A constant of satisfaction for the solver. It has been related to
cremonitoring process is also required as each move has ative thinking [23, 24] and includes an exploratory phase
consequences that can bring the solver closer to or further where divergent thinking takes place, especially during
to the desired goal state. It also requires mental flexibility the early stages of the problem solving process. This
aland thus, inhibitory control [15]. If a first chosen solution lows the person to produce new ideas or connect existing
seems to be inappropriate, the solver has to adapt his ideas. The second phase is the convergent thinking phase
strategy. where a solution should be elected by synthesizing,
analyzing and monitoring the matching degree of the current
2.2. Social learning result to the expected one. Although the experience of
insight is sudden and can seem disconnected from the
imSocial learning is a crucial component of human intel- mediately preceding thought, recent research shows that
ligence, allowing us to rapidly adapt to new scenarios, insight is the culmination of a series of brain states and
learn new tasks, and communicate knowledge that can processes operating at diferent time scales. Elucidation
be built on by others [16]. The work of Lev Vygotsky of these precursors suggests interventional opportunities
who put forward this view already in the 1920s takes for the facilitation of insight [3], including concurrent
into account factors such as the language development verbalization [25]. As for every problem solving, most
and cultural influences in the cognitive development of of these strategies rely on a constant restructuring of
children [
        <xref ref-type="bibr" rid="ref2">17</xref>
        ]. From his perspective, mental functioning the mental representation of the problem. One way in
and development rely on an interdependence between which explicit knowledge manifests itself is through the
individual and social processes. When learners, whatever formation of causal inferences and the generation of
extheir age, participate in joint activities, they gain new planations that, in research with children, are used for
abilities and strategies to better understand the world detecting gaps in their causal knowledge.
and adapt to it. This process is also mediated by signs
and tools such as language and mnemonic techniques. 2.4. Explanation generation
Vygotsky folds them in the category of semiotics means.
      </p>
      <p>They are considered as a cornerstone for knowledge co- Regarding explanation generation, there is a large body
construction and can help independent problem-solving of works in various fields. But it always implies an
exactivity. This leads to the diference of what a learner plainer and an explainee with their own respective
charcan do with or without help as he described under the acteristics. Of particular interest across the fields is the
concept of the Zone of Proximal Development (ZPD). The role of the Theory of Mind ie. the ability of a person
social interaction with the use of linguistic and cultural to attribute mental states to the consequent behaviours
tools facilitate the internalization of knowledge and its of herself or others [26]. The selection and evaluation
transformation into cognitive tools supporting the de- processes of explanations depend on the explainer and
velopment of new cognitive functions. The latter aspect explainee, but also on the characteristics of the context.
has been considered for the design of artificial agents The nature of an interaction for explaining is diferent
that are able to interact with others and internalize these in kindergarten between the teacher and a young child
interactions in a similar way as humans [18, 19]. than between the cockpit desk and the pilots during a
lfight. The role of beliefs has also been raised recently
as a cornerstone. An explanation does not necessarily
needs to be consistent with a person’s beliefs but should
help promoting a revision component [27] thus allowing
the evolution of the internal representations. Human
explanations from social sciences became an integrated
part of Artificial Intelligence (AI) through the XAI field
in order to provide explanatory agents and to facilitate
interactions between humans and machines.</p>
    </sec>
    <sec id="sec-2">
      <title>3. Methodological approach</title>
      <sec id="sec-2-1">
        <title>4.1.1. Data analysis</title>
        <sec id="sec-2-1-1">
          <title>For the elaboration of the data we used the approach of microgenetic analysis [30]. The microgenetic method is defined by three properties: (a) observations span a period of rapid change in competency; (b) the density of</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Empirical Studies: A Selection of Use Cases</title>
      <sec id="sec-3-1">
        <title>This section presents a line of empirical evidence that have contributed to our identification of behavioural indicators that facilitated the transition from tacit to explicti</title>
        <p>Code
Spontaneous musicking</p>
        <p>Sound exploration</p>
        <p>Assessment</p>
        <p>Reasoning
Deliberate musicking</p>
        <p>Planning
11.05 Computer-supported music composition was selected as
15.87 an open-ended task which does not include a predefined
27.38 objective final “solution”; rather, it involves
decision1183..606 making based on subjective criteria and self-regulated
14.04 goal identification and provides the context for the
emergence of a variety of processes and interactions. We
Table 1 identify two major findings relevant to the scope of this
The taxonomy of behaviours that emerged during the open- paper; first, despite the unstructured and the highly
exended self-regulated task of children’s collaborative music- ploratory nature of this task, we observed that children
making and the percentage of occurrence per behaviour. exhibited behaviours that correspond to “making” and to
“reflecting”. Spontaneous and exploratory actions were
mixed with deliberate actions and planning while the
latter were supported by assessment and reasoning.
Second, the collaborative setting of this study facilitated
children’s verbal interactions and negotiations during
their decision-making process and consequently their
reasoning and reflection on their actions. These process
correspond to the mobilization of theirverbal
metacognition part of which was the generation of explanations
during the negotiation of their task-related decisions.</p>
        <p>This means that given the opportunity (in this case
collaborative setting), children as young as 5 years old
acFigure 3: Average percentage of children’s behaviours in tively engage in self-initiated reflection on their actions
Study 1, Making (C1, C2, C5 and C6) and Reflecting (C3 and and imagine the future outcomes while being able to
C4). explain their reasoning to the collaborator. However,
we observed that they often lacked the verbal abilities
observations is high relative to the rate of change; and (c) and the terminology for accurate explanations. For this
the observations are subjected to an intensive, trial-trial reason, they mobilised other available modalities, such
analysis to infer the processes that give rise to change. as gestures, and used the afordances of the graphical
The annotation of the data was based on children’s verbal user interface of the tool provided to complement their
and non-verbal behaviours and the corpus included 7063 explanatory behaviours.
annotated behaviours. The taxonomy of the behaviours
that related to children’s cognitive processes and the 4.2. Study 2: An indication for the Aha!
percentage of their occurrence appear in Fig. 1.</p>
        <p>The results indicate that despite the fact that the par- moment
ticipant children were of a relatively young age - which is The goal of this experiment was to test the impact of
typically related to exploratory actions - the behaviours the type of a robot intervention on children’s
problemof deliberate musicking (C5) and planning (C6) appeared solving process. We used the cognitive task of the Tower
slightly more than the exploratory behaviours of sponta- of Hanoi (ToH) [31] which is used to measure children’s
neous musicking (C1) and sound exploration (C2). planning abilities and inhibitory control. To reach the</p>
        <p>Furthermore, a grouping of the behaviours that corre- optimal solution, it requires participants to involve
inspond to reflective actions (C3 and C4) and the ones that hibition of impulsive moves that superficially bring the
correspond to active music-making (C1, C2, C5 and C6) child closer to the goal, but are unhelpful for the
longerreveals that the “reflecting” behaviours occurred 46.18% term solution [32]. We designed a experiment with three
of the total cognitive behaviours, while the active music- phases: a baseline (single child), an intervention
(manipumaking behaviours occurred 53.82% (see Fig. 3). These lation of the robot’s behaviour) and an evaluation (child’s
results indicate that despite the young age of the partici- voluntary interaction with the robot) for  = 20
chilpants, reflecting and reasoning about the musical choices dren 5 to 7 years old. For the intervention phase, we had
appear as an integral part in children’s cognitive engage- two conditions; in Condition1, the robot and the child
ment with music-making. solved the task in a turn-taking setting and in
Condition2 we designed a child-initiated voluntary interaction
with the robot. In this paper, we focus on a single child’s
problem-solving process to explore behavioural
properties relevant for our understanding of the transition from
exploratory actions to the transformation of the
problem representation, which requires the involvement of
inhibitory control and the stabilization of the optimal
performance. The details of the study appear in [33].</p>
        <sec id="sec-3-1-1">
          <title>4.2.1. Data analysis</title>
          <p>We evaluated the task performance in relation to the
trajectory of optimal and suboptimal movements over the
course of the task. The optimal movements are defined
as the ones that lead to the solution of the task with
the minimum number of movements. In addition, we
measured the relevant speed of the movements in relation
to the baseline of each participant. Given the assumption
that during the task the children sustained the necessary
attention, we identify point A in Fig. 4 as the point
that separates the phase of mostly suboptimal moves
(red peaks) with the phase of mostly optimal movements
(blue peaks), which are also carried out faster.
instability, the incremental optimization and the
performance stabilization. During the exploratory phase, the
children were reinforced by the results of their actions
which eventually guided them to the restructuring of the
problem representation and consequently the use of the
strategy which is based on inhibitory control. After the
Aha! moment, we observe a stabilisation of the optimal
moves which indicates learning. One of the limitations of
this study was the fact that it was not designed in a way
to facilitate the child’s verbalisation of their thoughts,
reasoning and reflections. For this reason, we were not
able to make any inferences regarding the children’s
reasoning, their verbal metacognition and the generation
of possible explanations during the problem-solving
process.
4.3. Study 3: Social Interaction and</p>
          <p>Explanations</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>The purpose of study 3 was to explore the role of a social</title>
        <p>robot on children’s collective problem-solving and the
4.2.2. Reflection child-child social dynamics in a setting of two children
We observed exploratory behaviours that typically were and one robot (see Fig. 5). We built upon study 2 and
characterised by increased number of suboptimal moves. we used the same task, the ToH task and the same robot.
We identify as an Aha! moment, the point when a trans- We designed a controlled 2X2 experimental study with
formation of the mental representation of the problem  = 86 children who all participated in a baseline session
occurs which, in this task, is behaviourally manifested (without robot), an intervention (with the manipulation
by the mobilization of inhibition as a strategy for the of robot behaviour, in terms of its cognitive reliability
optimal solution of the task, meaning that the child in- and expressivity) and an evaluation session (with
childhibits the impulsive move and performs the less obvious initiated form of interaction) to solve the Tower of Hanoi
one that will lead to the stabilization of optimal solution task with an incremental dificulty level in the diferent
of the task (see point A in Fig. 4). This is a cognitive experimental configurations without any expert’s
interstrategy that in the age-group of the present studies does vention. For the purposes of this paper, we focus on the
not appear intuitively. As shown in Figure 1, behavioural ifndings on the patterns of children’s social interactions
properties that appear in the problem-solving process in and verbal negotiations and explanations during the
colthe context of the given tasks include the performance lective task performance. The detailed research design,
analysis and findings of the study appear in [ 22].
and planning appeared to be an integral part of the pro- as part of explanation generation. This was more evident</p>
        <sec id="sec-3-2-1">
          <title>4.3.1. Data analysis</title>
          <p>We observed that the setting of the study facilitated
childchild social interaction and verbal reflection, reasoning
cess which was lacking from study 2. To measure the
team disparity, we define social interaction,  , as the
number of task-related interactions between children.
 =
 1 +  2

where   with  = 1, 2 refers to the number of times
child  addresses their peer with a task-related verbal
or non-verbal (i.e. pointing and gestures) behaviours
and L refers to the number of movements needed by the
team to solve the task. Our analysis showed that children
had a higher  rate during the sessions with the robot,
namely the Intervention ( = 0.16,  = 0.14
) and the
Evaluation ( = 0.13,  = 0.092
) which difered
significantly from the Baseline session ( = 0.06,  = 0.09
with  = 0.08 and  = 0.015 . Among the verbal
manifestations we identified the utterances related to planning
as one of the strategies children used to negotiate for
the next movement on the ToH task. We identified the
balance between children in the planning of the
move)
ments, and defined a planning disparity metric, as the
absolute diference in the number of interactions
initivalues: the LA2 tries to solve the game alone while being able
to ask for help whenever its best action is not good enough
(plot not on logarithmic scale as the agent asks for help at
most 7 times)
catalytic for the facilitation of their task-related planning
in the sessions with the robot. One possible explanation
for this is the fact that one of the conditions involved
a robot that suggested suboptimal movements. In that
case, the children engaged in child-child negotiations
and explanation generation to collectively take a
decision for the next move. Our observations indicate that
two cognitive strategies were involved in children’s
explanations, planning as a part of an a priori explanation
of their reasoning for a certain decision and reflecting
as a part of an a posteriori explanation. We need yet to
analyse the association of the strategy of planning in
the context of explanatory behaviours and its relation
to a preceding Aha! moment. It should be noted that
additional non-verbal manifestations, such as pointing
and gestures, were mobilised in the cases that a child did
not have the verbal maturity to formulate the planning
or the explanation.
4.4. Study 4: Multi-agent setting
This study in [34] consists of the same non-open-ended
task (ToH) and collaborative setting: one learning agent
ning performing better ( = 19,  = 0.51,  = 0.40
compared to groups with an unbalanced planning
behaviour ( = 18,  = 1, 61,  = 0.98
). In this case,
planning was used as part of the explanation formation
which was observed to be one of the strategies for
children’s negotiations in problem-solving.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>4.3.2. Reflection</title>
          <p>Children’s social verbal and non-verbal interaction
during the problem-solving process in Study 3 appeared
there was a significant diference in task performance
ated by each child of the team: Our analysis showed that (LA) and one helping agent with focus on the voluntary
interaction among artificial agents. In order to explore
( = 297,  &lt; 0.001
) between teams with a balanced plan- if algorithms benefit from asking for help in
collabora) tive problem-solving, as children do, two hypotheses are
tested:
learning.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>H1: Canonical interventions from an expert speed up</title>
      </sec>
      <sec id="sec-3-4">
        <title>H2: Getting help on demand from an expert accelerates</title>
        <p>ifnding the optimal solution compared to not on demand.</p>
      </sec>
      <sec id="sec-3-5">
        <title>The expert intervention occurs in 2 diferent scenarios:</title>
      </sec>
      <sec id="sec-3-6">
        <title>1) LA1 solves the task in collaboration with the help</title>
        <p>ing agent in a “turn-taking” scenario, which results in
a canonical cognitive intervention from the expert. 2)
LA2 solves the task independently, having the option to
ask for help of the expert whenever (if) this is needed
resulting in an on demand intervention. Two parameters
are assessed: 1) Canonical intervention (help) rate (every
2, 3 or 4 turns), and 2) Ask-for-help threshold (from 0
to 1). The last parameter was created to simulate what
happens when a child asks for help: if the best policy
value is lower than the ask for help parameter, the expert
will play instead of the LA.
4.4.1. Data analysis
2. The lack of a natural language interactive
communication interface impedes questioning the
system about its confidence in real time, or
uncertainty estimates.
3. The dependence of a happy end solving the
task constraints explanations during the
learning phase. How can an agent communicate to
its developer or user its struggles in solving the
task when it has not yet achieved a satisfactory
performance? Could common AI practices (data
augmentation, fine-tuning), be accessible to the
agent for it to communicate and be able to choose
to change them?
4. Most deep RL models rely on baselines of other
agents to assess their worth, lacking comparisons
in multi-agent settings [36] with children
learning.
5. The inability of current RL algorithms to
communicate the continual progress beyond reporting
a sole reward value obtained at the end of a
convergence curve makes it challenging for LAs to
explain their skills to solve the credit assignment
problem, their dificulties or agility to complete
sub-tasks, their acting self-confidence, or learned
savvy behaviours.
6. The lack of alignment of explanations of LAs with
meta-learning and trustworthy AI dimensions
(such as the trust calibration meta-information
taxonomy [1]) should be accounted for in the
explanation generation process, in the same way
as mechanisms to ensure the reproducibility of
insights-built explanations.</p>
      </sec>
      <sec id="sec-3-7">
        <title>From reinforcement learning (RL) plots, as training episodes evolve as a function of the mean number of moves to solve the task, some interpretations are extracted:</title>
        <p>From scenario 1 it is observed that the LA is more
eficient when it is helped by the expert in a turn taking
scenario, and that it is even more efective when helped
every 4 turns rather than every 2. The importance of
exploring on its own is showcased by the agent, rather
than always having the optimal solution.</p>
        <p>Even when all approaches converge in both scenarios,
in scenario 2, the agent that asks for help becomes also
faster and more efective: help is most useful at the
beginning of learning. After asking for help many times
during the first episodes it starts solving the task by itself,
resulting in an increase of ineficient moves. The agent
seems to gain confidence in movements which, while
imperfect, allow the task resolution by exploring diferent
states.</p>
        <p>Compared to the LA not being helped, the
asking-forhelp agent is a lot more eficient, but there is not much
variation among the canonical and the help-asking con- Reflection
ifgurations. This is probably due to the rather simple This is the only study not involving kids, but
followsimulation of the trigger for the request of help. Simu- ing the same protocol as in the ToH studies. While RL
lating the child’s behaviour is a complex task and more model developers communicate a model learning
converemphasis needs to be placed on accurately describing it, gence due to reaching plateaus in learning curves, these
including how to simulate the ”asking for help” function. changes should as well reflect key changes in kid
beAdding mechanisms such as intrinsic motivation, about haviour. However, we showed this is not always evident
the LA’s desire to solve the game on its own, could make to map. Providing the learning agent with signals such as
the comparison more accurate. The agent asks for help the Aha! moment to adjust its self learning
/hyperparamwhen it considers that a movement is not good enough eter changes could be paramount to avoid blind manual
to be played, whereas the actual mechanisms that drive engineering (on e.g., reward function crafting) processes
the child to ask for help are more complex. where no common procedures exist. We believe these</p>
        <p>We identify emergent issues that should be part of the are the explainable dimensions that XAI for RL should
explanation for the design of a LA, and that are not part of work on (identifying Aha! moments, categorising
probthe modelled problem nor a representation is accessible lems, dificulty, environments, collaboration/competition
for the agent and thus, for its explanation: dimensions, etc).</p>
        <p>1. The lack of a multimodal input space [35] for the The dificulties to explain the learning process of a
agent to perform in the action-interaction loop single LA could be reduced by involving interaction with
of RL can restrain the agent from exhibiting the other agents. One could attend to social interaction [16]
correct behaviour and communicating as humans and social influence as intrinsic motivation [ 37] learning
would expect. metrics. Both showed to enhance learning in multiagent
settings.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Explanations should reflect the needs for these incen</title>
        <p>tives that agents depend on to progress. Once an agent
learned, it is not enough that the agent performs tasks
in less time, and better, but also that it uses other human
factors or social outcome metrics such as in [38, 39].</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusions, Limitations and</title>
    </sec>
    <sec id="sec-5">
      <title>Future work</title>
      <p>We presented a set of behavioural studies and an
experiment with a Q-learning algorithm in order to identify
behavioural properties that relate to the Aha! moment,
i.e., the moment of the restructuring of the problem
representation (Fig. 1). These behavioural properties appear
as part of the transition from exploratory to explanatory
behaviours (tacit to explicit knowledge). They include
task-related observations such as performance
instability, incremental optimization and stabilization as well as
verbal metacognitive manifestations (observed only in
two of our studies with children) that involve reasoning,
reflection and planning. These behavioural properties
seem to facilitate the Aha! moment and eventually the
generation of explanations by children.</p>
      <p>In current RL practice, detecting this moment normally
translates into observing convergence curves of the
objective functions being optimized (normally reaching a
plateau in cumulative reward or optimized loss, usually
both). This is an external signal not usually leveraged
by the agent. Although there are exceptions such as the
use of artificial curiosity signals for self learning of the
agent [40]), in regular AI model development practice,
we must highlight the need for easier mechanisms to
convey actions that demonstrate the dificulties of the
agent until convergence plateaus and/or a suficient level
of an XAI metric are reached.</p>
      <p>We summarise the main points we propose to consider
in approaches evaluating XAI, as follows:
1. The Aha! moment (or problem representation
restructuring) acts as the intermediate step between
non-explainable and explainable behaviours. In
a deeper view, since the explanation acts as an
interface between the model and a given target
audience, the Aha! moment is a trigger signal for
a model to start elaborating explanations. More
efort should be put into specifying the meaning
of the Aha! moment in various tasks that RL
models are currently tackling. Defining and
detecting high level policies characterising an Aha!
moment (e.g. in terms of key/exploratory action
sequences) can be signs we should be able to not
only programmatically detect, but also
communicate. In this way, we can achieve explainable and
reliable models, since Aha! signs must act as an
additional proxy to attain trustworthy systems.
2. Children understand but sometimes lack the
cognitive and metacognitive skills to explain.
Explanations are subject to both biological and artificial
systems’ understanding of properties of a given
task, and in young children explanations are
subject to their verbal abilities. Children often use
gestures such as pointing, which means that
explanations that can support human-AI interaction
are subject to tools responsible for social
interaction. Aiming towards human-level AI requires a
broader set of key social skills for complex
embodied communication in multimodal settings within
constantly evolving social worlds [41].
3. The hybrid use of strategies of “planning” and
“insight”: Questioning the false dilemma of
logical reasoning vs machine learning, we argue for
a synergy between these two paradigms in order
to obtain hybrid AI systems.
4. Both social interaction [16] and social influence
as intrinsic motivation [37] show to be enhancers
for learning in multi-agent settings.</p>
      <sec id="sec-5-1">
        <title>This paper is our first attempt to synthesise the results</title>
        <p>of our research on children’s problem solving in diferent
settings and combine them with our research on XAI.
However, due to space limitations, we are limited to
provide overviews without in-depth analysis. We aim to
tackle the latter in our future work.</p>
        <p>We hope this work is useful beyond developmental
robotics and AI, i.e., facilitating an efective and ethical
deployment of RL systems, e.g. from energy building
management to AI for health, where evaluating single
reward functions simply does not reflect nor assess the
complexity of the system nor the dificulties it has to deal
with.</p>
        <p>Future work should aim to involve more tangible
evaluation metrics that both 1) optimize technical
robustness more broadly, and 2) reflect a human-centered view
where machine learning factors are questioned,
monitored and explained in parallel ways to how children
learn. Evaluation mappings across human and machine
learning will allow us to better assess the trade-ofs
between AI assisted decision making and policies.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgments</title>
      <sec id="sec-6-1">
        <title>Charisi is supported by the HUMAINT project of the</title>
        <p>JRC; Díaz-Rodríguez by IJC2019-039152-I funded by
MCIN/AEI /10.13039/501100011033 by “ESF Investing in
your future” and Google Research Scholar Program; and
Merino by Programa Operativo FEDER Andalucia
20142020 through the project DeepBot (PY20_00817).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Learning (ICML)</surname>
          </string-name>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Vygotsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cole</surname>
          </string-name>
          , Mind in society: Develop[1]
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Cancro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Foulds</surname>
          </string-name>
          ,
          <article-title>Tell me something ment of higher psychological processes</article-title>
          , Harvard
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>that will help me trust you: A survey of trust cali-</article-title>
          university press,
          <year>1978</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>bration in human-agent interaction</article-title>
          , arXiv preprint [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Colas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Karch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Moulin-Frier</surname>
          </string-name>
          , P.-Y. Oudeyer,
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>arXiv:2205.02987</source>
          (
          <year>2022</year>
          ).
          <source>Vygotskian autotelic artificial intelligence: Lan</source>
          [2]
          <string-name>
            <given-names>A. B.</given-names>
            <surname>Arrieta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Díaz-Rodríguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Del</given-names>
            <surname>Ser</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Ben- guage and culture internalization for human-like</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>netot</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Tabik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Barbado</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>García</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gil-López</surname>
          </string-name>
          ,
          <source>AI</source>
          , arXiv preprint arXiv:
          <volume>2206</volume>
          .01134 (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Molina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Benjamins</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Explainable</surname>
            artifi- [19]
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Lindblom</surname>
          </string-name>
          , T. Ziemke, Social situatedness of natu-
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>portunities and challenges toward responsible AI</source>
          ,
          <source>Adaptive Behavior</source>
          <volume>11</volume>
          (
          <year>2003</year>
          )
          <fpage>79</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>Information fusion 58</source>
          (
          <year>2020</year>
          )
          <fpage>82</fpage>
          -
          <lpage>115</lpage>
          . [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Bruner</surname>
          </string-name>
          , G. Ross, The role of tutoring [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kounios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Beeman</surname>
          </string-name>
          ,
          <article-title>The aha! moment: The cog- in problem solving</article-title>
          .,
          <string-name>
            <surname>Child</surname>
            <given-names>Psychology</given-names>
          </string-name>
          &amp; Psychiatry
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>nitive neuroscience of insight, Current directions</article-title>
          &amp; Allied
          <string-name>
            <surname>Disciplines</surname>
          </string-name>
          (
          <year>1976</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>in psychological science 18</source>
          (
          <year>2009</year>
          )
          <fpage>210</fpage>
          -
          <lpage>216</lpage>
          . [21]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kerawalla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pearce</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. O'Connor</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Luckin</surname>
            , [4]
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Stuyck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Cleeremans</surname>
            , E. Van den Bussche, Aha!
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Yuill</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Harris</surname>
          </string-name>
          ,
          <article-title>Setting the stage for collabora-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>strained by cognitive load</article-title>
          ,
          <source>Cognition</source>
          <volume>219</volume>
          (
          <year>2022</year>
          )
          <article-title>of shared space</article-title>
          .,
          <source>in: AIED</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>842</fpage>
          -
          <lpage>844</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          104946. [22]
          <string-name>
            <given-names>V.</given-names>
            <surname>Charisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Merino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Escobar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Caballero</surname>
          </string-name>
          , [5]
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Slone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A developmental approach R.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Gómez,</surname>
          </string-name>
          <article-title>The efects of robot cogni-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>to machine learning?, Frontiers in psychology 8 tive reliability and social positioning on child-robot</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          (
          <year>2017</year>
          )
          <article-title>2124</article-title>
          . team dynamics, in: 2021 IEEE International Con[6]
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Lake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Ullman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Tenenbaum</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. J.</surname>
          </string-name>
          <article-title>ference on Robotics and Automation (ICRA)</article-title>
          , IEEE,
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Gershman</surname>
          </string-name>
          ,
          <article-title>Building machines that learn</article-title>
          and
          <source>think</source>
          <year>2021</year>
          , pp.
          <fpage>9439</fpage>
          -
          <lpage>9445</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>like people</article-title>
          ,
          <source>Behavioral and brain sciences 40</source>
          (
          <year>2017</year>
          ). [23]
          <string-name>
            <given-names>T. I.</given-names>
            <surname>Lubart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mouchiroud</surname>
          </string-name>
          , Creativity: A source [7]
          <string-name>
            <given-names>P.-Y.</given-names>
            <surname>Oudeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <article-title>How evolution may work of dificulty in problem solving, The psychology of</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>through curiosity-driven developmental process, problem solving (</article-title>
          <year>2003</year>
          )
          <fpage>127</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>Topics in Cognitive Science</source>
          <volume>8</volume>
          (
          <year>2016</year>
          )
          <fpage>492</fpage>
          -
          <lpage>502</lpage>
          . [24]
          <string-name>
            <given-names>W.</given-names>
            <surname>Carpenter</surname>
          </string-name>
          , The aha! moment: The science be[8]
          <string-name>
            <surname>M. M. Andersen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kiverstein</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Roep- hind creative insights</article-title>
          , in: Toward Super-Creativity-
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          of play (
          <year>2021</year>
          ).
          <article-title>Human-Machine Collaborations</article-title>
          ,
          <source>IntechOpen</source>
          ,
          <year>2019</year>
          . [9]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Bruner</surname>
          </string-name>
          , Organization of early skilled action, [25]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. N.</given-names>
            <surname>MacGregor</surname>
          </string-name>
          , Human performance on
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Child development</surname>
          </string-name>
          (
          <year>1973</year>
          )
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
          <article-title>insight problem solving: A review</article-title>
          ,
          <source>The Journal of</source>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Barrouillet</surname>
          </string-name>
          ,
          <source>Theories of cognitive development: Problem Solving</source>
          <volume>3</volume>
          (
          <year>2011</year>
          )
          <article-title>6</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>From piaget to today</article-title>
          ,
          <source>Developmental Review</source>
          <volume>38</volume>
          [26]
          <string-name>
            <given-names>T.</given-names>
            <surname>Miller</surname>
          </string-name>
          , Explanation in artificial intelligence: In-
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
          <article-title>sights from the social sciences</article-title>
          ,
          <source>Artificial</source>
          <volume>intelli</volume>
          [11]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Siegel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Magid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pelz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Tenenbaum</surname>
          </string-name>
          , gence
          <volume>267</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Schulz</surname>
          </string-name>
          ,
          <article-title>Children's exploratory play tracks the [</article-title>
          27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shvo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Q.</given-names>
            <surname>Klassen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>McIlraith</surname>
          </string-name>
          , Towards
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>cations 12</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          . ternational Workshop on Explainable, Transpar[12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Schulz</surname>
          </string-name>
          , Play, curiosity, and cogni- ent Autonomous Agents and
          <string-name>
            <surname>Multi-Agent</surname>
            <given-names>Systems</given-names>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>tion</surname>
          </string-name>
          , Annual Review of Developmental Psychology Springer,
          <year>2020</year>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>93</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <volume>2</volume>
          (
          <year>2020</year>
          )
          <fpage>317</fpage>
          -
          <lpage>343</lpage>
          . [28]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Yin</surname>
          </string-name>
          , et al.,
          <source>Design and methods</source>
          ,
          <source>Case study [13] A. Gopnik</source>
          ,
          <article-title>Childhood as a solution to explore- research 3 (</article-title>
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>exploit tensions</article-title>
          , Philosophical Transactions of the [29]
          <string-name>
            <given-names>V.</given-names>
            <surname>Charisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Liem</surname>
          </string-name>
          , E. Gomez, Novelty-based
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Royal Society</surname>
            <given-names>B</given-names>
          </string-name>
          375 (
          <year>2020</year>
          )
          <article-title>20190502. cognitive processes in unstructured music-making</article-title>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pelz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kidd</surname>
          </string-name>
          ,
          <article-title>The elaboration of exploratory settings in early childhood</article-title>
          , in: 2018 Joint IEEE
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>play</surname>
          </string-name>
          ,
          <source>Philosophical Transactions of the Royal Soci- 8th International Conference on Development and</source>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>ety B</source>
          <volume>375</volume>
          (
          <year>2020</year>
          )
          <fpage>20190503</fpage>
          .
          <article-title>Learning and Epigenetic Robotics (ICDL-EpiRob)</article-title>
          , [15]
          <string-name>
            <surname>J. M. Unterrainer</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          <string-name>
            <surname>Owen</surname>
            , Planning and
            <given-names>IEEE</given-names>
          </string-name>
          ,
          <year>2018</year>
          , pp.
          <fpage>218</fpage>
          -
          <lpage>223</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>problem solving: from neuropsychology to func</article-title>
          - [30]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Siegler</surname>
          </string-name>
          ,
          <article-title>icrogenetic analyses of learning</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <volume>99</volume>
          (
          <year>2006</year>
          )
          <fpage>308</fpage>
          -
          <lpage>317</lpage>
          . tion, and language, John Wiley &amp; Sons Inc,
          <year>2006</year>
          , p. [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ndousse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jaques</surname>
          </string-name>
          , Learning 464-510.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <article-title>social learning</article-title>
          , Internation Conference on Machine [31]
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Simon</surname>
          </string-name>
          ,
          <article-title>The functional equivalence of prob-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <article-title>lem solving skills</article-title>
          ,
          <source>Cognitive psychology 7</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          268-
          <fpage>288</fpage>
          . [32]
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Zook</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Davalos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. L.</given-names>
            <surname>DeLosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. P.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          and london tasks,
          <source>Brain and cognition 56</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          286-
          <fpage>292</fpage>
          . [33]
          <string-name>
            <given-names>V.</given-names>
            <surname>Charisi</surname>
          </string-name>
          , E. Gomez,
          <string-name>
            <given-names>G.</given-names>
            <surname>Mier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Merino</surname>
          </string-name>
          , R. Gomez,
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <issue>AI 7</issue>
          (
          <year>2020</year>
          )
          <fpage>15</fpage>
          . [34]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bennetot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Charisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Díaz-Rodríguez</surname>
          </string-name>
          , Should
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>ICRA</surname>
          </string-name>
          , Paris/Remote (
          <year>2020</year>
          ). [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Holzinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Emmert-Streib</surname>
          </string-name>
          , R. Cuc-
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <surname>intelligence</surname>
          </string-name>
          ,
          <source>Information Fusion</source>
          <volume>79</volume>
          (
          <year>2022</year>
          )
          <fpage>263</fpage>
          -
          <lpage>278</lpage>
          . [36]
          <string-name>
            <given-names>A.</given-names>
            <surname>Heuillet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Couthouis</surname>
          </string-name>
          , N. Díaz-Rodríguez,
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <source>Knowledge-Based Systems</source>
          <volume>214</volume>
          (
          <year>2021</year>
          )
          <fpage>106685</fpage>
          . [37]
          <string-name>
            <given-names>N.</given-names>
            <surname>Jaques</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lazaridou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Hughes</surname>
          </string-name>
          , C. Gulcehre,
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <year>2019</year>
          , pp.
          <fpage>3040</fpage>
          -
          <lpage>3049</lpage>
          . [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Perolat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Leibo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zambaldi</surname>
          </string-name>
          , C. Beattie,
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <source>Systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ). [39]
          <string-name>
            <given-names>A.</given-names>
            <surname>Heuillet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Couthouis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Díaz-Rodríguez</surname>
          </string-name>
          , Col-
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <source>Computational Intelligence Magazine</source>
          <volume>17</volume>
          (
          <year>2022</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          59-
          <fpage>71</fpage>
          . [40]
          <string-name>
            <given-names>D.</given-names>
            <surname>Pathak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Efros</surname>
          </string-name>
          , T. Darrell,
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <string-name>
            <surname>learning</surname>
          </string-name>
          , PMLR,
          <year>2017</year>
          , pp.
          <fpage>2778</fpage>
          -
          <lpage>2787</lpage>
          . [41]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kovač</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Portelas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          , P.-Y. Oudeyer,
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <source>arXiv:2107.00956</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>