<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Top-Down and Bottom-Up Interactions between Low-Level Reactive Control and Symbolic Rule Learning in Embodied Agents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Clément Moulin-Frier</string-name>
          <email>clement.moulinfrier@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xerxes D. Arsiwalla</string-name>
          <email>x.d.arsiwalla@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jordi-Ysard Puigbò</string-name>
          <email>jordiysard.puigbo@upf.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martì Sanchez-Fibla</string-name>
          <email>santmarti@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Armin Duff</string-name>
          <email>armin.duff@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universitat Pompeu Fabra &amp; ICREA</string-name>
          <email>paul.verschure@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Barcelona</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SPECS Lab, Universitat Pompeu Fabra</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Mammals bootstrap their cognitive structures through embodied interaction with the world. This raises the question: how the reactive control of initial behavior is recurrently coupled to the high-level symbolic representations they give rise to? We investigate this question in the framework of the "Distributed Adaptive Control" (DAC) cognitive architecture, where we study top-down and bottomup interactions between low-level reactive control and symbolic rule learning in embodied agents. Reactive behaviors are modeled using a neural allostatic controller, whereas high-level behaviors are modeled using a biologically-grounded memory network reflecting the role of the prefrontal cortex. The interaction of these modules in a closed-loop fashion suggests how symbolic representations might have been shaped from low-level behaviors and recruited for behavior optimization.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A major challenge in cognitive neuroscience is to propose a unified theory of cognition based on
general models of the brain structure [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Such a theory should be able to explain how specific
cognitive functions (e.g. decision making or planning) result from the particular dynamics of
an embodied cognitive architecture in a specific e nvironment. This has led to various proposals,
formalizing how cognition arises from interaction of functional modules. Early implementations of
cognitive architectures trace back to the era of Symbolic Artificial Intelligence, starting from the
General Problem Solver (GPS, [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]) and has been followed by a number of subsequent architectures
such as Soar [
        <xref ref-type="bibr" rid="ref12 ref13">13, 12</xref>
        ], ACT-R [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] and their follow-ups. These architectures are considered top-down
and representation-based in the sense that they consist of a complex representation of a task, which
Copyright © 2016 for this paper by its authors. Copying permitted for private and academic purposes.
has to be decomposed recursively into simpler ones to be executed by the agent. Although relatively
powerful at solving abstract symbolic tasks, top-down architectures have enjoyed very little success
at bootstrapping behavioral processes and taking advantage of the agent’s embodiment (nonetheless,
several interfaces with robotic embodiment have been proposed, see [
        <xref ref-type="bibr" rid="ref24 ref28">28, 24</xref>
        ]). In contrast to top-down
representation-based approaches, behavior-based robotics [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] emphasizes lower-level sensory-motor
control loops as a starting point of behavioral complexity that can be further extended by combining
multiple control loops together, e.g. as in the subsumption architecture [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Such approaches are
also known as bottom-up and they generally model behavior without relying on complex knowledge
representation and reasoning. This is a significant departure from Newell’s and Anderson’s views
of cognition expressed in GPS, Soar and ACT-R. Top-down and bottom-up approaches thus reflect
different aspects of cognition: high-level symbolic reasoning for the former and low-level embodied
behaviors for the latter. However, both aspects are of equal importance when it comes to defining
a unified theory of cognition. It is therefore a major challenge of cognitive science to unify both
approaches into a single theory, where (a) reactive control allows an initial level of complexity in
the interaction between an embodied agent and its environment and (b) this interaction provides the
basis for learning higher-level symbolic representations and for sequencing them in a causal way
for top-down goal-oriented control. We propose to split the problem of how neural and symbolic
approaches are integrated into the following three research questions:
      </p>
      <p>How are high-level symbolic representations shaped from low-level reactive behaviors in a
bottomup manner?</p>
      <p>How are those representations recruited in rule and plan learning?</p>
      <p>
        How do rules and plans modulate reactive behaviors through top-down control for realizing
longterm goals?
To address these questions, we adopt the principles of the Distributed Adaptive Control (DAC)
theory of the mind and brain [
        <xref ref-type="bibr" rid="ref30 ref31">31, 30</xref>
        ], which posits that cognition is based on the interaction of four
interconnected control loops operating at different levels of abstraction (Fig. 1). The first level is the
embodiment of the agent within its environment, with the sensors and actuators of the agent (called
the Somatic layer). The Somatic layer incorporates physiological needs of the agent (e.g. exploration
or safety) and which drives the dynamics of the whole architecture [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Extending behavior-based
approaches with drive reduction mechanisms, complex behavior is bootstrapped in DAC from the
self-regulation of an agent’s physiological needs when combined with reactive behaviors (the Reactive
layer). This reactive interaction with the environment drives learning processes for acquiring a state
space of the agent-environment interaction (the Adaptive layer) and the acquisition of higher-level
cognitive abilities such as abstract goal selection, memory and planning (the Contextual layer). These
high-level representations in turn modulate behavior at lower levels via top-down pathways shaped
by behavioral feedback. The control flow in DAC is therefore distributed, both from bottom-up
and top-down interactions between layers, as well as from lateral information processing into the
subsequent layers.
      </p>
      <p>
        In this paper, we present biologically-grounded neural models of the reactive and contextual layers
and discuss their possible integration to address the question of how the reactive control of initial
behavior is recurrently coupled to the high-level symbolic representations they give rise to. On one
hand, our model of the reactive layer relies on the concept of allostatic control. Sterling proposes
that allostasis drives regulation through anticipation of needs [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], allostasis is seen as a
reactive meta-regulation system of homeostatic loops that are possibly contradicting each other in a
winner-takes-all process, modulated by emotional and physiological feedback. On the other hand, we
present a neural model of the contextual layer for rule and plan learning grounded in the neurobiology
of the prefrontal cortex (PFC) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. It has been shown that representations of sensory states, actions
and their combinations can be found in the PFC [
        <xref ref-type="bibr" rid="ref16 ref9">16, 9</xref>
        ] and that is reciprocally connected to sensory
as well as motor areas [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This puts the PFC in a favorable position for the representation of
sensorymotor contingencies, i.e. patterns of sensory-motor dependencies [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], selected according to their
relevance in goal-oriented behavior. Its involvement in flexible cognitive control and planning (e.g.
[
        <xref ref-type="bibr" rid="ref10 ref5">10, 5</xref>
        ]) backs the hypothesis that sensory-motor contingencies promote flexible cognitive control and
planning. This paper aims at identifying the key neurocomputational challenges in integrating both
models in a complete cognitive architecture.
      </p>
      <p>
        The next section introduces a model of the reactive layer based on the concept of allostatic control
and is concerned with prioritization of multiple self-regulation loops. We then present an existing
model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] of the prefrontal cortex that is able to integrate sensory-motor contingencies in rules and
plans for long-term reward maximization. Finally, we discuss how both models can be integrated to
bridge the gap between low-level reactive control and high-level symbolic rule learning in embodied
agents. We will make a particular emphasis on how the acquired rules are able to inhibit the reactive
system to achieve long-term goals as proposed in theories on the neuropsychology of anxiety [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
and consciousness [
        <xref ref-type="bibr" rid="ref29 ref3 ref4">29, 3, 4</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>A neural model of allostatic control</title>
      <p>
        In this section we describe the concept of allostatic control and introduces a neural model for it. Based
on the principles of DAC, we consider an embodied agent endowed with physiological needs for e.g.
foraging or safety and self-regulating them through parallel drive reduction control loops. Drives aim
at self-regulating internal state variables within their respective homeostatic ranges. Such an internal
state variable could, for example, reflect the current glucose level in an organism, with the associated
homeostatic range defining the minimum and maximum values of that level. A drive for foraging
would then correspond to a self-regulatory mechanism where the agent actively searches for food
whenever its glucose level is below the homeostatic minimum, and stops eating even if food is present
whenever levels are above the homeostatic maximum. A drive is therefore defined as the real-time
control loop triggering appropriate behaviors whenever the associated internal state variable goes out
of its homeostatic range, as a way to self-regulate its value in a dynamic and autonomous way. It is
common for drives to be competing and we thus require a method to prioritize them. Consider for
example a child in front of a transparent box full of candies, with a caregiver telling her to not open
the box. Two drives are competing in such a situation: one for eating candies and one for obeying
to the caregiver. Depending on how fond of candies she is and of how strict the caregiver is, she
will choose to either break a social rule or to enjoy an extremely pleasant moment. Such regulation
conflicts can be solved through the concept of an allostatic controller [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], defined as a set of parallel
homeostatic control loops operating in real-time and dealing with their prioritization to ensure an
efficient global regulation of multiple internal state variables.
      </p>
      <p>The allostatic control model we introduce in this paper is composed of three subsystems (Fig. 2).
Motor primitive neurons are connected to the agent actuators through synaptic connections. For
example, on a 2-wheeled mobile robot, a "turn right" motor primitive would excite the left wheel
actuator and inhibit the right wheel one, whereas a "go forward" motor primitive would excite
both. A repertoire of behaviors (inner boxes inside the middle box) link sensation neurons (middle
nodes in the behavior boxes) to motor primitive neurons (right nodes in the behavior boxes) through
behavior-specific connections. In a foraging behavior for example, sensing food on the right would
activate a "turn right" motor primitive, whereas in an obstacle avoidance action, sensing an obstacle
on the right would activate a "turn left" motor primitive. In Figure 2, two behaviors are connected to
the two same motor primitives. In the general case however, they could only connect to a specific
relevant subset of the motor primitives. The agent’s exteroceptive sensors (vertical red bar on the
middle) are connected to sensation nodes in the behavior subsystem (middle nodes in each inner
box) by connections not shown in the figure. Each behavior is provided with an input activation
node (left node in each behavior box) which has a baseline activity (indicated by the number 1)
inhibiting the sensation nodes. Therefore, if no input is provided to an activation node, the sensation
nodes of the corresponding behavior are inhibited, preventing information to propagate to the motor
primitives. Drive nodes (left) can activate their associated behavior through the inhibition of the
corresponding activation nodes, that in turn disinhibit sensation nodes through a double inhibition
process, allowing sensory information to propagate to the motor primitives. In Figure 2, two drives
nodes are represented, each connected to its associated behavior which is supposed to reduce the
drive activity (e.g. a foraging drive connecting to food attraction behavior). Drive nodes are activated
through connections from the agent’s interoceptive sensors (vertical red bar on the left) and form a
winner-takes-all process through mutual inhibition. This way, drives compete against each other as in
the child-caregiver example presented above.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Contextual layer: a neural implementation of a rule learning system</title>
      <p>
        In the context of the DAC architecture, we use an existing biologically-grounded model for rule
learning and planning developed in our research group [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] based on key physiological properties of the
prefrontal cortex (PFC), i.e. reward modulated sustained activity and plasticity of lateral connectivity.
Recent proposals highlight the role of sensory–motor contingencies as the building blocks of many
intelligent behaviors including rule learning and planning. Sensory–motor contingencies combine
information about perceptual inputs and related motor actions forming internal states, which are
subsequently used to structure and plan behavior. The DAC architecture has been developed to
investigate how sensory–motor contingencies can be formed and exploited for behavioral control such
as rule learning and flexible planning. Contingencies formed at the level of the adaptive layer provide
inputs to the contextual layer, which acquires, retains, and expresses sequential representations using
systems for short-term and long-term memory. The PFC-grounded contextual layer consists of a
group of laterally connected memory-units. Each memory-unit is selective for one specific stimulus,
e.g. color, and can induce one specific action, e.g. "go left", "go forward" or "go right", forming a
sensory–motor contingency (Fig. 3). A memory-unit can be interpreted as a micro-column comprising
a number of neurons with the same coding properties. Sequential rules are expressed through the
coordinated activation of different memory-units in the correct order. This group of memory-units
forms the elementary substrate for the representation and expressions of rules.
      </p>
      <p>
        The activity of memory-units in this architecture is driven by perceptual inputs, observed reward and
state prediction through the lateral connectivity. Memory-units compete in a probabilistic selection
mechanism. The higher the activity, the higher the probability of being selected. The selected
memory-units propagate their activity and contribute to the final action of the agent. To express a rule,
the activity of these memory-units must be modulated in order to control the selection of specific units
contributing to the final action. The modulated activity of these units is influenced by two systems, the
lateral connectivity and the reward system. The lateral connectivity captures the context and the order
of the sequential rules and influences activity through trigger values. Trigger values allow to chain
through a specific sequence of memory-units. The reward system validates different rules represented
in the network and influences the activity through the reward value. The modulation of each
memoryunit activity is realized by multiplying the perceptual activity by the trigger value as well as the
reward value. This model has been validated in simulated robotic experiments where stimuli-response
associations have to be learned to maximize reward in a multiple T-maze environment, as well in
solving the Towers of London task. Moreover the model is able to re-adapt to a changing environment,
e.g. when the stimuli-response associations are modified on-line [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Drives
      </p>
      <p>Behaviors</p>
      <p>Motor primitives</p>
      <p>Body + Environment</p>
    </sec>
    <sec id="sec-4">
      <title>Bridging the gap between reactive control with continuous sensory-motor contingencies and symbolic prediction through rule learning</title>
      <p>We have presented in the two last sections two neural models.</p>
      <p>The first, that we call the reactive layer, implements an allostatic controller which deals with the
parallel execution of possibly conflicting self-regulation control loops in real-time. It is composed
of three subsystems: Drives that are activated through interoception (related e.g. to the organism’s
glucose level for a foraging drive), Behaviors implemented by reactive control loops (e.g. attraction
to food spot when the glucose level is low) and activation of Motor Primitives that send the required
commands to the agent’s actuators (e.g. turning left or turning right).</p>
      <p>The second, that we call the contextual layer, implements a rule learning system composed of
a large number of memory units that each encodes a specific sensory-motor contingency, i.e. is
sensitive to a particular sensory receptive field and activates a particular action. Lateral connectivity
between units, learned from the agent’s experience, allows the temporal chaining of sensory-motor
contingencies through action. Each unit is associated with a reward value that is also learned and
predict the (possibly delayed) reward expected by executing the corresponding action encoded by the
unit in the corresponding sensory state it is selective to. At a given time, the activity of each memory</p>
      <p>Perception e
Action m
unit is driven by the sensitivity of that unit to the current agent’s perception, the prediction of that
perception from the lateral connections and the reward value learned by the unit over time. The most
activated units then compete together to decide what next action should be performed to maximize
future rewards.</p>
      <p>
        In this section, we discuss how both layers can be integrated in a complete cognitive architecture
where contextual rules are learned from sensory-motor actions generated by reactive control, and
where the actions generated by the contextual layer in turn, modulates the activity of the reactive
system to achieve long-term goals. To illustrate the cognitive dynamics at work, let’s consider again
the example we have described above about a child that experiences a conflict between eating candies
and disobeying to a caregiver. How the reactive and contextual layer models we have presented in
the two last sections interact together in such a situation? Without any previous experience of it,
only the reactive layer generates behavior through the self-regulation of internal drives, here eating
candies and obeying the caregiver. Satisfying one or the other drive depends on their respective
current level as well as the parameters of their mutual inhibition (one could conceive it as personality,
rather a rebel or a good child). Since this situation occurs in various contexts, where the drive initial
activities differ, the consequences of both behaviors are experienced by the child. When she obeys
the caregiver, she is frustrated of not eating a candy. When she disobeys, she has an argument with
the caregiver and likely doesn’t eat any candy anyway according to how strict her opponent is. By
experiencing multiple times the consequences of the two actions, she will discover that respecting
the social rule is of better interest to maximize reward, consequently self-inhibiting her own drive
for eating candies by contextual top-down control acting on the reactive layer. Besides this real-life
example, such a top-down "behavioral inhibition system" has been proposed as a major component
of theories on the neuropsychology of consciousness [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] and anxiety [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        A computational implementation of the reactive and contextual layer integration has to solve the
issues of (a) preprocessing sensory-motor information generated by the reactive layer to provide
relatively abstract and stable perceptions of the environment at the contextual level, (b) modulating
the memory-unit activities through reward values derived from the drive levels, (c) modulating the
activity of drive levels from the output generated by the contextual layer to maximize reward.
Solving (a) requires the addition of an adaptive layer in the architecture that acquires a state space of
the agent-environment interaction through perceptual learning mechanism modulated by reward. In
previous computational models of the DAC architecture, this is accounted as associative learning [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
In the context of the two models presented in this paper, a solution consists of a direct connection
from the sensation neurons of the reactive layer to the perceptual activation of memory units as in Fig.
3. In the current version of the rule learning model, a large number of memory-units is generated
with pseudo-random perceptual sensitivity and motor specificity. Therefore the total number of units
has to be large compared to the number of units that are actually recruited for generating actions
after learning. Adaptively tuning memory-unit sensitivity would allow optimization of the size
of the network, e.g. through a learning rule able to attract perceptual sensitivity of memory units
according to their rewarding effect. Note that a similar learning rule can be applied for learning the
motor specificity of memory-units. Besides a direct connection between the neural units of both
layers, neural hierarchies would also improve rule learning by providing more abstract information
to memory-units. This can be achieved using recent advances in Deep Reinforcement Learning, as
e.g. in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], able to learn highly abstract perceptual representations from raw sensory data driven by
reward; or by adopting a more biologically-grounded approach using neurocomputational models of
perceptual learning grounded in the neurobiology of the cerebral cortex interaction with the amygdala
[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Neuromodulators independent from reward are also present during aversive events, sustained
attention or surprise, modulating not just memory formation but perception as well. The temporal
stabilization of sensory features have been proposed in [32]. Note that not only sensory information
can be abstracted and stabilized, but hierarchies of needs [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and behaviors as well.
To solve (b), reward values can directly be computed from drive activities: the lower the drive activity
the better it is satisfied. Drive-related activities can also modulate the reward value of memory
units according to the physiological context. This way, lateral connectivity among memory units is
coding for the causal link between sensory-motor contingencies, supposed to be context-independent,
whereas reward activation depends on the current configuration of drive levels (one could call it the
emotional state of the agent). This will allow the memory network to behave over different rule sets
according to the internal state of the agent (the rules maximizing reward when an agent is hungry are
not the same as when it is sleepy), avoiding learning interference between different situations.
Finally, regarding (c), we propose that the contextual layer’s activity modulates the reactive one by
directly acting on the drive levels, instead of acting latter in the reactive layer pipeline, e.g. on the
motor primitives. By directly modulating the drive levels, the contextual layer takes full control of
the reactive one by acting on it from the source (see Fig. 2. This allows the agent to self-inhibit some
of its own drives for maximizing reward on the long term as in the child-caregiver example.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper we have designed an allostatic controller and introduced a novel computational model
implementing it. This reactive layer allows the self-regulation of multiple drive-reduction control
loop operating in parallel. Then we have presented an existing computational model of a contextual
layer grounded in the neurobiology of the prefrontal cortex and able to learn sequential rules and
plans through experience. Finally, our main contribution in this paper has been to argue for both a
bottom-up and top-down interaction between low-level reactive control and high-level contextual
plans. Symbolic representations in our approach are learned as sensory-motor contingencies encoded
in discrete memory-units from the activity generated by the reactive layer. In turn, the actions
generated by the contextual layer modulates the activity of the reactive system through a top-down
pathway, inhibiting reactive drives to achieve long-term goals.
Both the reactive and contextual models are implemented and we are now working on their
computational integration. We have identified in the last section the main challenges in term of bottom-up
perceptual abstraction, multitask reward optimization as well as top-down drive modulation. This
will allow a computational implementation of a reactive agent, embodied in a physical mobile robot
and progressively acquiring contextual rules of its environment from experience, thus demonstrating
an increasingly rational behavior.</p>
      <p>
        We will also put a particular emphasis on applying this integrated cognitive architecture to study
the formation of social norms in multi-robot setups [
        <xref ref-type="bibr" rid="ref18 ref19">19, 18</xref>
        ]. The long term goal is to understand
how the constraints imposed by a multi-agent environment favor conscious experience. The research
direction we adopt is based on the hypothesis that social norms are needed for the evolution of large
multi-agent groups and that the formation of those social norms requires each individual to take
conscious control of its own drive system [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>Acknowledgments
Work supported by ERC’s CDAC project:"Role of Consciousness in Adaptive Behavior"
(ERC-2013ADG 341196); &amp; EU projects Socialising Sensori-Motor Contingencies
socSMC-641321—H2020FETPROACT-2014 &amp; What You Say Is What You Did WYSIWYD (FP7 ICT 612139).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Anderson</surname>
          </string-name>
          .
          <article-title>The Architecture of Cognition</article-title>
          . Harvard University Press,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bothell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Byrne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Douglass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lebiere</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qin</surname>
          </string-name>
          .
          <article-title>An integrated theory of the mind</article-title>
          .
          <source>Psychological review</source>
          ,
          <volume>111</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1036</fpage>
          -
          <lpage>1060</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X. D.</given-names>
            <surname>Arsiwalla</surname>
          </string-name>
          , I. Herreros,
          <string-name>
            <given-names>C.</given-names>
            <surname>Moulin-Frier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanchez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Is consciousness a control process?</article-title>
          <source>In International Conference of the Catalan Association for Artificial Intelligence</source>
          , pages
          <fpage>233</fpage>
          -
          <lpage>238</lpage>
          . IOS,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X. D.</given-names>
            <surname>Arsiwalla</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Herreros</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>On three categories of conscious machines</article-title>
          .
          <source>In Conference on Biomimetic and Biohybrid Systems</source>
          , pages
          <fpage>389</fpage>
          -
          <lpage>392</lpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>W. F.</given-names>
            <surname>Asaad</surname>
          </string-name>
          , G. Rainer, and
          <string-name>
            <given-names>E. K.</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>Task-specific neural activity in the primate prefrontal cortex</article-title>
          .
          <source>Journal of Neurophysiology</source>
          ,
          <volume>84</volume>
          (
          <issue>1</issue>
          ):
          <fpage>451</fpage>
          -
          <lpage>459</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Brooks</surname>
          </string-name>
          .
          <article-title>A robust layered control system for a mobile robot</article-title>
          .
          <source>IEEE Journal on Robotics and Automation</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>14</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Brooks</surname>
          </string-name>
          .
          <article-title>Intelligence without representation</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>47</volume>
          (
          <issue>1-3</issue>
          ):
          <fpage>139</fpage>
          -
          <lpage>159</lpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Duff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Fibla</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>A biologically based model for the integration of sensory-motor contingencies in rules and plans: A prefrontal cortex based extension of the distributed adaptive control architecture</article-title>
          .
          <source>Brain research bulletin</source>
          ,
          <volume>85</volume>
          (
          <issue>5</issue>
          ):
          <fpage>289</fpage>
          -
          <lpage>304</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fuster</surname>
          </string-name>
          . The Prefrontal Cortex: Anatomy, Physiology, and
          <article-title>Neurophysiology of the Frontal Lobe</article-title>
          . Lippincott-William &amp; Wilkins, Philadelphia,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J. M. Fuster</surname>
            ,
            <given-names>G. E.</given-names>
          </string-name>
          <string-name>
            <surname>Alexander</surname>
          </string-name>
          , et al.
          <article-title>Neuron activity related to short-term memory</article-title>
          .
          <source>Science</source>
          ,
          <volume>173</volume>
          (
          <issue>3997</issue>
          ):
          <fpage>652</fpage>
          -
          <lpage>654</lpage>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Gray</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>McNaughton</surname>
          </string-name>
          .
          <article-title>The neuropsychology of anxiety: An enquiry into the function of the septo-hippocampal system</article-title>
          .
          <source>Number 33</source>
          . Oxford university press,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Laird</surname>
          </string-name>
          .
          <article-title>The Soar Cognitive Architecture</article-title>
          . MIT Press,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Laird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Newell</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Rosenbloom. SOAR</surname>
          </string-name>
          :
          <article-title>An architecture for general intelligence</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>33</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>64</lpage>
          ,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Marcos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ringwald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Duff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sánchez-Fibla</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>The hierarchical accumulation of knowledge in the distributed adaptive control architecture</article-title>
          .
          <source>In Computational and robotic models of the hierarchical organization of behavior</source>
          , pages
          <fpage>213</fpage>
          -
          <lpage>234</lpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Maslow</surname>
          </string-name>
          .
          <article-title>A theory of human motivation</article-title>
          .
          <source>Psychological review</source>
          ,
          <year>1943</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>E. K.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Desimone</surname>
          </string-name>
          .
          <article-title>Activity of neurons in anterior inferior temporal cortex during a short-term memory task</article-title>
          .
          <source>The Journal of Neuroscience</source>
          ,
          <volume>13</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1460</fpage>
          -
          <lpage>1478</lpage>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mnih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Rusu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Veness</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Bellemare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riedmiller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Fidjeland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Ostrovski</surname>
          </string-name>
          , et al.
          <article-title>Human-level control through deep reinforcement learning</article-title>
          .
          <source>Nature</source>
          ,
          <volume>518</volume>
          (
          <issue>7540</issue>
          ):
          <fpage>529</fpage>
          -
          <lpage>533</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Moulin-Frier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanchez-Fibla</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Autonomous development of turntaking behaviors in agent populations: a computational study</article-title>
          .
          <source>In IEEE International Conference on Development and Learning</source>
          , ICDL/Epirob, Providence (RI), USA,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Moulin-Frier</surname>
          </string-name>
          and
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Two possible driving forces supporting the evolution of animal communication. comment on" towards a computational comparative neuroprimatology: Framing the language-ready brain" by michael a</article-title>
          .
          <source>arbib. Physics of life reviews</source>
          ,
          <volume>16</volume>
          :
          <fpage>88</fpage>
          -
          <lpage>90</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Newell</surname>
          </string-name>
          .
          <article-title>Unified theories of cognition</article-title>
          . Harvard University Press,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Newell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Shaw</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Simon</surname>
          </string-name>
          .
          <article-title>Report on a general problem-solving program</article-title>
          .
          <source>IFIP Congress</source>
          , pages
          <fpage>256</fpage>
          -
          <lpage>264</lpage>
          ,
          <year>1959</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>J. K. O'Regan</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Noë</surname>
          </string-name>
          .
          <article-title>A sensorimotor account of vision and visual consciousness</article-title>
          .
          <source>The Behavioral and brain sciences</source>
          ,
          <volume>24</volume>
          (
          <issue>5</issue>
          ):
          <fpage>939</fpage>
          -
          <lpage>73</lpage>
          ; discussion 973-1031, oct
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>J.-Y. Puigbò</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Moulin-Frier</surname>
            , and
            <given-names>P. F.</given-names>
          </string-name>
          <string-name>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Towards self-controlled robots through distributed adaptive control</article-title>
          .
          <source>In Conference on Biomimetic and Biohybrid Systems</source>
          , pages
          <fpage>490</fpage>
          -
          <lpage>497</lpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>J.-Y. Puigbo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Pumarola</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Angulo</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Tellez</surname>
          </string-name>
          .
          <article-title>Using a cognitive architecture for general purpose service robot control</article-title>
          .
          <source>Connection Science</source>
          ,
          <volume>27</volume>
          (
          <issue>2</issue>
          ):
          <fpage>105</fpage>
          -
          <lpage>117</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>J.-Y. Puigbò</surname>
            , G. Maffei,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ceresa</surname>
          </string-name>
          , G.
          <string-name>
            <surname>-B. M.</surname>
          </string-name>
          <article-title>A</article-title>
          ., and
          <string-name>
            <given-names>V.</given-names>
            <surname>P.F.M.J.</surname>
          </string-name>
          <article-title>Learning relevant features through a two-phase model of conditioning</article-title>
          .
          <source>IBM Journal of Research</source>
          and Development, Special issue on Computational Neuroscience, Accepted, in Press.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>M. Sanchez-Fibla</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Bernardet</surname>
            , E. Wasserman,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Pelc</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Mintz</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Jackson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lansink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Pennartz</surname>
            , and
            <given-names>P. F.</given-names>
          </string-name>
          <string-name>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Allostatic control for robot behavior regulation: a comparative rodent-robot study</article-title>
          .
          <source>Advances in Complex Systems</source>
          ,
          <volume>13</volume>
          (
          <issue>03</issue>
          ):
          <fpage>377</fpage>
          -
          <lpage>403</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sterling</surname>
          </string-name>
          .
          <article-title>Allostasis: a model of predictive regulation</article-title>
          .
          <source>Physiology &amp; behavior</source>
          ,
          <volume>106</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>G.</given-names>
            <surname>Trafton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hiatt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Harrison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tamborello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khemlani</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Schultz. ACT-R/E: An</surname>
          </string-name>
          <article-title>Embodied Cognitive Architecture for Human-Robot Interaction</article-title>
          .
          <source>Journal of Human-Robot Interaction</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>30</fpage>
          -
          <lpage>55</lpage>
          ,
          <year>Mar 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Verschure</surname>
          </string-name>
          .
          <article-title>Synthetic consciousness: the distributed adaptive control perspective</article-title>
          .
          <source>Phil. Trans. R. Soc. B</source>
          ,
          <volume>371</volume>
          (
          <issue>1701</issue>
          ):
          <fpage>20150448</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>P. F. M. J. Verschure</surname>
            ,
            <given-names>C. M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Pennartz</surname>
            , and
            <given-names>G. Pezzulo.</given-names>
          </string-name>
          <article-title>The why, what, where, when and how of goal-directed choice: neuronal and computational principles</article-title>
          .
          <source>Philosophical Transactions of the Royal Society B: Biological Sciences</source>
          ,
          <volume>369</volume>
          (
          <issue>1655</issue>
          ):
          <fpage>20130483</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>P. F. M. J. Verschure</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Voegtlin</surname>
            , and
            <given-names>R. J.</given-names>
          </string-name>
          <string-name>
            <surname>Douglas</surname>
          </string-name>
          .
          <article-title>Environmentally mediated synergy between perception and behaviour in mobile robots</article-title>
          .
          <source>Nature</source>
          ,
          <volume>425</volume>
          (
          <issue>6958</issue>
          ):
          <fpage>620</fpage>
          -
          <lpage>624</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>