<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Developmental AI: The paradox of ravenous intelligent agents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michelangelo Diligenti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Gori</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Maggini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DII - University of Siena</institution>
        </aff>
      </contrib-group>
      <fpage>25</fpage>
      <lpage>27</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In spite of extraordinary achievements in specific tasks,
nowadays intelligent agents are still striving for acquiring
a truly ability to deal with many challenging human
cognitive processes, especially when a mutable environment is
involved. In the last few years, the progressive awareness
on that critical issue has led to develop interesting bridging
mechanisms between symbolic and sub-symbolic
representations and to develop new theories to reduce the huge gap
between most approaches to learning and reasoning. While the
search for such a unified view of intelligent processes might
still be an obliged path to follow in the years to come, in
this paper, we claim that we are still trapped in the insidious
paradox that feeding the agent with the available information,
all at once, might be a major reason of failure when
aspiring to achieve human-like cognitive capabilities. We claim
that the children developmental path, as well as that of
primates, mammals, and of most animals might not be primarily
the outcome of biologic laws, but that it could be instead the
consequence of a more general complexity principle,
according to which the environmental information must properly
be filtered out so as to focus attention on “easy tasks.” We
claim that this leads necessarily to stage-based
developmental strategies that any intelligent agent must follow, regardless
of its body.</p>
    </sec>
    <sec id="sec-2">
      <title>Developmental path and focus of attention</title>
      <p>There a number of converging indications that most of
nowadays approaches to learning and reasoning have been
bouncing against the same wall (in the case of sequential
information, see e.g. [Frasconi et al., 1995]). This is especially
clear when facing cognitive tasks that involve both learning
and reasoning capabilities, that is when symbolic and
subsymbolic representations of the environment need to be
properly bridged. A unified approach to embrace the behavior of
intelligent agents involved in both perceptual and symbolic
information is based on expressing learning data and explicit
knowledge by constraints [Diligenti et al., 2010]. Following
that framework, let us consider tasks that can be formalized
by expressing a parsimonious solution consistent with a given
set of constraints C = f 1; : : : ; qg. It is worth mentioning
that the in the case of supervised learning, the parsimonious
satisfaction of the constraints is reduced to a finite collection
of points according to the classic statistical framework
behind kernel machines. We consider an agent which operates
dynamically in a mutable environment where, at each time
t, it is expected to access only a limited subset Ct CU of
constraints, where CU can be thought of as the universal set
of constraints. Of course, any agent of relevant interest might
be restricted to acquire a limited set of constraints C, so as
8t 2 T : Ct C CU . Instead of following a
developmental path, one could think of agents that acquire C all at
once.</p>
      <p>Definition 2.1 A ravenous agent is one which accesses the
whole constraint set at any step, that is one for which
8t 2 T : Ct = C.</p>
      <p>At first a glance, ravenous agents seem to have more
chances to develop an efficient and effective behavior, since
they can access all the information expressed by C at any
time. However, when bridging symbolic and sub-symbolic
models one often faces the problem of choosing a
developmental path. It turns out that accessing all the information
at once might not be a sound choice in terms of complexity
issues.</p>
      <sec id="sec-2-1">
        <title>The paradox of ravenous agents: Ravenous agents</title>
        <p>are not the most efficient choice to achieve a parsimonious
constraint consistency.</p>
        <p>To support the paradox, we start noting that
hierarchical modular architectures used in challenging perceptual
tasks like vision and speech are just a way to introduce
intermediate levels of representation, so as to focus on simplified
tasks. For example, in speech understanding, phonemes
and words could be intermediate steps for understanding
and take decisions accordingly. Similarly, in vision, SIFT
features could be an intermediate representation to achieve
the ability to recognize objects. However, when looking for
deep integration of sub-symbolic and symbolic levels the
issue is more involved and mostly open. We discuss three
different contests that involve different degree of symbolic
and sub-symbolic representations.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Developmental paths</title>
        <p>Reasoning in the environment - When thinking of
circumscription and, in general, of a non-monotonic
reasoning, one immediately realizes that we are addressing
issues that are outside the perimeter of ravenous agents.
For example, we can start from a default assumption,
that is the typical “bird flies.” This leads us to conclude
that if a given animal is known to be a bird, and nothing
else is known, it can be assumed to be able to fly. The
default assumption can be retracted in case we
subsequently learn that the animal is a penguin. Something
similar happens during the learning of the past tense
of English verbs [Rumelhart and McClelland, 1986],
that is characterized by three stages: memorization of
the past tense of a few verbs, application of the rule
of regular verbs to all verbs and, finally, acquisition of
the exceptions. Related mechanisms of retracting
previous hypotheses arise when performing abductive
reasoning. For example, the most likely explanation when
we see wet grass is that it rained. This hypothesis,
however, must be retracted if we get to know that the real
cause of the wet grass was simply a sprinkler. Again,
we are in presence of non-monotonic reasoning.
Likewise, if a logic takes into account the handling of
something which is not known, it should not be monotonic.
A logic for reasoning about knowledge is the
autoepistemic logic, which offers a formal context for the
representation and reasoning of knowledge about knowledge.
Once again, we rely on the assumption of not to
construct ravenous agents, which try to grasp all the
information at once, but on the opposite, we assume that the
agent starts reasoning with a limited amount of
information on the environment, and that there is a mechanism
for growing up additional granules of knowledge.
Nature helps developing vision - The visual behavior of
some animals seems to indicate that motion plays a
crucial role in the process of scene understanding. Like
other animals, frogs could starve to death if given only a
bowlful of dead flies 1, whereas they catch and eat
moving flies [Lettvin et al., 1968]. This suggests that their
excellent hunting capabilities depend on the acquisition
of a very good vision of moving objects only. Saccadic
eye movements play an important role in facilitating
human vision and, in addition, the autonomous motion is
of crucial importance for any animal in vision
development. Birds, some of which exhibit proverbial abilities
to discover their preys (e.g. eagles, terns), are known to
detect slowly moving objects, an ability that is likely to
have been developed during evolution, since the objects
they typically see are far away when flying. A detailed
investigations on fixational eye movements across
vertebrate indicates that micro-saccades appear to be more
important in foveate than afoveate species and that
saccadic eye movements seem to play an important role in
perceiving static images [Martinez-Conde and Macknik,
2008]. When compared with other animals, humans are
likely to perform better in static - or nearly static -
vision simply because they soon need to look at objects
and pictures thoughtfully. However, this comes at a late
1In addition to their infrared vision, snakes are also known to
react much better to quick movements.
stage during child development, jointly with the
emergence of other symbolic abilities. The same is likely to
hold for amodal perception and for the development of
strange perceptive behavior like popular Kanizsa
triangle [Kanizsa, 1955]. We claim that image
understanding is very hard to attack, and that the fact that humans
brilliantly solve the problem might be mostly due to the
natural embedding of pictures into visual scenes.
Human perception of static images seems to be a higher
level quality that might have been acquired only after
having gained the ability of detecting moving objects.
In addition, the saccadic eye movements might suggest
that static images are just an illusory perception, since
human eyes always perceive moving objects. As a
consequence, segmentation and recognition are only
apparently separate processes: They could be in fact two faces
of the same medal that are only regarded as separate
phases mostly because of the bias induced by years of
research in pattern recognition aimed at facing specific
perceptual tasks. Interestingly, animals which develop
a remarkable vision system, and especially humans, in
their early stages characterized by scarce motion
abilities, deal primarily with moving objects from a fixed
background, which facilitates their segmentation and,
consequently, their recognition. Static and nearly static
vision comes later in human developments, and it does
not arise at all in many animals like frogs. The vision
mechanisms seems to be the outcome of complex
evolutive paths (e.g. our spatial reasoning inferences, which
significantly improve vision skills, only emerge in late
stage of development). We subscribe claims from
developmental psychology according to which, like for other
cognitive skills, its acquisition follows rigorous
stagebased schemes [Piaget, 1961]. The motion is likely to be
the secret of vision developmental plans: The focus on
quickly moving objects at early stages allows the agent
to ignore complex background.</p>
        <p>On the bridge between learning and logic
Let us consider the learning task sketched in Fig. 2,
where supervised examples and FOL predicates can be
expressed in the same formalism of constraints. There
is experimental evidence to claim that a ravenous agent
which makes use of all the constraints C (supervised
pairs and FOL predicates) is not as effective as one
which focuses attention on the supervised examples and,
later on, continues incorporating the predicates
[Diligenti et al., 2010]. Basically, the developmental path
which first favors the sub-symbolic representations leads
to a more effective solution. The effect of this
developmental plan is to break up the complexity of
learning jointly examples and predicates. Learning turns out
to be converted into an optimization problem, which is
typically plagued by the presence of sub-optimal
solutions. The developmental path which enforces the
learning from examples at the first stage is essentially a way
to circumvent local minima.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>
        This paper supports the position that stage-based learning, as
discussed in developmental psychology is not the outcome
of biology, but is instead the consequence of optimization
principles and complexity issues that hold regardless of the
body. This position is supposed to re-enforce recent studies
on developmental AI more inspired to studies in cognitive
development
        <xref ref-type="bibr" rid="ref7">(see e.g. [Sloman, 2009])</xref>
        and is somehow
coherent with the growing interest in deep architectures and
learning [Bengio et al., 2009].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>We thank Luciano Serafini for quence processing: a constrained non-deterministic approach</article-title>
          .
          <source>Knowledge-Based Systems</source>
          ,
          <volume>8</volume>
          (
          <issue>6</issue>
          ):
          <fpage>313</fpage>
          -
          <lpage>332</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Kanizsa</source>
          , 1955]
          <string-name>
            <given-names>G.</given-names>
            <surname>Kanizsa. Margini</surname>
          </string-name>
          quasi
          <article-title>-percettivi in campi con stimolazione omogenea</article-title>
          .
          <source>Rivista di Psicologia</source>
          ,
          <volume>49</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7</fpage>
          -
          <lpage>30</lpage>
          ,
          <year>1955</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Lettvin et al.,
          <year>1968</year>
          ]
          <string-name>
            <given-names>J.Y.</given-names>
            <surname>Lettvin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.H.</given-names>
            <surname>Maturana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.S.</given-names>
            <surname>McCulloch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.H.</given-names>
            <surname>Pitts</surname>
          </string-name>
          .
          <article-title>What the frog's eye tells the frog's brain</article-title>
          .
          <source>In Reprinted from - The Mind: Biological</source>
          Approaches to its functions - Eds
          <string-name>
            <surname>Willian C. Corning</surname>
          </string-name>
          , Martin Balaban, pages
          <fpage>233</fpage>
          -
          <lpage>358</lpage>
          .
          <year>1968</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Martinez-Conde and Macknik</source>
          , 2008]
          <string-name>
            <given-names>S.</given-names>
            <surname>Martinez-Conde</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.L.</given-names>
            <surname>Macknik</surname>
          </string-name>
          .
          <article-title>Fixational eye movements across vertebrates: Comparative dynamics, physiology, and perception</article-title>
          .
          <source>Journal of Vision</source>
          ,
          <volume>8</volume>
          (
          <issue>14</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Piaget</source>
          , 1961]
          <string-name>
            <given-names>J.</given-names>
            <surname>Piaget</surname>
          </string-name>
          . La psychologie de l'intelligence.
          <source>Armand Colin</source>
          , Paris,
          <year>1961</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Rumelhart and McClelland</source>
          , 1986]
          <string-name>
            <given-names>D.E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.L.</given-names>
            <surname>McClelland</surname>
          </string-name>
          .
          <article-title>On learing the pat tense of english verbs</article-title>
          .
          <source>In Parallel Distributed Processing</source>
          , Vol.
          <volume>2</volume>
          , pages
          <fpage>216</fpage>
          -
          <lpage>271</lpage>
          .
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Sloman</source>
          ,
          <year>2009</year>
          ]
          <string-name>
            <given-names>A</given-names>
            <surname>Sloman</surname>
          </string-name>
          .
          <article-title>Ontologies for baby animals and robots. from baby stuff to the world of adult science: Developmental ai from a kantian viewpoint</article-title>
          .
          <source>Technical report</source>
          , University of Birmingham,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>