<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Neural Sub-Symbolic Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreas Wichert</string-name>
          <email>andreas.wichert@ist.utl.pt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics INESC-ID / IST - Technical University of Lisboa Portugal</institution>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>The sub-symbolical representation often corresponds to a pattern that mirrors the way the biological sense organs describe the world. Sparse binary vectors can describe sub-symbolic representation, which can be efficiently stored in associative memories. According to the production system theory, we can define a geometrically based problemsolving model as a production system operating on sub-symbols. Our goal is to form a sequence of associations, which lead to a desired state represented by sub-symbols, from an initial state represented by sub-symbols. We define a simple and universal heuristics function, which takes into account the relationship between the vector and the corresponding similarity of the represented object or state in the real world.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        One form of distributed representation corresponds to a
pattern that mirrors the way the biological sense organs describe
the world. Sense organs sense the world by receptors. By the
given order of the receptors the living organisms experience
the reality as a simple Euclidian geometrical world. Changes
in the world correspond to the changes in the distributed
representation. Prediction of these changes by the nervous
system corresponds to a simple geometrical reasoning process.
Mental imagery problem solving is an example for a complex
geometrical problem- solving. It is described by a sequence
of associations, which progressively change the mental
imagery until a desired solution of a problem is formed. For
example, do the skis fit in the boot of my car? Mental
representations of images retain the depictive properties of the
image itself as perceived by the eye [Kosslyn, 1994]. The
imagery is formed without perception through the
construction of the represented object from memory. Symbols on the
other hand are not present in the world; they are the
constructs of human mind to simplify the process of problem
solving. Symbols are used to denote or refer to something
other than them, namely other things in the world
        <xref ref-type="bibr" rid="ref31">(according to the pioneering work of Tarski [Tarski, 1956])</xref>
        . They
are defined by their occurrence in a structure and by a formal
language, which manipulates these structures [Simon, 1991;
      </p>
      <p>Newell, 1990]. In this context, symbols do not by themselves,
represent any utilizable knowledge. They cannot be used for
a definition of similarity criteria between themselves. The
use of symbols in algorithms which imitate human intelligent
behavior led to the famous physical symbol system
hypothesis by Newell and Simon (1976) [Newell and Simon, 1976]:
The necessary and sufficient condition for a physical system
to exhibit intelligence is that it be a physical symbol system.
We do not agree with the physical symbol system hypothesis.
Instead we state that the actual perception of the world and
manipulation in the world by living organisms lead to the
invention or recreation of an experience, which at least in some
respects, resembles the experience of actually perceiving and
manipulating objects in the absence of direct sensory
stimulation. This kind of representation is called sub-symbolic.
Subsymbolic representation implies heuristic functions. Symbols
liberate us from the reality of the world although they are
embodied in geometrical problem solving through the usage of
additional heuristics functions. Without the use of heuristic
functions real world problems become intractable.</p>
      <p>The paper is organized a follows: We review the
representation principles of objects by features as used in cognitive
science. In the next step we indicate how the
perceptionoriented representation is build on this approach. We define
the sparse sub-symbolical representation. Finally, we will
introduce the sub-symbolical problem solving which relies on
a sensorial representation of the reality.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Sub-symbols</title>
      <p>Perception-oriented representation is an example of
subsymbolical representation, such as the representation of
numbers by the Oksapmin tribe of Papua New Guinea. The
Oksapmin tribe of Papua New Guinea counts by associating a
number with the position of the body [Lancy, 1983]. The
sub-symbolical representation often corresponds to a pattern
that mirrors the way the biological sense organs describe the
world. Vectors represent patterns. A vector is only a
subsymbol if there is a relationship between the vector and the
corresponding similarity of the represented object or state in
the real world through sensors or biological senses. Feature
based representation is an example of sub-symbolical
representation.
Objects can be described by a set of discrete features, such
as red, round and sweet [Tversky, 1977; McClelland and
Rumelhart, 1985]. The similarity between them can be
defined as a function of the features they have in common
[Osherson, 1995; Sun, 1995; Goldstone, 1999; Gilovich,
1999]. The contrast model of Tversky [Tversky, 1977] is one
well-known model in cognitive psychology [Smith, 1995;
Opwis and Plo¨ tzner, 1996] which describes the similarity
between two objects which are described by their features. An
object is judged to belong to a verbal category to the extent
that its features are predicted by the verbal category
[Osherson, 1987]. The similarity of a category C and of a feature
set F is given by the following formula, which is inspired by
the contrast model of Tversky [Tversky, 1977; Smith, 1995;
Opwis and Plo¨ tzner, 1996],
|C ∩ F |</p>
      <p>|C|
Sim(C, F ) =
|C| is the number of the prototypical features that define
the category a. The present features are counted and
normalized so that the value can be compared. This is a very simple
form of representation. A binary vector in which the positions
represent different features can represent the set of features.
For each category a binary vector can be defined. Overlaps
between stored patterns correspond to overlaps between
categories.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>The Lernmatrix</title>
      <p>The Lernmatrix, also simply called “associative memory”
was developed by Steinbuch in 1958 as a biologically
inspired model from the effort to explain the psychological
phenomenon of conditioning [Steinbuch, 1961; 1971]. Later this
model was studied under biological and mathematical aspects
by Willshaw [Willshaw et al., 1969] and G. Palm [Palm,
1982; 1990].</p>
      <p>Associative memory is composed of a cluster of units.
Each unit represents a simple model of a real biological
neuron. The Lernmatrix was invented in by Steinbuch, whose
goal was to produce a network that could use a binary version
of Hebbian learning to form associations between pairs of
binary vectors, for example each one representing a cognitive
entity. Each unit is composed of binary weights, which
correspond to the synapses and dendrites in a real neuron. They are
described by wij ∈ {0, 1} in Figure 1. T is the threshold of
the unit. We call the Lernmatrix simply associative memory if
no confusion with other models is possible [Anderson, 1995a;
Ballard, 1997].</p>
      <p>The patterns, which are stored in the Lernmatrix, are
represented by binary vectors. The presence of a feature is
indicated by a ‘one’ component of the vector, its absence through
a ‘zero’ component of the vector. A pair of these vectors is
associated and this process of association is called learning.
The first of the two vectors is called the question vector and
the second, the answer vector. After learning, the question
vector is presented to the associative memory and the answer
vector is determined by the retrieval rule.
wmn
T
ym</p>
      <p>Learning In the initialization phase of the associative
memory, no information is stored. Because the information is
represented in weights, they are all initially set to zero. In the
learning phase, pairs of binary vector are associated. Let ~x be
the question vector and ~y the answer vector, the learning rule
is:
winjew
1
wiojld
if yi · xj = 1
otherwise.</p>
      <p>This rule is called the binary Hebbian rule [Palm, 1982].
Every time a pair of binary vectors is stored, this rule is used.
Retrieval In the one-step retrieval phase of the associative
memory, a fault tolerant answering mechanism recalls the
appropriate answer vector for a question vector ~x. For
the presented question vector ~x, the most similar learned
x~l question vector regarding the Hamming distance is
determined and the appropriate answer vector y~ is identified.
For the retrieval rule, the knowledge about the correlation
of the components is sufficient. The retrieval rule for the
determination of the answer vector ~y is:
yi =
where T is the threshold of the unit. The threshold is set as
proposed by [Palm et al., 1997] to the maximum of the sums
Pjn=1 wij xj :
n
T := max nX wij xj o.</p>
      <p>1≤i≤m
j=1</p>
      <p>Only the units that are maximal correlated with the
question vector are set to one.</p>
      <p>Storage capacity For an estimation of the asymptotic
number of vectorpairs (~x, ~y) which can be stored in an associative
memory before it begins to make mistakes in retrieval phase,
it is assumed that both vectors have the same dimension n.
It is also assumed that both vectors are composed of M 1s,
which are likely to be in any coordinate of the vector. In
this case it was shown [Palm, 1982; Hecht-Nielsen, 1989;
Sommer, 1993] that the optimum value for M is
approximately</p>
      <p>.</p>
      <p>M = log2(n/4)
(5)
and that approximately [Palm, 1982; Hecht-Nielsen, 1989]
.</p>
      <p>L = (ln 2)(n2/M 2)
(6)
of vector pairs can be stored in the associative memory. This
value is much greater then n if the optimal value for M is
used. In this case, the asymptotic storage capacity of the
Lernmatrix model is far better than those of other
associative memory models, namely 69.31%. This capacity can be
reached with the use of sparse coding, which is produced
when very small number of 1s is equally distributed over
the coordinates of the vectors [Palm, 1982; Stellmann, 1992].
For example an optimal code is defined as following; in the
vector of the dimension n=1000000 M=18 ones should be
used to code a pattern. The real storage capacity value is
lower when patterns are used which are not sparse or are
strongly correlated to other stored patterns. Usually
suboptimal sparse codes a sufficiently good to be used with the
associative memory. An example of a suboptimal sparse code
is the representation of words by context-sensitive letter units
[Wickelgren, 1969; 1977; Rumelhart and McClelland, 1986;
Bentz et al., 1989]. The ideas for the used robust
mechanism come from psychology and biology [Wickelgren, 1969;
1977; Rumelhart and McClelland, 1986; Bentz et al., 1989].
Each letter in a word is represented as a triple, which
consists of the letter itself, its predecessor, and its successor. For
example, six context-sensitive letters encode the word desert,
namely: de, des, ese, ser, ert, rt . The character “ ” marks
the word beginning and ending. Because the alphabet is
composed of 26+1 characters, 273 different context-sensitive
letters exist. In the 273 dimensional binary vector each position
corresponds to a possible context-sensitive letter, and a word
is represented by indication of the actually present
contextsensitive letters. We demonstrate the principle of sparse
coding by an example of the visual system and visual scene
representation.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Sparse features</title>
      <p>In hierarchical models of the visual system [Riesenhuber and
Poggio, 1999],[Fukushima, 1980], [Fukushima, 1989],
[Cardoso and Wichert, 2010] the neural units have a local view
unlike the common fully-connected networks. The receptive
fields of each neuron describe this local view. During the
categorization the network gradually reduces the information
from the input layer through the output layer. Integrating
local features into more global features does this. Supposed in
the lower layer tow cells recognize two categories at
neighboring position, and these two categories are integrated into a
more global category. The first cell is named α the second β.
The numerical code for α and β may represent the position of
each cell. A simple code would indicate if a cell is active or
not. One indicates active, zero not active. For c cells we could
indicate this information by a binary vector of dimension c.
For an image of size x × y a cell covers the image X times.
A binary vector that describes that image using the cell
representation has the dimension c × X . For example gray images
of the size 128 × 96 resulting in vectors of dimension 12288
can be covered with:
• 3072 masks M of the size of a size 2 × 2 resulting in a
binary vector that describes that image has the dimension
c1 × 3072, X1 = 3072 (see Figure 2 (a) ).
• 768 masks M of the size of a size 4 × 4 resulting in a
binary vector that describes that image has the dimension
c2 × 768, X2 = 768 (see Figure 2 (b) ).
• 192 masks M of the size of a size 8 × 8 resulting in a
binary vector that describes that image has the dimension
c3 × 192, X3 = 192 (see Figure 3 (a) ).
• 48 masks M of the size of a size 16 × 16 resulting in a
binary vector that describes that image has the dimension
c4 × 48, X4 = 48 (see Figure 3 (b) ).
A s s o c i a t iv e
m e m o r y</p>
      <p>The ideal c value for a sparse code is related to M
log2(n/4).</p>
      <p>X = log2(X · c/4)
2X = X · c/4
c =
4 · 2X</p>
      <p>X
The ideal value for c grows exponentially in relation to X .
Usually the used value for c is much lower then the ideal
value resulting in a suboptimal sparse code. The
representation of images by masks results in a suboptimal code. The
optimal code is approached with the size of masks, the bigger
the mask, the smaller the value of X . The number of pixels
inside a mask grows quadratic. A bigger masks implies the
ability to represent more distinct categories, which implies a
bigger c.</p>
      <p>An ideal value for c is possible, if the value for X &lt;&lt; 100.
Instead of covering an image by masks, we indicate the
present objects. Objects and their position in the visual field
can represent a visual scene. A sub-vector of the vector
representing the visual scene represents each object. For
example, if there is a total of 10 objects, the c value is 409. To
represent 409 different categories of objects at different
positions resulting in 4090 dimensional binary vector. This vector
409!
could represent (409−20)! different visual states of the world.
The storage capacity of the associative memory in this case
would be around 159500 patterns, which is 28 times bigger
as the number of the units (4090).
2.4</p>
    </sec>
    <sec id="sec-5">
      <title>Problem Solving</title>
      <p>Human problem solving can be described by a
problembehavior graph constructed from a protocol of the person
talking aloud, mentioning considered moves and aspects of
the situation. According to the resulting theory, searching
whose state includes the initial situation and the desired
situation in a problem space [Newell, 1990; ?] solves problems.
This process can be described by the production system
theory. The production system in the context of classical
Artificial Intelligence and Cognitive Psychology is one of the
most successful computer models of human problem
solving. The production system theory describes how to form a
sequence of actions, which lead to a goal, and offers a
computational theory of how humans solve problems [Anderson,
1995b]. Production systems are composed of if-then rules
that are also called productions. A rule [contains several if
patterns and one or more then patterns. A pattern in the
context of rules is an individual predicate, which can be negated
together with arguments. A rule can establish a new
assertion by the then part (its conclusion) whenever the if part (its
premise) is true. One of the best-known cognitive models,
based on the production system, is SOAR. The SOAR state,
operator and result model was developed to explain human
problem-solving behavior [Newell, 1990]. It is a hierarchical
production system in which the conflict-resolution strategy
is treated as another problem to be solved. All satisfied
instances of rules are executed in parallel in a temporary mode.
After the temporary execution, the best rule is chosen to take
action. The decision takes place in the context of a stack of
.
=
(7)
earlier decisions. Those decisions are rated utilizing
preferences and added to the stack by chosen rules. Preferences
are determined together with the rules by an observer using
knowledge about a problem.</p>
      <p>According to the production system theory, we can define
a geometrically based problem-solving model as a
production system operating on vectors of fixed dimensions. Instead
of rules, we use associations and vectors represent the states.
Our goal is to form a sequence of associations, which lead to a
desired state represented by a vector, from an initial state
represented by a vector. Each association changes some parts of
the vector. In each state, several possible associations can be
executed, but only one has to be chosen. Otherwise, conflicts
in the representation of the state would occur. To perform
these operations, we divided a vector representing a state into
sub-vectors. An association recognizes some sub-vectors in
the vector and exchanges them for different sub-vectors. It is
composed of a precondition of fixed arranged m sub-vectors
and a conclusion. Suppose a vector is divided into n
subvectors with n &gt; m. An association recognizes m different
sub-vectors and exchanges them for different m sub-vectors.
To recognize m sub-vectors out of n sub-vectors we perform
a permutation p(n,m) and verify if each permutation
corresponds to a valid precondition of an association. For
example, if there is a total of 7 elements and we are selecting a
sequence of three elements from this set, then the first
selection is one from 7 elements, the next one from the remaining
6, and finally from the remaining 5, resulting in 7 * 6 * 5 =
210, see Figure 4.</p>
      <p>Out of several possible associations, we chose the one,
which modifies the state in such a way that it becomes more
similar to the desired state according to the Equation 1. The
desired state corresponds to the category of Equation 1, each
feature represents a possible state. The states are represented
by sparse features. With the aid of this heuristic hill
climbing is performed. Each element represents an object. Objects
are represented by some dimensions of the space and form a
sub-space by themselves, see Figure 4.</p>
      <p>The computation can be improved by a simple and
universal heuristics function, which takes into account the
relationship between the vector and the corresponding
similarity of the represented states see Figure 5 and Figure 6. The
heuristics function makes a simple assumption that the
distance between the states in the problem space is related to the
similarity of the vectors representing the states.</p>
      <p>The similarity between the corresponding vectors can
indicate the distance between the sub-symbols representing the
state. Empirical experiments in popular problem-solving
domains of Artificial Intelligence, like robot in a maze, block
world or 8-puzzle indicated that the distance between the
states in the problem space is actually related to the similarity
between the images representing the states [Wichert, 2001;
Wichert et al., 2008; Wichert, 2009].
3</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>Living organisms experience the world as a simple. The
actual perception of the world and manipulation in the world
by living organisms lead to the invention or recreation of an
experience that, at least in some respects, resembles the
experience of actually perceiving and manipulating objects in the
absence of direct sensory stimulation. This kind of
representation is called sub-symbolic. Sub-symbolic representation
implies heuristic functions. The assumption that the distance
between states in the problem space is related to the
similarity between the sub-symbols representing the states is only
valid in simple cases. However, simple cases represent the
majority of exiting problems in domain. Sense organs sense
the world by receptors which a part of the sensory system
and the nervous system. Sparse binary vectors can describe
sub-symbolic representation, which can be efficiently stored
in associative memories. A simple code would indicate if a
receptor is active or not. One indicates active, zero not active.
For c receptors we could indicate this information by a binary
vector of dimension c with only one ”1”, the bigger the c, the
sparser the code. For receptors in X positions the sparse code
results in c × X dimensional vector with X ones.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work was supported by Fundao para a Cencia e
Tecnologia (FCT) (INESC-ID multiannual funding) through the
PIDDAC Program funds.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>[Anderson</surname>
          </string-name>
          , 1995a]
          <string-name>
            <surname>James</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Anderson</surname>
          </string-name>
          .
          <article-title>An Introduction to Neural Networks</article-title>
          . The MIT Press,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>[Anderson</surname>
          </string-name>
          , 1995b]
          <string-name>
            <surname>John R. Anderson</surname>
            . Cognitive Psychology and
            <given-names>its Implications. W. H.</given-names>
          </string-name>
          <string-name>
            <surname>Freeman</surname>
          </string-name>
          and Company,
          <source>fourth edition</source>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Ballard</source>
          , 1997]
          <string-name>
            <surname>Dana</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Ballard</surname>
          </string-name>
          .
          <article-title>An Introduction to Natural Computation</article-title>
          . The MIT Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Bentz et al.,
          <year>1989</year>
          ]
          <string-name>
            <given-names>Hans J.</given-names>
            <surname>Bentz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Hagstroem</surname>
          </string-name>
          , and Gu¨nther Palm.
          <article-title>Information storage and effective data retrieval in sparse matrices</article-title>
          .
          <source>Neural Networks</source>
          ,
          <volume>2</volume>
          (
          <issue>4</issue>
          ):
          <fpage>289</fpage>
          -
          <lpage>293</lpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Cardoso and Wichert</source>
          , 2010]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          and
          <string-name>
            <given-names>A</given-names>
            <surname>Wichert</surname>
          </string-name>
          .
          <article-title>Neocognitron and the map transformation cascade</article-title>
          .
          <source>Neural Networks</source>
          ,
          <volume>23</volume>
          (
          <issue>1</issue>
          ):
          <fpage>74</fpage>
          -
          <lpage>88</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Fukushima</source>
          , 1980]
          <string-name>
            <given-names>K.</given-names>
            <surname>Fukushima</surname>
          </string-name>
          .
          <article-title>Neocognitron: a self organizing neural network model for a mechanism of pattern recognition unaffected by shift in position</article-title>
          .
          <source>Biol Cybern</source>
          ,
          <volume>36</volume>
          (
          <issue>4</issue>
          ):
          <fpage>193</fpage>
          -
          <lpage>202</lpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Fukushima</source>
          , 1989]
          <string-name>
            <given-names>K.</given-names>
            <surname>Fukushima</surname>
          </string-name>
          .
          <article-title>Analisys of the process of visual pattern recognition by the neocognitron</article-title>
          .
          <source>Neural Networks</source>
          ,
          <volume>2</volume>
          :
          <fpage>413</fpage>
          -
          <lpage>420</lpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Gilovich</source>
          , 1999]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Gilovich</surname>
          </string-name>
          .
          <source>Tversky. In The MIT Encyclopedia of the Cognitive Sciences</source>
          , pages
          <fpage>849</fpage>
          -
          <lpage>850</lpage>
          . The MIT Press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Goldstone</source>
          , 1999]
          <string-name>
            <given-names>Robert</given-names>
            <surname>Goldstone</surname>
          </string-name>
          .
          <article-title>Similarity. In The MIT Encyclopedia of the Cognitive Sciences</article-title>
          , pages
          <fpage>763</fpage>
          -
          <lpage>765</lpage>
          . The MIT Press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [
          <string-name>
            <surname>Hecht-Nielsen</surname>
          </string-name>
          ,
          <year>1989</year>
          ]
          <string-name>
            <given-names>Robert</given-names>
            <surname>Hecht-Nielsen</surname>
          </string-name>
          .
          <source>Neurocomputing. Addison-Wesley</source>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Kosslyn</source>
          , 1994]
          <string-name>
            <surname>Stephen</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Kosslyn</surname>
          </string-name>
          .
          <article-title>Image and Brain, The Resolution of the Imagery Debate</article-title>
          . The MIT Press,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>[Lancy</source>
          , 1983]
          <string-name>
            <given-names>D.F.</given-names>
            <surname>Lancy</surname>
          </string-name>
          .
          <article-title>Cross-Cultural Studies in Cognition and Mathematics</article-title>
          . Academic Press, New York,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[McClelland and Rumelhart</source>
          , 1985]
          <string-name>
            <given-names>J.L.</given-names>
            <surname>McClelland</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.E.</given-names>
            <surname>Rumelhart</surname>
          </string-name>
          .
          <article-title>Distributed memory and the representation of general and specific memory</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          ,
          <volume>114</volume>
          :
          <fpage>159</fpage>
          -
          <lpage>188</lpage>
          ,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>[Newell and Simon</source>
          , 1976]
          <string-name>
            <given-names>A.</given-names>
            <surname>Newell</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Simon</surname>
          </string-name>
          .
          <article-title>Computer science as empirical inquiry: symbols and search</article-title>
          .
          <source>Communication of the ACM</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>113</fpage>
          -
          <lpage>126</lpage>
          ,
          <year>1976</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>[Newell</source>
          , 1990]
          <string-name>
            <given-names>Allen</given-names>
            <surname>Newell</surname>
          </string-name>
          . Unified Theories of Cognition. Harvard University Press,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Opwis and Plo¨ tzner, 1996]
          <article-title>Klaus Opwis and Rolf Plo¨ tzner</article-title>
          . Kognitive Psychologie mit dem Computer. Spektrum Akademischer Verlag, Heidelberg Berlin Oxford,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>[Osherson</source>
          , 1987]
          <string-name>
            <given-names>Daniel N.</given-names>
            <surname>Osherson</surname>
          </string-name>
          .
          <article-title>New axioms for the contrast model of similarity</article-title>
          .
          <source>Journal of Mathematical Psychology</source>
          ,
          <volume>31</volume>
          :
          <fpage>93</fpage>
          -
          <lpage>103</lpage>
          ,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Osherson</source>
          , 1995]
          <string-name>
            <given-names>Daniel N.</given-names>
            <surname>Osherson</surname>
          </string-name>
          .
          <article-title>Probability judgment</article-title>
          . In Edward E. Smith and Daniel N. Osherson, editors,
          <source>Thinking</source>
          , volume
          <volume>3</volume>
          ,
          <string-name>
            <surname>chapter</surname>
            <given-names>two</given-names>
          </string-name>
          , pages
          <fpage>35</fpage>
          -
          <lpage>75</lpage>
          . MIT Press, second edition,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Palm et al.,
          <year>1997</year>
          ]
          <string-name>
            <given-names>G.</given-names>
            <surname>Palm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schwenker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.T.</given-names>
            <surname>Sommer</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Strey</surname>
          </string-name>
          .
          <article-title>Neural associative memories</article-title>
          . In A. Krikelis and C. Weems, editors,
          <source>Associative Processing and Processors</source>
          , pages
          <fpage>307</fpage>
          -
          <lpage>325</lpage>
          . IEEE Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Palm</source>
          , 1982]
          <article-title>Gu¨ nther Palm</article-title>
          .
          <source>Neural Assemblies, an Alternative Approach to Artificial Intelligence</source>
          . Springer-Verlag,
          <year>1982</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>[Palm</source>
          , 1990]
          <article-title>Gu¨ nther Palm. Assoziatives Geda¨chtnis und Gehirntheorie</article-title>
          .
          <source>In Gehirn und Kognition</source>
          , pages
          <fpage>164</fpage>
          -
          <lpage>174</lpage>
          . Spektrum der Wissenschaft,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>[Riesenhuber and Poggio</source>
          , 1999]
          <string-name>
            <given-names>M.</given-names>
            <surname>Riesenhuber</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Poggio</surname>
          </string-name>
          .
          <article-title>Hierarchical models of object recognition in cortex</article-title>
          .
          <source>Nature Neuroscience</source>
          ,
          <volume>2</volume>
          :
          <fpage>1019</fpage>
          -
          <lpage>1025</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Rumelhart and McClelland</source>
          , 1986]
          <string-name>
            <surname>D.E.</surname>
          </string-name>
          <article-title>Rumelhart and McClelland. On learning the past tense of english verbs</article-title>
          . In J.L.
          <string-name>
            <surname>McClelland</surname>
            and
            <given-names>D.E</given-names>
          </string-name>
          . Rumelhart, editors,
          <source>Parallel Distributed Processing</source>
          , pages
          <fpage>216</fpage>
          -
          <lpage>271</lpage>
          . MIT Press,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>[Simon</source>
          , 1991]
          <article-title>Herbert A. Simon. Models of my Life</article-title>
          . Basic Books, New York,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>[Smith</source>
          ,
          <year>1995</year>
          ]
          <string-name>
            <given-names>Edward E.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Concepts and categorization</article-title>
          . In Edward E. Smith and Daniel N. Osherson, editors,
          <source>Thinking</source>
          , volume
          <volume>3</volume>
          ,
          <string-name>
            <surname>chapter</surname>
            <given-names>one</given-names>
          </string-name>
          , pages
          <fpage>3</fpage>
          -
          <lpage>33</lpage>
          . MIT Press, second edition,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>[Sommer</source>
          , 1993]
          <string-name>
            <surname>Friedrich</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Sommer</surname>
          </string-name>
          .
          <article-title>Theorie neuronaler Assoziativspeicher</article-title>
          .
          <source>PhD thesis</source>
          , Heinrich-HeineUniversita¨t Du¨ sseldorf, Du¨ sseldorf,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <source>[Steinbuch</source>
          , 1961]
          <string-name>
            <given-names>K.</given-names>
            <surname>Steinbuch</surname>
          </string-name>
          .
          <source>Die Lernmatrix. Kybernetik</source>
          ,
          <volume>1</volume>
          :
          <fpage>36</fpage>
          -
          <lpage>45</lpage>
          ,
          <year>1961</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>[Steinbuch</source>
          , 1971]
          <string-name>
            <given-names>Karl</given-names>
            <surname>Steinbuch</surname>
          </string-name>
          .
          <source>Automat und Mensch</source>
          . Springer-Verlag,
          <source>fourth edition</source>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>[Stellmann</source>
          , 1992]
          <string-name>
            <given-names>Uli</given-names>
            <surname>Stellmann</surname>
          </string-name>
          .
          <article-title>A¨ hnlichkeitserhaltende Codierung</article-title>
          .
          <source>PhD thesis</source>
          , Universita¨t Ulm, Ulm,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>[Sun</source>
          , 1995]
          <string-name>
            <given-names>Ron</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>A two-level hybrid architecture for structuring knowledge for commonsense reasoning</article-title>
          . In Ron Sun and Lawrence A. Bookman, editors,
          <source>Computational Architectures Integrating Neural and Symbolic Processing, chapter 8</source>
          , pages
          <fpage>247</fpage>
          -
          <lpage>182</lpage>
          . Kluwer Academic Publishers,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>[Tarski</source>
          , 1956]
          <string-name>
            <given-names>Alfred</given-names>
            <surname>Tarski</surname>
          </string-name>
          . Logic, Semantics,Metamathematics. Oxford University Press, London,
          <year>1956</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <source>[Tversky</source>
          ,
          <year>1977</year>
          ]
          <string-name>
            <given-names>A.</given-names>
            <surname>Tversky</surname>
          </string-name>
          .
          <article-title>Feature of similarity</article-title>
          .
          <source>Psychological Review</source>
          ,
          <volume>84</volume>
          :
          <fpage>327</fpage>
          -
          <lpage>352</lpage>
          ,
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [Wichert et al.,
          <year>2008</year>
          ]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wichert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Pereira</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Carreira</surname>
          </string-name>
          .
          <article-title>Visual search light model for mental problem solving</article-title>
          .
          <source>Neurocomputing</source>
          ,
          <volume>71</volume>
          (
          <fpage>13</fpage>
          -15):
          <fpage>2806</fpage>
          -
          <lpage>2822</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>[Wichert</source>
          , 2001]
          <string-name>
            <given-names>Andrzej</given-names>
            <surname>Wichert</surname>
          </string-name>
          .
          <article-title>Pictorial reasoning with cell assemblies</article-title>
          .
          <source>Connection Science</source>
          ,
          <volume>13</volume>
          (
          <issue>1</issue>
          ),
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <source>[Wichert</source>
          , 2009]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Wichert</surname>
          </string-name>
          .
          <article-title>Sub-symbols and icons</article-title>
          .
          <source>Cognitive Computation</source>
          ,
          <volume>1</volume>
          (
          <issue>4</issue>
          ):
          <fpage>342</fpage>
          -
          <lpage>347</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>[Wickelgren</source>
          , 1969]
          <article-title>Wayne A. Wickelgren. Contextsensitive coding, associative memory, and serial order in (speech)behavior</article-title>
          .
          <source>Psychological Review</source>
          ,
          <volume>76</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <source>[Wickelgren</source>
          , 1977]
          <string-name>
            <surname>Wayne</surname>
            <given-names>A. Wickelgren. Cognitive</given-names>
          </string-name>
          <string-name>
            <surname>Psychology</surname>
          </string-name>
          . Prentice-Hall,
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [Willshaw et al.,
          <year>1969</year>
          ]
          <string-name>
            <given-names>D.J.</given-names>
            <surname>Willshaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.P.</given-names>
            <surname>Buneman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.C.</given-names>
            <surname>Longuet-Higgins</surname>
          </string-name>
          .
          <article-title>Nonholgraphic associative memory</article-title>
          .
          <source>Nature</source>
          ,
          <volume>222</volume>
          :
          <fpage>960</fpage>
          -
          <lpage>962</lpage>
          ,
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>