<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>D. Billman, E. Heit, Observational Learning from Internal Feedback: A Simulation of
an Adaptive Learning Method, Cognitive Science</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1007/978-3-030-12800-5_3</article-id>
      <title-group>
        <article-title>Towards Conceptual Logic Tensor Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lucas Bechberger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Cognitive Science, Osnabrück University</institution>
          ,
          <addr-line>Wachsbleiche 27, 49090 Osnabrück</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>12</volume>
      <issue>1988</issue>
      <fpage>11</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>The symbol grounding problem refers to the necessity of grounding abstract symbolic knowledge (such as encoded in formal ontologies) in the real world through perception and action. The cognitive framework of conceptual spaces provides a potential way for solving the symbol grounding problem by proposing an intermediate representational layer: Concepts are represented by regions in low-dimensional similarity spaces, which are in turn grounded in subsymbolic processing. Logic tensor networks provide a general mechanism for learning membership functions in the presence of both bottom-up information (i.e., training examples) and top-down constraints in the form of logical rules. In this paper, we propose to combine logic tensor networks with conceptual spaces in order to ground predicates from the symbolic layer in conceptual regions while taking into account logical constraints from abstract background knowledge. We discuss several potential membership functions for concepts and argue that this approach can be used to provide a cognitive grounding for formal ontologies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Formal ontologies provide us with ways of encoding knowledge in a logical format. Large-scale
applications such as Google’s knowledge graph1 illustrate the usefulness of such a structured
representation for encoding entities, classes, and their respective relations. However, ontologies
as used in current technical systems sufer from the symbol grounding problem [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]: The
symbols they contain are not directly linked to the real world, but are usually defined based on
other symbols. Research in knowledge graph embeddings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], where entities are identified with
points in a feature space, provide only a partial solution to this problem, since the dimensions
of this feature space are usually not tied to perception and action.
      </p>
      <p>
        Deep neural networks have become the predominant approach in many machine learning
tasks, including the areas of computer vision [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and natural language processing [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. They
are able to extract compact representations from raw perceptual input without the need for
extensive manual feature engineering. However, their recent successes have been accompanied
with an urge for more human-like, explainable AI [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], since their overall input-output mapping
is opaque and cannot be easily analyzed or interpreted by human experts.
      </p>
      <p>
        Neural-symbolic integration [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] ofers the possibility to combine the interpretability of
symbolic systems with the learning capabilities of artificial neural networks. Also the area of
cognitive AI [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], i.e., intelligent systems inspired by findings from cognitive psychology, can
help to align artificial systems more closely with human cognition.
      </p>
      <p>
        The cognitive framework of conceptual spaces [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] unifies aspects of both the neural-symbolic
tradition and cognitive AI. It proposes a geometric representation of conceptual knowledge
based on psychological similarity spaces and ofers an intermediate level of representation
between the connectionist and the symbolic approach. The individual dimensions spanning
such a conceptual space correspond to cognitively meaningful features of the inputs (such as
hue, saturation, and brightness for colors). Concepts (such as the color blue) can then be defined
as convex regions in this space. Conceptual spaces thus provide an indirect way of grounding
symbolic descriptions in perception. They have seen a wide variety of applications in artificial
intelligence, linguistics, psychology, and philosophy [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
      </p>
      <p>
        Logic tensor networks [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ] (LTNs) are a type of neural network which uses fuzzy
membership functions in order to ground symbolic predicates in a given embedding space. When
optimizing the parameters of these membership functions, LTNs can take into account both
bottom-up information in the form of training examples and top-down constraints in the form
of general logical rules. For instance, conceptual hierarchies from an ontology can be used to
enforce a subsethood relation between the respective membership functions.
      </p>
      <p>
        In this paper, we propose to use logic tensor networks to ground ontologies in conceptual
spaces. Since conceptual spaces are based on meaningful dimensions, this aids the
interpretability of the resulting embedding. Conceptual spaces can be grounded in psychological dissimilarity
ratings [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ] or the features extracted by deep neural networks from raw perceptual inputs.
Therefore, a successful grounding of a given ontology in a conceptual space indirectly solves the
symbol grounding problem in a cognitively plausible way. Moreover, by explicitly considering
relations between classes, logic tensor networks can help to leverage background knowledge in
order to learn conceptual regions.
      </p>
      <p>The remainder of this paper is structured as follows: In Section 2, we describe both conceptual
spaces and logic tensor networks in more detail. In Section 3, we then argue why a combination
of these two frameworks seems promising and discuss possible implementations of diferent
membership functions. Section 4 then concludes this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>In the following, we will first give a general overview of the conceptual spaces framework
(Section 2.1), before introducing logic tensor networks (Section 2.2).</p>
      <sec id="sec-2-1">
        <title>2.1. Conceptual Spaces</title>
        <p>
          A conceptual space as proposed by Gärdenfors [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is a similarity space spanned by a small
number of interpretable, cognitively relevant quality dimensions (e.g., temperature, time, hue,
pitch). One can measure the distance between two observations with respect to each of these
dimensions and aggregate them into a global notion of semantic distance. Semantic similarity
is then defined as an exponentially decaying function of distance, i.e., (, ) = − · (,)
with a sensitivity parameter  &gt; 0.
        </p>
        <p>
          The overall conceptual space can be structured into so-called domains, which represent, for
example, diferent perceptual modalities such as color, shape, taste, and sound. The color domain,
for instance, can be represented by the three dimensions hue, saturation, and brightness, while
the sound domain is spanned by the dimensions pitch and loudness. Based on psychological
evidence [
          <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
          ], distance within a domain is measured with the Euclidean metric, while the
Manhattan metric is used to aggregate distances across domains.
        </p>
        <p>Gärdenfors defines properties like red, round, and sweet as convex regions within a single
domain (namely, color, shape, and taste, respectively). A property thus corresponds to a set of
observations from a single perceptual modality. Concept hierarchies are an emergent property
of this spatial representation: If the sky blue region is a subset of the blue region, this implicitly
encodes that sky blue is a special shade of blue. Based on properties, Gärdenfors now defines
full-fleshed concepts like apple or dog by using one convex region per domain, a set of salience
weights (which represent the relevance of the given domain to the given concept), and
information about cross-domain correlations. The apple concept may thus be represented by the
regions red, sweet, and round in the domains of color, taste, and shape, respectively.</p>
        <p>There are in principle three ways of obtaining the dimensions of a conceptual space [9,
Sections 1.7, 1.9, and 6.5]: Firstly, if the domain of interest is well understood, one can manually
define the dimensions of the similarity space.</p>
        <p>A second approach is based on machine learning algorithms for dimensionality reduction.
For instance, unsupervised artificial neural networks (ANNs) such as autoencoders or
selforganizing maps can be used to find a compressed representation for a given set of input stimuli.
This task is however solved by minimizing a mathematical error function which seems to be
not satisfactory from a psychological point of view.</p>
        <p>
          A third popular way of obtaining a conceptual similarity space is based on psychological
dissimilarity ratings. These dissimilarity ratings are collected for a fixed set of stimuli in a
psychological experiment. They are then converted into an -dimensional geometric
representation of the stimulus set by using a technique called “multidimensional scaling” (MDS),
which ensures that geometric distances between pairs of stimuli reflect their psychological
dissimilarity [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. While the similarity spaces produced by MDS are grounded in psychological
experiments, they do not readily generalize to unseen stimuli [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>
          Recently, a hybrid approach has been proposed [
          <xref ref-type="bibr" rid="ref19 ref20 ref21 ref22">19, 20, 21, 22</xref>
          ], where MDS is used to initialize
the similarity space and ANNs are then trained to generalize the mapping to novel inputs.
        </p>
        <p>
          Gärdenfors argues that the convexity requirement relates conceptual spaces to the prototype
theory of concepts [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], which assumes that concept membership is based on similarity to
a prototype. This can explain why some members of a category are deemed to be more
typical than others. Gärdenfors [9, Section 3.8] now argues that if concepts are represented
by convex regions, a prototype can be obtained by computing the center of gravity for the
conceptual region. Conversely, Gärdenfors [9, Section 3.9] shows that by assuming a
prototypebased representation, one can easily generate convex regions. For instance, if color properties
such as red and orange are represented by their prototypical points in color space (e.g., their
corresponding focal colors), one can partition the overall space into convex regions by assigning
each point in the space to its closest prototype. This way of partitioning a space is called a
Voronoi tessellation and will be discussed in more detail in Section 3.2.
        </p>
        <p>Since conceptual spaces can be interpreted as an intermediate layer of representation between
the traditional symbolic and subsymbolic layers, they can also help to solve the symbol grounding
problem: Individual observations, which correspond to high-dimensional activation vectors in
the subsymbolic layer, are represented by points in the lower-dimensional conceptual space
and can be mapped onto constants and variables from the symbolic layer. Predicates from
the symbolic layer (such as apple and red) can be mapped onto concepts and properties in the
conceptual layer. The symbols from the symbolic layer can therefore be indirectly grounded in
subsymbolic perception through the conceptual layer.</p>
        <p>Conceptual spaces in their original formulation focus mostly on concepts that can be defined
based on perceptual properties. Relations between objects can be represented using product
spaces, for instance by defining longerThan as a convex region in R+ × R+ (where each
dimension represents the length of one individual object) [9, Section 3.10.1]. Relational concepts
like robber or seat can be represented based on their respective roles in events like robbing and
sitting (which involve agent, patient, theme, action, and result) [24, Sections 6.7 and 8.4].</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Logic Tensor Networks</title>
        <p>
          Logic Tensor Networks (LTNs) [
          <xref ref-type="bibr" rid="ref12 ref13 ref25 ref26">12, 13, 25, 26</xref>
          ] provide a principled way of using neural
computations to connect feature spaces with symbolic rules through fuzzy logic.2
        </p>
        <p>
          Logic Tensor Networks integrate knowledge representation, learning, and reasoning using
a diferentiable fuzzy first-order logic language called “Real Logic”. Real Logic is a first-order
language containing constant symbols (representing individual observations such as Bob or
Paris), function symbols (representing mappings between observations, e.g., homeTownOf ),
predicate symbols (representing concepts and relations such as lawyer and livesIn), and variable
symbols (representing lists of observations) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. All of these language constituents are typed
with respect to a set of domains : We can require that Bob belongs to the domain of people and
Paris to the domain of cities, while the function homeTownOf takes only people as input and
returns cities. The individual parts of the language can now be combined into formulas such as
lawyer(Bob), or ∀ : (lawyer() → homeTownOf() = Paris). These formulas are constructed
using logical connectives (such as →) and quantifiers (such as ∀) and have a fuzzy degree of
truth in the interval [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ].
        </p>
        <p>
          In order to relate the semantics of the logical language to actual data points, Real Logic
makes use of a so-called grounding function  , which maps terms (i.e., constants, variables,
and results of function applications) onto points in a feature space, and both functions and
predicates onto neural networks. The networks implementing predicates are required to return
a value from the interval [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ] and can thus be interpreted as defining a fuzzy membership
function of the respective concept or relation in the given feature space. Relations such as
2See https://github.com/logictensornetworks/logictensornetworks for the implementation of this framework.
livesIn are implemented as fuzzy regions in a product space of domains (in this case people and
cities). Overall, the grounding associates any formula expressible in the language with a real
number in [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ], representing its degree of truth.
        </p>
        <p>
          This is done as follows (see Figure 1): First, all terms (such as Bob) are grounded into vectors,
and all function and predicate symbols (such as homeTownOf an livesIn) are grounded into
their respective neural networks. The structure of the symbolic formula then determines which
neural network is applied to which feature vector. If predicates such as married(, ) are applied
to variables such as  = (Bob, John) and  = (Alice, Mary, Susan), the resulting grounding is a
matrix containing the degree of truth for each possible combination of observations [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>
          In order to ground logical connectives such as ∧, ∨, ¬, and →, the corresponding operators
from fuzzy logic are used. Badreddine et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] note that many standard operators from fuzzy
logic are not well-suited for the context of neural networks, since they may cause vanishing
or exploding gradients. They recommend using the product norm (, ) =  ·  for
implementing the conjunction and its complement (, ) =  +  −  ·  for the disjunction.
For the negation,  () = 1 −  is used.
        </p>
        <p>
          Since Real Logic is a first-order language, it also needs to provide a grounding for the universal
and the existential quantifier. Mapping ∀ to the minimum over all entries of the variable  may
be a straightforward choice, but does not tolerate exceptions and may thus not be suitable for
real-life applications, if one assumes a certain amount of noise both in the background
knowledge and in the empirical data (e.g., incorrect labels) [
          <xref ref-type="bibr" rid="ref26 ref27">26, 27</xref>
          ]. Instead, the generalized mean
1
(1, . . . , ) = ︁( 1 ∑︀=1 )︁  is used, where the parameter  controls the “strictness”
of the aggregator.3 The current version of the framework [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] proposes to use diferent variants
of the generalized mean for grounding both the universal and the existential quantifier. This
of course breaks the duality between the existential and the universal quantifier, but seems
to be necessary to enable robust gradient-based learning. One could envision to use both an
exception-tolerant and a rigid classical version for both quantifiers. One would then however
need to specify which version to apply in which contexts.
        </p>
        <p>3Note that 1 corresponds to the standard arithmetic mean and − 1 to the harmonic mean. If  =
(1, . . . , ) is a diference vector, then () is equivalent to a Minkowski metric.</p>
        <p>
          Satisfiability specifies the degree to which a grounding  satisfies a given set  of formulas
by simply aggregating the truth values of all formulas  ∈  [
          <xref ref-type="bibr" rid="ref12 ref25 ref26">12, 25, 26</xref>
          ]. Donadello et al. [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]
propose to use the generalized mean also for this purpose, since using a conjunction over the
formulas can lead to undesired behavior in gradient-based optimization.
        </p>
        <p>
          The knowledge represented in logic tensor networks consists of both the formulas  in the
logical language (corresponding to symbolic top-down information) and the grounding 
obtained from observations (corresponding to subsymbolic bottom-up information) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. One
can encode diferent types of constraints into the system: For instance, one can explicitly fix the
grounding for some of the symbols (e.g., mapping a given constant to a concrete feature vector).
Also a parametric definition of predicates is possible by specifying the structure of the respective
neural network, but leaving its exact parameter settings undetermined. Moreover, diferent
kinds of formulas can be used to constrain the system: Factual propositions such as lawyer(Bob)
encode facts about individual constants (which corresponds to providing labels for training
examples), while generalized propositions such as ∀ : (lawyer() → homeTownOf() = Paris)
allow to specify data-independent general background knowledge.
        </p>
        <p>
          Learning in logic tensor networks takes place through gradient descent on the parameter
values  of the grounding  in order to maximize the satisfiability of the overall set of formulas
 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. In practice, maximizing satisfiability may need to be accompanied by a regularization
term on the parameters  to prevent overfitting [
          <xref ref-type="bibr" rid="ref13 ref26">13, 26</xref>
          ]. Once a grounding has been established,
it can be used for answering concrete queries about the truth value of a given formula or about
the embedding of a given term.
        </p>
        <p>
          Since LTNs unify subsymbolic machine learning aspects with symbolic logical constraints,
they can be applied to a variety of problems. Badreddine et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] have given a principled
overview of diferent tasks that can be solved with LTNs. These include classification, regression,
clustering, semi-supervised pattern recognition, embedding learning, and knowledge base
completion. LTNs have also been used to learn transitive predicates (such as hyponymOf )
for simple ontologies based only on one-hop examples [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Moverover, Bianchi et al. [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]
have recently illustrated the capability of LTNs to connect pre-trained entity embeddings with
a subset of the DBpedia ontology [30]. Other practical applications include semantic image
interpretation [
          <xref ref-type="bibr" rid="ref26 ref27">26, 27, 31</xref>
          ], incorporating fairness constraints into deep neural networks [32],
and supplementing reinforcement learning algorithms with semantic knowledge [33].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Towards a Fruitful Combination</title>
      <p>Logic tensor networks ofer the possibility to combine bottom-up information in the form of
training examples with top-down information in the form of general rules. They thus make an
ideal candidate for closing the gap between the conceptual and the symbolic layer.</p>
      <p>In Section 3.1, we show how logic tensor networks can reflect the general properties of the
conceptual spaces framework. Afterwards, we investigate diferent membership functions,
using a distinction into partitional (Section 3.2) and nonpartitional (Section 3.3) approaches.</p>
      <sec id="sec-3-1">
        <title>3.1. General Considerations</title>
        <p>
          The knowledge approach to concepts from psychology [34, Chapter 6] emphasizes the crucial
role of world knowledge in the learning and application of concepts and can thus be linked
to both formal ontologies and the influence of logical rules on the learning process in LTNs.
Moreover, LTNs take into account information about points in a feature space, and the grounding
of predicates usually gives rise to a membership function with one or more receptive fields.
LTNs can therefore also be related to the prototype [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] and exemplar theories [35] of concepts.
        </p>
        <p>These observations and the combination of bottom-up and top-down information make LTNs
quite interesting from the perspective of conceptual spaces: If we use a conceptual space as
a feature space, then the LTN can implement a two-way interaction between the conceptual
and the symbolic layer. Just as with the conceptual spaces framework, observations can be
represented by points and concepts can be represented as regions in the feature space. Moreover,
LTNs are able to encode the domain structure of a conceptual space and they use a similar way
of encoding simple relations as regions in a product space. Finally, the usage of fuzzy sets and
fuzzy logic allows us to represent vague conceptual boundaries.</p>
        <p>How exactly can we apply LTNs to conceptual spaces? Both properties and concepts can
be represented by predicates with a convex membership function. While properties refer only
to a single domain, concepts are defined on a concatenation of domains. We can require that
the predicate red() is defined on the three-dimensional color domain, while the predicate
apple( = (, , )) involves the domains of color, taste, and shape. Individual observations
can then be represented by one point per domain. Function symbols could potentially be used to
represent actions and changes: Applying a function symbol like lift could for instance translate
into a simple vector addition in the location domain that increases the altitude of the given
object. Finally, basic relations are defined by considering regions in product spaces of multiple
domains. In order to represent more complex relational knowledge, one could try to use the
event structure proposed by Gärdenfors [24, Chapter 9].</p>
        <p>An advantage of using logic tensor networks for grounding ontologies in conceptual spaces is
their large variety of inference and learning methods. In addition to learning conceptual regions
from observations, they can for instance also generate an embedding of an unobserved object
based on a symbolic description. For example, “object  is a red apple” can be translated into a
point in the conceptual space by maximizing the satisfiability of apple() ∧ red(). This spatial
representation of object  can then in turn be used to make further inferences, for instance
about the taste domain (e.g., by evaluating sweet()). Thus, the geometric embedding can give
rise to common-sense inferences not easily realizable within the symbolic layer.</p>
        <p>Moreover, LTNs do not only provide an embedding of entities and classes, but they are also
able to enforce the validity of general rules, which may reduce the required number of training
examples. This makes them especially attractive for bridging the conceptual and the symbolic
layer, since they can harness the whole expressivity of formal ontologies in order to guide the
machine learning process. This also related to embodied and enactivist approaches to cognition
[36], which assume that top-down information strongly influences bottom-up perception and
conceptualization of the environment. One may furthermore speculate that the enforcement of
general logical rules can help to prevent catastrophic interference, where continued learning
causes a neural network to forget previously learned knowledge [37].</p>
        <p>
          While LTNs have already been used in the context of ontologies [
          <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
          ], their underlying
Real Logic is not intended as a language for writing domain ontologies. Moreover, to the best
of our knowledge, its formal properties (e.g., expressivity, decidability, or complexity) have
not been thoroughly analyzed, yet. For our current purposes, Real Logic is merely used as a
translation tool for encoding relevant domain knowledge from a given ontology as constraints
for a machine learning process and for extracting structured knowledge (which may then be
added to the ontology) from a trained machine learning system. The combination of conceptual
spaces and LTNs is thus in principle applicable to any symbolic language.
        </p>
        <p>Finally, we would like to mention the recent work by Singh et al. [38], who combine the
computational power of deep ANNs with a psychological model of categorization. Their
endto-end model learns both a similarity space and a prototype-based categorization model at
once. One could envision a similar application of LTNs: The input domain contains raw images,
which are then mapped by a function symbol (implemented as deep neural network) into a
low-dimensional conceptual space. In this conceptual space, one can then define membership
functions for the diferent concepts under consideration. The whole system could then be
trained based on labeled examples, but also using additional background knowledge based on
human similarity ratings and general rules from the symbolic layer. In the terms of cognitive
psychology, this would result in a combination of prototype theory (represented by convex
membership functions) with the knowledge view on concepts (represented by the presence of
constraints from background knowledge), spanning all three layers of representation.</p>
        <p>When viewed from the perspective of cognitive science, logic tensor networks can of course
not be labeled as a cognitively plausible learning mechanism: They rely on batch-processing
large amounts of (typically labeled) data with gradient descent. One can of course argue that
LTNs are not used to model the human concept acquisition process itself, but rather to take a
shortcut to the resulting concept inventory. However, it would certainly also be interesting to
extend LTNs such that they can work in an incremental way. A potential example application
in this context are language games [39], where a population of agents needs to negotiate a
common conceptualization of the world, receiving only indirect feedback through the success
or failure of their interactions.</p>
        <p>In the following, we will take a look at diferent membership functions from the conceptual
spaces literature and discuss their applicability in logic tensor networks. For illustration
purposes, we will consider a one-dimensional conceptual space with two concepts 1 and 2 as
well as three data points 1, 2, 3, which are supposed to belong to 1, but not to 2. Since
the parameters of the membership functions are optimized through gradient descent, we will
especially focus on their derivatives.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Partitional Membership Functions</title>
        <p>Let us first consider membership functions which partition the underlying conceptual space, i.e.,
which aim to assign each point to exactly one concept. We start with Gärdenfors’ approach of
identifying concepts with a prototypical point and creating a Voronoi tessellation of the space
[9, Section 3.9]: One starts from a set of prototypical points 1, . . . ,  for the  concepts under
consideration. Each point  in the conceptual space is then assigned to its closest prototype 
based on the distances (, ). As Gärdenfors [9, Section 3.9] argues, a Voronoi tessellation
based on the Euclidean metric partitions the overall space into convex regions.</p>
        <p>Since the Voronoi tessellation gives us a partitioning of the overall space, the membership
function of each concept  is constant almost everywhere and undefined on the border line to
a neighboring conceptual region (see Figure 2a). Therefore, the derivative of this membership
function with respect to any variable is either zero or undefined, which is highly problematic
for gradient descent. For example, consider the point 3, which is currently misclassified as
belonging to 2 instead of 1. In gradient-based optimization, the prototypes 1 and 2 are
updated by computing the derivative  (3) and then slightly increasing or decreasing the

value of , depending on the sign of the derivative and whether we want to increase or decrease
 (3). In the case of Figure 2a, we however note that both derivatives are zero – small changes
to 1 and 2 do not result in any changes to  (3). Thus, gradient descent is incapable of
making any update to the prototypes.</p>
        <p>In order to make gradient-based learning possible, we need a soft version of the Voronoi
approach. We can express the classification decision of the Voronoi tessellation as follows
(where  &gt; 0 is a sensitivity parameter):</p>
        <p>=  (,  ) =  (−  · (,  ))
Instead of the crisp  function (which results in flat membership values), we can now
apply the so called softmax function, which is commonly used in neural networks to provide
an output probability distribution over a set of mutually exclusive classes:

 () = ∑︀ 
Here,  () gives the probability for class , given a vector of raw confidence values .
If we combine this with the Voronoi tessellation approach, we obtain a soft Voronoi tessellation
with the following membership function:
− · (,)
 () =  (−  · (, )) = ∑︀ − · (,)
As we can see, the numerator reflects the semantic similarity of  and , while the denominator
is the sum over all similarities to all prototypes. The resulting membership function can thus
also be interpreted as a normalized version of semantic similarity. Figure 2b illustrates this
membership function: We now have a continuous transition from high membership values to
low membership values. Moreover, the derivative of this membership function is defined on the
whole conceptual space and nonzero in all cases - even points such as 1 have a very small, but
nonzero derivative.</p>
        <p>We should highlight at this point that we assume that the same sensitivity parameter  is used
for all concepts. If we allow diferent values 1 ̸= 2, we can control the size of the respective
conceptual regions (smaller values of  leading to larger regions). However, these diferent
values may cause some unintended efects. For instance, Figure 3a illustrates the case where
1 ≫ 2, which causes the membership function  2() to be no longer convex.</p>
        <p>Generalized Voronoi tessellations [9, Section 4.9] allow to encode diferently sized conceptual
regions by considering prototypical regions  instead of prototypical points . These
prototypical regions are usually represented as disks with a central point  and a radius . Based on these
prototypical regions, one can now generate a generalized Voronoi tessellation by assigning each
point  in the conceptual space to the concept whose prototypical region is closest. In the case
of disks, this corresponds to finding the concept  for which (, ) = max(0, (, ) − )
is smallest. Concepts with larger prototypical regions (as reflected through a larger value of )
thus result in larger conceptual regions in the generalized Voronoi tessellation (see Figure 2c).
Again, by using the   instead of the  function, this can be generalized to a soft
notion of concept membership (see Figure 2d).</p>
        <p>Also Douven et al. [40] consider prototypical regions instead of prototypical points. However,
they create all possible Voronoi diagrams by picking a single point  ∈  for all prototypical
regions . These individual Voronoi tessellations are then aggregated into a so called “collated
Voronoi diagram”: A point  is assigned to concept  if and only if it has been assigned to
 in all individual Voronoi diagrams. Douven et al. identify borderline cases as points  that
belong to diferent conceptual regions for diferent Voronoi diagrams. These borderline points
are not assigned to any concept and represent vagueness in concept boundaries. In Figure 2e,
we again note that the derivative is zero within the conceptual regions and undefined in the
border area.</p>
        <p>Decock and Douven [41] extend the work of Douven et al. [40] by providing a degree of
1 ≫ 2. (b) Workaround for zero gradient inside prototypical regions.
membership for borderline cases. They define the membership of a point  to a concept 
as the fraction of individual Voronoi diagrams for which  belongs to the conceptual region
of . Decock and Douven note that if the prototypical regions  have an infinite number of
points, then the membership function is s-shaped (cf. Figure 2f). However, we can observe
that the membership function is flat for large parts of the conceptual space, namely, for all
non-borderline points. This is again highly problematic for gradient descent.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Nonpartitional Membership Functions</title>
        <p>The usage of Voronoi tessellations for conceptual spaces has not been without challenge in the
literature. For instance, Lewis and Lawry [42] argue that partitioning the conceptual space may
be adequate for individual domains such as color, but that it is not suitable for a combination of
multiple domains. It seems implausible that every single point in a high-dimensional space has
to be assigned to exactly one category: On the one hand, some regions in the overall conceptual
space may not be covered by any existing concept. Points in such regions should be recognized
as outliers or members of a novel, previously unknown category. On the other hand, conceptual
regions may also overlap, for instance in order to represent conceptual hierarchies.</p>
        <p>Lewis and Lawry [42] have also made a general proposal for nonpartitional membership
functions: A point  in the conceptual space is said to belong to concept  if its distance to the
prototypical region  is not greater than a threshold distance  . Lewis and Lawry assume that
the threshold   is not known, but that a probability distribution   over its possible values is
available. The degree of membership of a point  to a concept  is then given by the probability
of (, ) being smaller than  :
 () = P ((, ) ≤  ) =</p>
        <p>( ) 
∫︁ ∞
(,)
Lewis and Lawry are in general open to diferent forms for the probability distribution  . If we
use  ( ) =  · − ·  , then concept membership reflects similarity to the prototypical region:
 () =
∫︁ ∞
(,)
 · − ·     = [︀ − − ·   ]︀  →∞
 =(,) = 0 −
︁(
− − · (,)︁)
= − · (,)
The shape of the resulting similarity function for  = {} (i.e., prototypical points) is illustrated
in Figure 4a. As we can see, all points in the similarity space receive a non-zero membership
value. Moreover, the derivative of the membership function is defined for all points except
for the prototypes 1 and 2. In practical applications, this theoretical shortcoming can be
overcome by defining the derivative in this point to equal zero. Furthermore, we are able to
control the size of the conceptual regions by choosing diferent sensitivity parameters 1 ̸= 2.</p>
        <p>However, we can also note that the derivative of the membership function is proportional to
the membership value itself: The largest derivatives are observed for the points with the highest
membership in the concept. Since gradient descent algorithms typically take into account not
only the direction, but also the magnitude of the gradient, this can lead to undesired efects.
Consider for instance the point 2 in Figure 4a, which has a fairly high membership to 1. The
derivative  1(2) is quite large and will thus cause the gradient descent algorithm to increase
1
1 considerably. In the resulting configuration,  1(2) may however be smaller than before,
since 1 may have moved considerably past 2. This seems to be a major shortcoming of this
similarity-based approach to concept membership.</p>
        <p>The examples by Lewis and Lawry [42] often make use of uniform distributions  ( ) =
  (0, ). As we can see in Figure 4b, the membership curve has in this case a triangular
shape and its derivative is therefore constant for all points with a partial membership. However,
both concept membership and its derivative are zero for most parts of the similarity space.</p>
        <p>These considerations can of course also be generalized to prototypical regions , which
subsumes our own formalization of the conceptual spaces framework [43, 44, 45]: There exists a
well-defined region with full membership, which in our case is based on the union of axis-aligned
cuboids. Membership is then defined as similarity to this prototypical region.</p>
        <p>In Figure 4c, we can see two problems with this approach: On the one hand, we again have
the problem of large derivatives for large membership values as already discussed for Figure
4a. On the other hand, the membership function is constant for all points in the prototypical
region, hence, the derivative is zero. If an observation such as 3 is confidently misclassified as
belonging to 2, then gradient descent is not able to move 2 away from 3.</p>
        <p>The problem of a zero derivative could be circumvented as follows: We define a new
membership function  ′() := (1 −  ) ·  () for some small  &gt; 0. Furthermore, we identify the
central point  ∈ . The membership value for  ∈  is then increased based on its distance
to , such that  ′() = 1 and that  ′() = 1 −  for points on the border of . This provides
a small slope for the membership function inside the prototypical region and thus a nonzero
derivative (cf. Figure 3b). However, it remains to be seen whether such a workaround is useful
in practice.</p>
        <p>Despite these shortcomings, there are however reasonably strong arguments for using a
membership function like the one proposed in our formalization: Firstly, by using a union of
axis-aligned cuboids, our formalization is able to represent correlations between domains. This
is an important aspect of human conceptualization [46, 47] which is not captured by any of
the aforementioned approaches. Secondly, one can apply a variety of operations defined in the
context of our formalization in order to reason on the learned concepts. For instance, relations
such as conceptual similarity and conceptual betweenness are not defined in LTNs, but they
become immediately available with the use of our proposed formalization of concepts. Thirdly,
logical formulas in LTNs always have to be evaluated on a set of data points which requires
that one keeps all examples in memory. Our formalization on the other hand provides closed
formulas for computing the validity of such logical formulas – the original data points are
not needed any more and the computation can potentially be faster. However, the operations
defined in our formalization are based on the minimum norm, while LTNs are commonly used
with the product norm. Therefore, the numeric results of the computations might difer.</p>
        <p>Motivated by the problem of large gradients for large membership values, we also consider
multivariate Gaussian functions, whose membership value can be defined as follows with a
symmetric, positive semi-definite matrix Σ :</p>
        <p>() = − 21 (− ) Σ− 1(− )</p>
        <p>Figure 4d illustrates the usage of such Gaussian functions in our one-dimensional similarity
space.4 As one can see, this type of membership function does not sufer from the gradient size
problem as identified in Figure 4a: The derivative is small both for points with a very low and
for points with a very high membership. It is largest for points with an intermediate level of
membership, i.e., points that are currently treated as borderline members. Another advantage of
multivariate Gaussian functions is that they are able to encode correlations between dimensions
as well as diferent distribution widths through their covariance matrix Σ .</p>
        <p>2
4We can model this with  ( ) =   2 · − 2 2 in the one-dimensional case using the approach by Lewis and
Lawry [42]</p>
        <p>However, the usage of Gaussians in the context of conceptual spaces is somewhat
unsatisfactory from a theoretical standpoint. The notion of similarity is not based on the
Euclidean distance  (, ) = √︀∑︀( − )2, but on the squared Mahalanobis distance
 (, ) = √︀( − ) Σ − 1( − ). Applying the Mahalanobis distance corresponds to
transforming the similarity space with the covariance matrix, and then computing the Euclidean
metric in the transformed space. This implicit transformation of the similarity space would in
our opinion cause a major modification of the original framework. Nevertheless, the simplicity
and computational attractiveness of multivariate Gaussians make them an interesting candidate
for experimental investigations, such that one should not hastily dismiss them.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions</title>
      <p>In this paper, we have introduced both conceptual spaces and logic tensor networks. We
have argued that a combination of the two frameworks is a promising direction of research:
Conceptual spaces can provide a grounding for the feature spaces considered in logic tensor
networks and allow us to use relatively simple membership functions for representing predicates.
Logic tensor networks on the other hand can help us to bridge the gap between the conceptual
and the symbolic layer by learning concepts not only based on labeled examples, but also
based on general logical top-down constraints. Moreover, the resulting system can provide
a cognitive grounding for formal ontologies: Individual concepts from the ontology can be
grounded in regions of a conceptual space, whose dimensions are grounded in psychological
data and/or perceptual sensor information. Moreover, this grounding can take into account the
most important part of ontologies, namely the relations between concepts. Since the envisioned
system unifies both bottom-up and top-down processes, the information from the conceptual
layer can furthermore give rise to additional rules for the symbolic ontology.</p>
      <p>Moreover, we have also discussed several possible membership functions for concepts in
conceptual spaces and their applicability to gradient-based optimization methods. If we consider
partitional approaches, a soft version of generalized Voronoi tessellations seems to be most
promising: It is capable of representing conceptual regions of diferent sizes and comes with
a derivative that is guaranteed to be non-zero everywhere. If we are however interested in
nonpartitional membership functions, multivariate Gaussians seem to be preferable from a
computational point of view: They are able to explicitly encode correlations between dimensions,
they can take into account concepts of varying size, and they provide a meaningful non-zero
gradient everywhere. Nevertheless, also the membership function of our own formalization of
the conceptual spaces framework should be explored, since it provides us with a large number
of operations for downstream reasoning processes.</p>
      <p>Our proposal has so far been only a theoretical one. In order to evaluate its actual merit,
practical experiments need to be conducted. Ideally, these experiments should consider all
membership functions discussed in this paper in order to confirm or refute our theoretical
analyses. There are several data sets that can serve as test beds for a first study, including the
conceptual spaces extracted by Banaee et al. [48] and Derrac and Schockaert [49], as well as the
robotics data set by Spranger et al. [50]. Since the strength of LTNs stems from their ability
to incorporate top-down rules to compensate for scarce training data, especially the movie
spaces from Derrac and Schockaert [49] are relevant: Each movie is annotated with its genres,
a set of plot keywords, at its age restriction. Using techniques such as the apriori algorithm
[51], one can extract rules from the co-occurrence statistics of the labels and then simulate
few-shot learning [52] by showing only a small part of the available examples, but providing
the general rules as additional constraints to the system. After such initial experiments, studies
with actual ontologies are needed in order to ensure that all relevant pieces of ontological
information (especially relations of varying complexity) can be adequately encoded by the
proposed approach.
Rules and Reasoning, Springer International Publishing, Cham, 2019, pp. 161–170.
[30] S. Auer, C. Bizer, G. Kobilarov, J. Lehmann, R. Cyganiak, Z. Ives, DBpedia: A Nucleus for
a Web of Open Data, in: K. Aberer, K.-S. Choi, N. Noy, D. Allemang, K.-I. Lee, L. Nixon,
J. Golbeck, P. Mika, D. Maynard, R. Mizoguchi, G. Schreiber, P. Cudré-Mauroux (Eds.), The
Semantic Web, Springer Berlin Heidelberg, Berlin, Heidelberg, 2007, pp. 722–735.
[31] I. Donadello, L. Serafini, A. d’Avila Garcez, Logic Tensor Networks for Semantic Image
Interpretation, in: Proceedings of the Twenty-Sixth International Joint Conference on
Artificial Intelligence (IJCAI-17), 2017.
[32] B. Wagner, A. d’Avila Garcez, Neural-Symbolic Integration for Fairness in AI, in: A. Martin,
K. Hinkelmann, H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.), Proceedings
of the AAAI 2021 Spring Symposium on Combining Machine Learning and Knowledge
Engineering (AAAI-MAKE 2021), 2021.
[33] S. Badreddine, M. Spranger, Injecting Prior Knowledge for Transfer Learning into
Reinforcement Learning Algorithms using Logic Tensor Networks (2019). arXiv:1906.06576.
[34] G. Murphy, The Big Book of Concepts, MIT Press, 2002.
[35] D. L. Medin, M. M. Schafer, Context Theory of Classification Learning, Psychological</p>
      <p>Review 85 (1978) 207.
[36] A. K. Engel, A. Maye, M. Kurthen, P. König, Where’s the Action? The Pragmatic Turn
in Cognitive Science, Trends in Cognitive Sciences 17 (2013) 202–209. doi:10.1016/j.
tics.2013.03.006.
[37] M. McCloskey, N. J. Cohen, Catastrophic Interference in Connectionist Networks: The
Sequential Learning Problem, volume 24 of Psychology of Learning and Motivation, Academic
Press, 1989, pp. 109–165. doi:10.1016/S0079-7421(08)60536-8.
[38] P. Singh, J. Peterson, R. Battleday, T. Grifiths, End-to-end Deep Prototype and Exemplar
Models for Predicting Human Behavior, in: Proceedings for the 42nd Annual Meeting of
the Cognitive Science Society, 2020.
[39] L. Steels, The Talking Heads Experiment: Origins of Words and Meanings, Language</p>
      <p>Science Press, 2015.
[40] I. Douven, L. Decock, R. Dietz, P. Égré, Vagueness: A Conceptual Spaces Approach, Journal
of Philosophical Logic 42 (2011) 137–160. doi:10.1007/s10992-011-9216-0.
[41] L. Decock, I. Douven, What Is Graded Membership?, Noûs 48 (2014) 653–682. doi:10.</p>
      <p>1111/nous.12003.
[42] M. Lewis, J. Lawry, Hierarchical Conceptual Spaces for Concept Combination, Artificial</p>
      <p>Intelligence 237 (2016) 204–227. doi:10.1016/j.artint.2016.04.008.
[43] L. Bechberger, K.-U. Kühnberger, A Thorough Formalization of Conceptual Spaces, in:
G. Kern-Isberner, J. Fürnkranz, M. Thimm (Eds.), KI 2017: Advances in Artificial
Intelligence: 40th Annual German Conference on AI, Dortmund, Germany, September 25–29,
2017, Proceedings, Springer International Publishing, 2017, pp. 58–71. doi:10.1007/
978-3-319-67190-1_5.
[44] L. Bechberger, K.-U. Kühnberger, Formal Ways for Measuring Relations between Concepts
in Conceptual Spaces, Expert Systems 0 (2018) e12348. doi:10.1111/exsy.12348, e12348
EXSY-Apr-18-107.R1.
[45] L. Bechberger, K.-U. Kühnberger, Formalized Conceptual Spaces with a Geometric
Representation of Correlations, Springer International Publishing, Cham, 2019, pp. 29–58.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Harnad</surname>
          </string-name>
          ,
          <source>The Symbol Grounding Problem, Physica D: Nonlinear Phenomena</source>
          <volume>42</volume>
          (
          <year>1990</year>
          )
          <fpage>335</fpage>
          -
          <lpage>346</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0167</fpage>
          -
          <lpage>2789</lpage>
          (
          <issue>90</issue>
          )
          <fpage>90087</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gärdenfors</surname>
          </string-name>
          ,
          <article-title>How to Make the Semantic Web More Semantic</article-title>
          ,
          <source>in: Formal Ontology in Information Systems</source>
          ,
          <year>2004</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O. Lütfü</given-names>
            <surname>Özçep</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leemhuis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wolter</surname>
          </string-name>
          ,
          <article-title>Cone Semantics for Logics with Negation</article-title>
          , in: C.
          <string-name>
            <surname>Bessiere</surname>
          </string-name>
          (Ed.),
          <source>Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, International Joint Conferences on Artificial Intelligence Organization</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1820</fpage>
          -
          <lpage>1826</lpage>
          . doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2020</year>
          /252, main track.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          ,
          <source>in: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          , Language Models are Few-Shot
          <string-name>
            <surname>Learners</surname>
          </string-name>
          (
          <year>2020</year>
          ). arXiv:
          <year>2005</year>
          .14165.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G.</given-names>
            <surname>Marcus</surname>
          </string-name>
          , E. Davis,
          <string-name>
            <surname>Rebooting</surname>
            <given-names>AI</given-names>
          </string-name>
          :
          <string-name>
            <surname>Building Artificial Intelligence We Can Trust</surname>
          </string-name>
          , Pantheon,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. d.</given-names>
            <surname>Garcez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Besold</surname>
          </string-name>
          , L. De Raedt,
          <string-name>
            <given-names>P.</given-names>
            <surname>Földiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Icard</surname>
          </string-name>
          , K.-U. Kühnberger,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Lamb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Miikkulainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <article-title>Neural-Symbolicy Learning and Reasoning: Contributions and Challenges</article-title>
          ,
          <source>in: AAAI 2015 Spring Symposium on Knowledge Representation and Reasoning: Integrating Symbolic and Neural Approaches</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lieto</surname>
          </string-name>
          ,
          <source>Cognitive Design for Artificial Minds, Routledge</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gärdenfors</surname>
          </string-name>
          , Conceptual Spaces: The Geometry of Thought, MIT Press,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zenker</surname>
          </string-name>
          , P. Gärdenfors (Eds.),
          <source>Applications of Conceptual Spaces</source>
          , Springer Science + Business Media,
          <year>2015</year>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -15021-5.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kaipainen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zenker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hautamäki</surname>
          </string-name>
          , P. Gärdenfors (Eds.),
          <source>Conceptual Spaces: Elaborations and Applications</source>
          , volume
          <volume>405</volume>
          , Springer,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Serafini</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>d'Avila Garcez, Logic Tensor Networks: Deep Learning and Logical Reasoning from Data and Knowledge (</article-title>
          <year>2016</year>
          ). arXiv:
          <volume>1606</volume>
          .
          <fpage>04422</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Badreddine</surname>
          </string-name>
          , A.
          <string-name>
            <surname>d'Avila Garcez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Spranger</surname>
          </string-name>
          , Logic Tensor Networks,
          <year>2021</year>
          . arXiv:
          <year>2012</year>
          .13635.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bechberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Scheibel</surname>
          </string-name>
          ,
          <article-title>Analyzing Psychological Similarity Spaces for Shapes</article-title>
          , in: M.
          <string-name>
            <surname>Alam</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Braun</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          Yun (Eds.),
          <source>Ontologies and Concepts in Mind and Machine</source>
          , Springer International Publishing, Cham,
          <year>2020</year>
          , pp.
          <fpage>204</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>I.</given-names>
            <surname>Borg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Groenen</surname>
          </string-name>
          ,
          <source>Modern Multidimensional Scaling: Theory and Applications</source>
          , Springer Series in Statistics, 2nd ed., Springer-Verlag New York,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Attneave</surname>
          </string-name>
          , Dimensions of Similarity,
          <source>The American Journal of Psychology</source>
          <volume>63</volume>
          (
          <year>1950</year>
          )
          <fpage>516</fpage>
          -
          <lpage>556</lpage>
          . doi:
          <volume>10</volume>
          .2307/1418869.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>R. N.</given-names>
            <surname>Shepard</surname>
          </string-name>
          ,
          <article-title>Attention and the Metric Structure of the Stimulus Space</article-title>
          ,
          <source>Journal of Mathematical Psychology</source>
          <volume>1</volume>
          (
          <year>1964</year>
          )
          <fpage>54</fpage>
          -
          <lpage>87</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0022</fpage>
          -
          <lpage>2496</lpage>
          (
          <issue>64</issue>
          )
          <fpage>90017</fpage>
          -
          <lpage>3</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>R. M. Battleday</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Peterson</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          <string-name>
            <surname>Grifiths</surname>
          </string-name>
          ,
          <article-title>From Convolutional Neural Networks to Models of Higher-Level Cognition (and Back Again)</article-title>
          ,
          <source>Annals of the New York Academy of Sciences</source>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1111/nyas.14593.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bechberger</surname>
          </string-name>
          , K.-U. Kühnberger, Generalizing Psychological Similarity Spaces to Unseen Stimuli - Combining
          <source>Multidimensional Scaling with Artificial Neural Networks</source>
          , Springer International Publishing, Cham,
          <year>2021</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -69823-
          <issue>2</issue>
          _
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bechberger</surname>
          </string-name>
          , K.-U. Kühnberger,
          <string-name>
            <surname>Mapping Line Drawings Into Shape Space - Combining Convolutional Neural</surname>
          </string-name>
          <article-title>Networks with Psychological Similarity Spaces, Machine Learning (under review).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Sanders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Nosofsky</surname>
          </string-name>
          ,
          <article-title>Using Deep-Learning Representations of Complex Natural Stimuli as Input to Psychological Models of Classification</article-title>
          ,
          <source>in: Proceedings of the 2018 Conference of the Cognitive Science Society</source>
          , Madison.,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Sanders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Nosofsky</surname>
          </string-name>
          ,
          <article-title>Training Deep Networks to Construct a Psychological Feature Space for a Natural-Object Category Domain</article-title>
          ,
          <source>Computational Brain &amp; Behavior</source>
          <volume>3</volume>
          (
          <year>2020</year>
          )
          <fpage>229</fpage>
          -
          <lpage>251</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>E.</given-names>
            <surname>Rosch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. B.</given-names>
            <surname>Mervis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. D.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Boyes-Braem</surname>
          </string-name>
          ,
          <article-title>Basic Objects in Natural Categories</article-title>
          ,
          <source>Cognitive Psychology 8</source>
          (
          <year>1976</year>
          )
          <fpage>382</fpage>
          -
          <lpage>439</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0010</fpage>
          -
          <lpage>0285</lpage>
          (
          <issue>76</issue>
          )
          <fpage>90013</fpage>
          -
          <lpage>x</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gärdenfors</surname>
          </string-name>
          ,
          <source>The Geometry of Meaning: Semantics Based on Conceptual Spaces</source>
          , MIT Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>L.</given-names>
            <surname>Serafini</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. S.</surname>
          </string-name>
          <article-title>d'Avila Garcez, Learning and Reasoning with Logic Tensor Networks</article-title>
          , Springer International Publishing, Cham,
          <year>2016</year>
          , pp.
          <fpage>334</fpage>
          -
          <lpage>348</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>319</fpage>
          -49130-1_
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>L.</given-names>
            <surname>Serafini</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Donadello</surname>
          </string-name>
          , A. d. Garcez,
          <article-title>Learning and Reasoning in Logic Tensor Networks: Theory and Application to Semantic Image Interpretation</article-title>
          ,
          <source>in: Proceedings of the Symposium on Applied Computing</source>
          , SAC '17,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2017</year>
          , p.
          <fpage>125</fpage>
          -
          <lpage>130</lpage>
          . doi:
          <volume>10</volume>
          .1145/3019612.3019642.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>I.</given-names>
            <surname>Donadello</surname>
          </string-name>
          , L. Serafini,
          <article-title>Compensating Supervision Incompleteness with Prior Knowledge in Semantic Image Interpretation</article-title>
          , in: 2019
          <source>International Joint Conference on Neural Networks (IJCNN)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bianchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          ,
          <article-title>On the Capabilities of Logic Tensor Networks for Deductive Reasoning</article-title>
          ,
          <source>in: Proceedings of the AAAI Spring Symposium on Combining Machine Learning with Knowledge Engineering</source>
          , AAAI-MAKE,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bianchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hitzler</surname>
          </string-name>
          , L. Serafini,
          <article-title>Complementing Logical Reasoning with Sub-symbolic Commonsense</article-title>
          , in: P. Fodor,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Calvanese</surname>
          </string-name>
          , D. Roman (Eds.),
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>