<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Capabilities of Logic Tensor Networks for Deductive Reasoning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Federico Bianchi</string-name>
          <email>federico.bianchi@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pascal Hitzler</string-name>
          <email>pascal.hitzler@wright.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright held by the author(s). In A. Martin, K. Hinkelmann, A. Gerber</institution>
          ,
          <addr-line>D. Lenat, F. van Harmelen, P. Clark (Eds.)</addr-line>
          ,
          <institution>Proceedings of the AAAI 2019 Spring Symposium on Combining Machine Learning with Knowledge Engineering (AAAI-MAKE 2019). Stanford University</institution>
          ,
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Milan-Bicocca</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Wright State University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Neural-symbolic integration is a field in which classical symbolic knowledge mechanisms are combined with neural networks. This is done to provide satisfactory computational capabilities from the network side and to exploit the descriptive power of symbolic reasoning. Logic Tensor Networks (LTNs) are a deep learning model that can be used to combine data with fuzzy logic to provide inferences and reasoning mechanisms over data. While LTNs have been shown effective in some contexts no detailed analysis on their capabilities for deductive logical reasoning has been conducted. In this paper we explore the capabilities and the limitations of LTNs in terms of deductive reasoning.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Neural-symbolic learning and reasoning
        <xref ref-type="bibr" rid="ref10 ref2">(Garcez, Lamb, and
Gabbay 2008; Besold et al. 2017)</xref>
        involves integrating
standard logical reasoning with neural networks with the aim
of providing fast and robust computational methods for
reasoning and explanation over data. Logic Tensor Networks
(LTNs) are a deep learning model that comes from the
neural-symbolic field: it integrates both logic and data in
a neural network to provide support for neural
symboliclearning and reasoning
        <xref ref-type="bibr" rid="ref21">(Serafini and Garcez 2016)</xref>
        . LTNs
use first-order fuzzy logic to express knowledge about the
world: using fuzzy logic over classical first-order logic
allows us to represent truth using continuous values in the
interval [0; 1] to represent the degree of truth.
      </p>
      <p>Input to LTNs are data and axioms over (fuzzy)
firstorder predicate logic, e.g., parent(Ann, Susan), 8x; y :
parent(x; y) ! ancestor(x; y). Two key components of
logic tensor networks are the grounding of formulas and the
learning by best satisfiability. With formula grounding we
refer to the mapping of formulas to a vector space. For
example, constants are mapped to n-dimensional vectors while
function symbols are mapped to linear functions. A neural
network can be used to compute the degree of truth of a
given formula considering the embedded representation of
constants and symbols.</p>
      <p>
        Deep learning models
        <xref ref-type="bibr" rid="ref11 ref21">(Goodfellow, Bengio, and
Courville 2016)</xref>
        usually learn by optimizing a function; in
LTNs this task is replaced with the task of best satisfiability:
the model has to optimize the representation of each atom,
function and predicate in such a way that the satisfiability
of each formula is maximized. In this way the network
learns the best possible parameters to represent both data
and axioms.
      </p>
      <p>
        The main advantages of LTNs are the following: i) it is
possible to express knowledge using logical axioms over
data ii) it is possible to tackle and solve standard machine
learning tasks (e.g., classification) and iii) provide
explanations using fuzzy logic over the trained network. Indeed,
after training it is possible to make fuzzy inferences over
data to obtain the degree of truth with respect to certain
predicates. The model was tested with promising results
on simple reasoning tasks
        <xref ref-type="bibr" rid="ref21">(Serafini and Garcez 2016)</xref>
        and
on semantic image interpretation
        <xref ref-type="bibr" rid="ref2 ref6">(Donadello, Serafini, and
d’Avila Garcez 2017)</xref>
        .
      </p>
      <p>
        An initial exploration of the reasoning capabilities of
LTNs was done on the well-known
smoker-friends-andcancer dataset
        <xref ref-type="bibr" rid="ref21">(Serafini and Garcez 2016)</xref>
        . The dataset
contains data about two groups of people for which friend
relationships and smoking habits are given, while the fact of
having cancer or not is only given for people in the first
group. Axioms related to smoking properties (i.e.,
smoking implies cancer) are given to the network. The network
learns to predict if people in the second group have cancer
having learned the patterns present in the first group. More
recently, LTNs were used on a semantic image interpretation
task in which they learned to classify bounding boxes of
images with the help of background knowledge
        <xref ref-type="bibr" rid="ref2 ref6">(Donadello,
Serafini, and d’Avila Garcez 2017)</xref>
        . Still, an in-depth
analysis of the deductive reasoning capabilities of LTNs remains
to be done.
      </p>
      <p>In this work we explore LTNs in the context of
reasoning tasks, showing insights and properties of the model. We
introduce two simple datasets that contain relationships and
we define additional axioms over these datasets. These two
datasets are used to evaluate deductive reasoning
capabilities. We also perform some experiments on the
computation time that is required to learn model parameters. Our
results show that LTNs are a good model that can fit well
the data and that is able to do simple deductive inferences.
The real added value of the model is that it lends itself to
explanations, since it allows us to do after-training fuzzy
inferences over the data. Nevertheless, the model generates
some errors, in particular when multi-hop inferences are to
be drawn, and thus some refinements over the general model
might be required to improve the results.</p>
      <p>The rest of the paper is organized as follows: in Section
2 we describe LTNs showing the basic definitions and the
learning process, in Section 3 we introduce our
experimental setting and we describe and evaluate the results of our
experiments. Section 4 contains other related work. Finally,
we end the paper in Section 5 with some conclusions and
future work.</p>
    </sec>
    <sec id="sec-2">
      <title>Logic Tensor Networks</title>
      <p>
        LTNs use first-order fuzzy logic
        <xref ref-type="bibr" rid="ref19">(Petr 1998)</xref>
        and embed
atoms, functions, and predicates in a vector space. LTNs
are inspired by Neural Tensor Networks
        <xref ref-type="bibr" rid="ref22">(Socher et al. 2013)</xref>
        that have been shown to be effective in natural logic
reasoning tasks
        <xref ref-type="bibr" rid="ref4 ref5">(Bowman, Potts, and Manning 2015)</xref>
        . In the
following sections we will give a short primer on logic tensor
networks and their learning methodology. More details on
LTNs can be found in the paper in which they were first
introduced
        <xref ref-type="bibr" rid="ref21">(Serafini and Garcez 2016)</xref>
        . To describe LTNs we
will follow the definitions given by Serafini and Garcez.
      </p>
      <sec id="sec-2-1">
        <title>Logic</title>
        <p>LTNs are implemented over a logic called Real Logic that
is described by a language L that contains a set of constants
C, a set of function symbols F and a set of predicates P .
In this language rules from fuzzy logic apply and
connectives are interpreted as binary operations over real numbers
in [0; 1]. For example t-norms are used in place of the
conjunction from classical logic. The t-norm is an operation
[0; 1]2 ! [0; 1] and different versions of the operation exist
(Lukasiewicz, Go¨del and product t-norms are some possible
examples). Once the t-norm is chosen also the other
connectives can be defined with respect to it. Thus, the use of
t-norms and the other fuzzy connectives allows us to operate
on real-values in the interval [0; 1].</p>
      </sec>
      <sec id="sec-2-2">
        <title>Grounding</title>
        <p>Each element of the language L is grounded in the vector
space. Constants are mapped to vectors in Rm while
function symbols are mapped to functions in the vector space.
An n-ary function symbol is mapped to an n-ary function
Rk n ! Rm. Predicates are mapped to functions with
codomain in [0; 1]: Rm n ! [0; 1]; the predicate is mapped to
a fuzzy subset that defines the degree of truth (membership
to the set) for that predicate given its arguments.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Networks</title>
        <p>The dimensionality of the vector of the constant is an
hyperparameter of the model. While constants are mapped to
vectors, functions and predicates are mapped to actual
operations over the vector space. We will use G(f ) and G(P )
to identify groundings of functions and predicates.
Function symbols are implemented as linear functions: given f
a symbol function of arity m and v1; : : : ; vm 2 Rn are the
groundings of m terms then the grounding for the symbol
function f can be expressed as:</p>
        <p>G(f )(v1; : : : ; vm) = Mf v + Bf
(1)
where v = hv1; : : : ; vmi, Mf is a transformation matrix and
Bf is the bias. This operation can be encoded into a
onelayer neural network.</p>
        <p>
          Predicates are instead mapped to neural tensor
operations
          <xref ref-type="bibr" rid="ref22">(Socher et al. 2013)</xref>
          , the output of the neural tensor
network is given in input to a sigmoid such that the final
output of the predicate is a value in the interval [0; 1]. The
tensor operation is the following:
        </p>
        <p>G(P )(v) = (uTP (tanh(vT W P[1:k]v + VP v + BP ))) (2)
is the sigmoid function while W , V , B and u are
parameters to be learned by the network while k corresponds to
the layer size of the tensor and is an hyper-parameter in the
network.</p>
        <p>Quantifiers like 8 in fuzzy logic are defined with
aggregation functions (like the min): this should consider an
aggregation over an infinite number of instances, making it
impossible to compute. Thus, quantifiers are implemented as
aggregation operations over a subset of the domain space
Rk. Different possible implementations can be used to
implement the aggregation for the universal quantifiers, for
example mean, min and hmean (harmonic mean).</p>
      </sec>
      <sec id="sec-2-4">
        <title>Learning to Satisfy Formulas</title>
        <p>
          LTNs reduce the learning problem to a maximum
satisfiability problem: the task is to find groundings for
atoms, predicates and formulas that maximize the
satisfiability of a given formula. For example, given the formula
parent(Susan; Ann), which describes the fact that Susan
is one of Ann’s parents, the network will try to optimize the
groundings of the predicate parent (i.e., the parameters in the
tensor layer) and the groundings of Susan and Ann (i.e.,
their respective two vectors) in such a way that the degree
of truth of the formula is close to 1. Thus, the groundings
are both the embedded representation of the atoms and the
parameters in the networks that represent both functions and
predicates; the values of these components can be learned
through the use of back-propagation
          <xref ref-type="bibr" rid="ref11 ref21">(Goodfellow, Bengio,
and Courville 2016)</xref>
          . The output of the learning process is a
satisfiability score (in the interval [0; 1]) that can be
considered similar to the value of the loss function in a standard
deep learning setting.
        </p>
        <p>
          We show an example of how grounding and the
satisfiability are combined. For compactness, in this example we
will identify the grounding of each element with a G as a
superscript: given the formula P (x; y) ^ R(w; z), the
groundings for the constants x, w, z, and y are retrieved (denoted
with xG). P and R are grounded to the respective
operations: P G(xG; yG) ^ RG(wG; zG). The output of both
predicates is a real value in [0,1] that can be aggregated with the
use of the t-norm. LTNs will learn to optimize the
groundings in such a way that the final value is close to 1 (i.e., the
formula is satisfied).
In this experimental section we aim to obtain answers for the
following questions: i) what can LTNs learn and ii) how fast
is the LTNs learning phase. To allow easy replication of our
experiments we will first describe the datasets we use and
then we will introduce some details on the general
methodology we have followed during our experiments. Details that
are related to a particular experiment will be given in the
related section. For our experiment we use the original LTNs
TensorFlow implementation1 provided by the authors
          <xref ref-type="bibr" rid="ref21">(Serafini and Garcez 2016)</xref>
          . Datasets, code and results are
available online with specific instructions on how to repeat our
experiments2. We briefly summarize here the four
experiments we ran:
        </p>
        <p>Experiments 1 and 2 will concentrate on a knowledge
base completion task in which we will give to the network
only true predicates and some axioms;
Experiment 3 will compare LTNs with a simple deep
learning baseline to provide insights about the strength
and the limits of the model;
Experiment 4 will show computational times related to
experiments on learning with LTNs.</p>
      </sec>
      <sec id="sec-2-5">
        <title>Definitions</title>
        <p>By KBS we denote an input (starting) knowledge base, and
KB will denote the corresponding completed knowledge
base (i.e., with all relevant logical consequences added).
KBT denotes the set of all inferences not in KBS , i.e.,
KBT = KB n KBS . In the experiment we will often show
the performance over both KB and KBT by putting results
related to KBT within parentheses.</p>
      </sec>
      <sec id="sec-2-6">
        <title>Datasets</title>
        <p>We use mainly two datasets for our experiments, the first
one, called dataset A, represents a taxonomy that mainly
contains hierarchies of classes (inspired by the DBpedia
Ontology3). The taxonomy contains 25 nodes. Each node but
one (the root) has an outgoing edge to its superclass (i.e.,
Cat is connected to Feline). Figure 1 shows the taxonomy
used in the experiment.</p>
        <p>The second dataset P is a parent-ancestor dataset that
contains 17 nodes. Edges connect parents to one or more
children, for a total of 22 parental relationships. Figure 2 shows
the parental relationships.</p>
        <p>
          We will test these two datasets on tasks in which we will
have heavily unbalanced classes (more negative examples
than positive ones). While our datasets are small compared
to the ones currently used for knowledge base completion
tasks
          <xref ref-type="bibr" rid="ref3">(Bordes et al. 2013)</xref>
          , we think that the results of our
experiments can point out interesting capabilities of LTNs:
can LTNs perform deductive reasoning over these simple
datasets? Moreover, using these small datasets, results can
be manually inspected to better understand where and how
1https://github.com/logictensornetworks
2https://github.com/vinid/ltns-experiments
3http://dbpedia.org
Bank
        </p>
        <p>Company</p>
        <p>Organization Agent</p>
        <p>Lizard
Thing</p>
        <p>Reptile
SnakeCrocodile</p>
        <p>Animal
Human Mammal</p>
        <p>Dog</p>
        <p>FelineSquirreDlolphin
Cat</p>
        <p>BaldEagle</p>
        <p>LilBird
Bird
Fish</p>
        <p>Eagle</p>
        <p>Shark
BlueFish
Given a dataset we define a set of axioms and we test a
knowledge base completion task, showing hyper-parameter
details. LTNs will receive in input data under the form
of predicates (e.g., parent(Ann,Susan)) and axioms (e.g.,
8x; y : parent(x; y) ! ancestor(x; y)); the network will
learn groundings for all the parameters and in the testing
phase we will analyze predictions over data. Since
different configurations of hyper-parameters are possible we run
multiple models and we re-run each model multiple times
(to check variations due to random initialization). After a
first phase of trial and error we set as static the following
parameters: optimizer RMSprop, bias 1e 5, learning rate
0:01, decay 0:9. We tested three different aggregation
functions for the universal quantifiers (harmonic mean, mean and
min), two tensor layer sizes (10 and 20) for the tensor
network and two embedding sizes for constants (10 and 20
dimensions).</p>
        <p>
          Evaluation Measures To evaluate the models we use the
Mean Absolute Error (MAE), Matthews correlation
coefficient
          <xref ref-type="bibr" rid="ref16">(Matthews 1975)</xref>
          that is often regarded as stable
when classes are unbalanced, F1 score, precision, and
recall. When we compute MAE we will compute the absolute
distance between the fuzzy predictions and the actual true
values; this will give us the possibility of understanding how
good are models with a continuous error value. When
computing the other measures we will round the scores to the
nearest integer in such a way that we compare only binary
scores. We consider prediction values higher than 0.5 as 1
and vice-versa. While this is a strong approximation over the
degree of truth given by fuzzy logic it is still useful to
understand the performance of the model. We will also report
accuracy to summarize the performance of the model when
necessary. In general, we select the best model for each
experiment by considering the one with the highest F1 score.
        </p>
      </sec>
      <sec id="sec-2-7">
        <title>Experiment 1: Taxonomy Reasoning</title>
        <p>For the A dataset we ask the LTNs to learn the following
axioms:
8a; b; c 2 A : (sub(a; b) ^ sub(b; c)) ! sub(a; c)
8a 2 A : :sub(a; a)
8a; b : sub(a; b) ! :sub(b; a)
Where sub identifies the subclass relation in the dataset (e.g.,
sub(Cat; F eline)). The objective of this experiment is to
see if LTNs can generate the transitive closure starting from
a dataset using the axioms. Data contained in the A dataset
is our KBS while the edges needed to generate the
transitive closure will be our KBT . We compare the predictions of
LTNs (computed as the prediction over sub(x; y) given x; y)
with the actual transitive closure of the graph. We recall that
KBS contains only true predicates (e.g., sub(Cat; F eline))
while we ask the model to perform inferences also over
predicates that are false (e.g., we evaluate sub(F eline; Cat)
expecting a value close to 0).</p>
        <p>Table 1 shows results for the knowledge completion tasks
of the top performing model and one of the worst
performing ones: the top performing model had a satisfiability equal
to 0.99 while one of the worst ones had a satisfiability of
0.56. The top-performing model was initialized with a layer
size in the tensor network of 20 and a dimension of the
embeddings equal to 20; the best universal aggregator was the
mean aggregator.</p>
        <p>The best model over KB is able to fit well the data
since the F1 measure show good performance over the
entire knowledge base (F1 = 0.64). LTNs are prone to generate
false positives: the model generates 36 false positives with
respect to 55 true positives and 26 false negatives with
respect to 459 true negatives.</p>
        <p>The performance drops when we consider only KBT
elements for testing (F1 = 0.51), this means that LTNs are in
this case not able to capture some more complicated
inferences.</p>
        <p>Still, the approach is better than a binary random
baseline. The accuracy of the model with the best satisfiability is
0.89, while a naive classifier that predicts only zeros would
have reached an accuracy equal to 0.85. This is important to
remark since the two classes are ill-balanced.</p>
        <p>Qualitative Analysis Analyzing the prediction of LTNs
we found that in some cases the model correctly predicts
multi-hop logical inferences (e.g., sub(Cat, Animal) close to
1), but fails on other simple inferences (e.g., sub(Cat, Bird)
close to 1). When there is not enough information regarding
the relationship between two elements (e.g., Cat and Bird)
the model has difficulties to predict the correct answer.</p>
      </sec>
      <sec id="sec-2-8">
        <title>Summary of the outcomes</title>
        <sec id="sec-2-8-1">
          <title>LTNs fit the data well;</title>
        </sec>
        <sec id="sec-2-8-2">
          <title>Multi-hop inferences tend to be more difficult; As expected performance increases with satisfiability.</title>
        </sec>
      </sec>
      <sec id="sec-2-9">
        <title>Experiment 2: Ancestors Reasoning</title>
        <p>For the P dataset we train LTNs with the following axioms:
8a; b 2 P : parent(a; b) ! ancestor(a; b)
8a; b; c 2 P : (ancestor(a; b) ^ parent(b; c)) !
ancestor(a; c)
8a 2 P : :parent(a; a)
8a 2 P : :ancestor(a; a)
8a; b 2 P : parent(a; b) ! :parent(b; a)
8a; b 2 P : ancestor(a; b) ! :ancestor(b; a)
Thus, we combine the knowledge of these axioms with
the data of the parental relationships. We distinguish
two different relationships in this dataset parent (i.e.,
parent(x; y) means that x is a parent of y) and ancestor
(i.e., ancestor(x; y) means that x is a ancestor of y).</p>
        <p>KBS contains only the parental relationships shown in
Figure 2 (e.g., parent(C; I)). The task we will test is to
infer the complete knowledge base for the ancestor predicate,
to which we will refer to as KBa; therefore, we would like
LTNs to learn if an ancestor relationship is true or false for
two given nodes only from axioms and parental data. The
representation for the ancestor predicate should be
generated from the knowledge in the axioms since no data about
it is provided.</p>
        <p>We will also test how the model performs over the set
of ancestor formulas that require multi-hop inferences to
be inferred (i.e., those that cannot be directly inferred from
8a; b 2 P : parent(a; b) ! ancestor(a; b)), we will
refer to this as KBaT : those ancestors pairs for which the
parent pair is false (e.g., ancestor(C; S)). As before, we recall
that KBp contains only true predicates (e.g., parent(C; I))
while we ask the model to perform inferences over the
ancestor dataset (KBa) that also contains predicates that
should be inferred as false (e.g., ancestor(I; C)).</p>
        <p>We do this to understand if LTNs are able to pass the
information from the parent predicate to the ancestor predicate
and if this is enough to give to the network the possibility
of making even more complex inferences that are related to
chains of ancestors.</p>
        <p>The best performing model for this task (with hmean, 10
dimensional embeddings, 10 neural tensor layers) over KBa
had an F1 score of 0.77. If we do not consider the
ancestor predicates that can be directly inferred from the axioms
(KBaT ), the model correctly infers 22 ancestors while
generating 25 false positives: the F1 is equal to 0.62. Again, the
network seems to be able to fit the data quite well, but it still
generates errors on multi-hop inferences.</p>
        <p>As another experiment over satisfiability, in Figure 3 we
show the relation between the MAE computed on KBa and
the level of satisfiability. To draw this figure we run
multiple experiments with LTNs and computed the mean MAE
by aggregating the satisfiability levels rounded to 2 decimal
digit. It is clear that the error decreases with the increase of
the satisfiability level and thus LTNs are able to learn and
infer some knowledge. This proves again that the model is
able to learn the originally not known ancestor relationships
from the combination of data and rules.</p>
        <p>Comparison With Added Axioms To provide a better
understanding of this experiment we decided to add two
axioms to the previous set. These two axioms explicitly state
the relationships between parents and ancestors:
8a; b; c 2 P : (ancestor(a; b) ^ ancestor(b; c)) !
ancestor(a; c)
8a; b; c 2 P
ancestor(a; c)
: (parent(a; b) ^ parent(b; c))
!</p>
        <p>Table 2 shows the comparison between the approach
without the new axioms (Six Axioms) and with the new axioms
(Eight Axioms) on the ancestor dataset. Performances were
computed on the two models with the highest satisfiability
(both around 0.99). The top-performing models for both Six
Axioms and Eight Axioms were initialized with a layer size in
the tensor network of 10 and a dimension of the embeddings
equal to 10; the best universal aggregator was the hmean
aggregator. Results show that the new axioms are beneficial for
the network, that is actually able to learn well the
relationships. Still, the precision over KBaT has increased by 0:19
points (the difference between the results within
parentheses).</p>
        <p>One interesting result about this is related to the fact that
the network is able to learn a good representation for the
ancestor predicate just from the axioms.</p>
        <p>Qualitative Analysis LTNs allows us to do fuzzy
inferences after training. The model is able to answer queries on
fuzzy formulas that were not in the original training data.
For example, 8a; b : ancestor(a; b) ! :parent(b; a) has
generally a value close to 1 in our experiments.</p>
      </sec>
      <sec id="sec-2-10">
        <title>Summary of the outcomes</title>
        <p>Satisfiability is strongly related with performance of the
model: the higher the satisfiability the lower the error;
LTNs learn to pass information quite efficiently
(information on parent(x; y) is passed to ancestor(x; y)). Still,
some more complicated inferences are difficult;
More axioms increase the performance of the model.</p>
      </sec>
      <sec id="sec-2-11">
        <title>Experiment 3: Comparison with a Multi-Input</title>
      </sec>
      <sec id="sec-2-12">
        <title>Network</title>
        <p>In this experiment, we want to compare LTNs with a
simple deep learning architecture on a common task. Starting
from the complete knowledge base of parents and
ancestors we randomly divide data into the training set and test
set. Training data consists of 100 parent predicates (both
true and false) and 100 ancestor predicates4 (both true and
false); test set contains 189 parent predicates and 189
ancestor predicates5. We thus tackle this problem by considering a
classification setting that can be solved with the use of deep
learning models.</p>
        <p>We built a simple multi-input architecture that took as
input three one-hot encoded representations of the pairs of
atoms and the predicates (e.g., Susan, Ann, parent). This is
not the most optimized architecture to solve this task, but it
is useful to understand the performance of LTNs compared
with classical deep learning approaches. We trained the
network using binary cross-entropy and the RMSprop gradient
optimization algorithm over 5,000 epochs with a 20%
validation split. To reduce possible effects of overfitting we use
4Note that the training set contains very few examples that are
positive</p>
        <p>5We tested different random subsets of training and testing, but
the results tend to be similar
L2 regularization (we experimentally found that results were
better with it than without). We show this architecture in
Figure 4 where we also show the dimensions of the layers.</p>
        <p>The network is trained to detect if a predicate, given two
constants, is true or false (binary outcome). LTNs are
implicitly trained on the same task: we train the network over
best satisfiability given the data in input and the six axioms
used in the previous setting.</p>
        <p>The performance of the models is computed over the 189
ancestor test predicates. We ignore the parent predicates in
this setting because there is little to no knowledge about how
to predict if a parental relationship in the test set is true or
false from the dataset.</p>
        <p>Results show that the multi-input network achieves an
accuracy equal to 0.84 while accuracy for LTNs was around
0.89; while the accuracies are comparable an in-depth
analysis with other measures revealed that the recall for the
multiinput was 1 and its precision was 0.12, while LTNs had a
lower recall (0.66) but a much higher precision (0.57). A
naive model that predicts only zeros (since classes are
unbalanced) would have reached an accuracy equal to 0.84. The
multi-input architecture tends to overfit in this task in which
most of the classes are 0. It is anyway important to note that
it is difficult for the multi-input architecture to understand
the task, while LTNs are helped by the axioms.</p>
        <p>However, the results show that while LTNs are good for
learning logical rules, their accuracy is still comparable to
the one obtained by neural-networks. Moreover, the
multiinput architecture would require more control on overfitting,
while the logical axioms used in LTNs seems to provide a
natural way to define some constraints over the vector space
and to reduce possible overfitting. Nevertheless, different
deep learning architectures with a different set of
parameters might generate better results.</p>
        <p>Using classical neural networks we lose the ability to
define high-level semantics to the data. For example, LTNs can
be used in combination with quantifiers to make inferences
over data using new axioms on which the network was not
trained (e.g., 8x; y : parent(x; y) ! :ancestor(y; x) has
a high truth value).</p>
        <p>
          As shown in the recent work on LTNs on semantic image
interpretation one key element of success might be the use
of LTNs over deep learning architectures
          <xref ref-type="bibr" rid="ref2 ref6">(Donadello,
Serafini, and d’Avila Garcez 2017)</xref>
          ; this would allow
augmenting data with semantic information that will make it possible
to explain predictions.
        </p>
      </sec>
      <sec id="sec-2-13">
        <title>Summary of the outcomes</title>
        <p>Results show that performance on this simple task is
comparable to a naive network;
Axioms in LTNs seem to provide a useful way of
defining constraints over the space of the solutions that might
reduce the possibility of overfitting;
The main advantage of LTNs resides in the possibility of
making inferences after training.</p>
      </sec>
      <sec id="sec-2-14">
        <title>Experiment 4: Time to Learn</title>
        <p>In this last experiment, we investigate how fast LTNs are in
the learning context. We consider the following
experimental setting: we generate a range of N constant and N
predicates and we evaluate different combinations of them. We
divide this experiment in three by considering unary, binary
and ternary predicates of the following from 8x : predn(x),
8x; y : predn(x; y), 8x; y; z : predn(x; y; z), we therefore
test only predicates that are universally quantified. We
compute 5,000 training epochs to learn the parameters of 4, 8, 12,
20, 30 constants with 4, 8, 12, 20, 30 (universally quantified)
predicates of arity one, two and three: this means that in the
setting with 4 constants and 8 predicates of arity 3 we
introduce 4 constants (a; b; c; d) in the model and 8 predicates
(pred1; pred2; : : : ; pred8) and each predicate is universally
quantified (e.g., 8x; y; z : pred1(x; y; z)). Size of the
embedded representation in this experiment is 10. Experiments
were run using a compiled version of Tensorflow on an i7
machine.</p>
        <p>Analysis Figures 5, 6, 7 show the seconds needed to
complete the learning phase for each setting. While it is clear that
constants have an influence on computational time (since
they are training data) we can also state that predicates and
their arity have a notable computational impact upon the
learning phase. With a low number of constants and
predicates (e.g., 4) the training time is not much different in all
the settings, but as soon as the number of constants increases
the model requires more time to learn. The arity of the
predicates seem to be the element with the higher impact on the
learning time: this is an expected result since the universal
quantifier has to cover multiple elements in the ternary case.
Since experiments were run on a CPU we expect training
time to be shorter on GPU 6.</p>
        <p>Time to learn the parameters is highly influenced by the
arity of the predicates;</p>
      </sec>
      <sec id="sec-2-15">
        <title>Other Experimental Notes</title>
        <p>In this section, we briefly describe other experimental results
that are interesting for the community. While the
following assertions are derived from empirical experiments they
might still be useful for the reader who wants to start using
LTNs.</p>
        <p>LTNs as all deep learning model suffers from
optimization problems: in our experiments we often found the model
reaching local minima. Global optimization tools might help
in a better parameter optimization search.</p>
        <p>In our experiments LTNs often predicted the class Cat to
be a subclass of the class Bird. This error might be due to
missing knowledge in the KB. The network is not able to
understand the difference between the two since they come
from different branches of the taxonomy. In general, it seems
that LTNs predict many false positives, while they are better
in detecting true negatives. This seems due to the fact that
true negatives in our experiments can be directly inferred
from the axioms: for example, 8a : :ancestor(a; a) gives a
good amount of information to the model about the fact that
each constant occurring in both parameters of the predicate
ancestor should generate negative values.</p>
        <p>If the model fits the data too well (i.e., it overfits) the
performance over the test set decreases. While this is a common
event for machine learning models and there are techniques
to prevent this, applying these to LTNs is not so
straightforward: cross-validation would require us to provide
completeness information to the training set, that would bias the
reasoning task.</p>
        <p>We tested different sets of hyper-parameters and we
release results on the tested tasks online. While this was not
the primary scope of the paper it is still important to
estimate the effects of the hyper-parameters to fully evaluate
the approach. Nevertheless, we empirically find out that
increasing the layers of the tensor network and the size of the
embeddings makes the model much more difficult to
optimize.</p>
        <p>After paper acceptance a new version of LTNs was
released by the original authors: this last version is easier to
optimize and shows a slight increase in performance over
the F1 measure.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>In this section we summarize some related approaches that
have been introduced in the state of the art. We refer
to Garcez, Lamb, and Gabbay; Besold et al. for discussions
6to show an effective comparison between different predicates
we decided to show results computed with a CPU: with the GPU
it was more difficult to highlight the differences between these
experiments
of different neural-symbolic approaches proposed in
literature: in this section we will only discuss a few of these
approaches and we will also describe some related methods.</p>
      <p>
        One of the main points of discussion that has involved the
artificial community in the last decades is the relationship
between symbolic artificial intelligence and connectivist
(i.e., related to neural networks) artificial intelligence
        <xref ref-type="bibr" rid="ref18">(Minsky 1991)</xref>
        . In recent years deep learning approaches have
shown great computational capabilities
        <xref ref-type="bibr" rid="ref11 ref21">(Goodfellow,
Bengio, and Courville 2016)</xref>
        , but still these approaches do not
achieve the same reasoning and knowledge transformation
abilities that symbolic approaches show. On the other hand,
symbolic artificial intelligence suffers from computational
limits and the knowledge acquisition bottleneck, i.e., the
need to generate high-quality knowledge bases, which is
usually done manually. A different voice in this group comes
from the neural-symbolic field, where the task is to bring
together the two worlds of symbolic artificial intelligence and
neural networks
        <xref ref-type="bibr" rid="ref10 ref13 ref8 ref9">(Garcez, Gabbay, and Broda 2002;
Hammer and Hitzler 2007; Garcez, Lamb, and Gabbay 2008;
Garcez et al. 2015)</xref>
        .
      </p>
      <p>
        In the current work we have explored only LTNs, but
there are different approaches in the field that have been
introduced. One of the most famous approaches of
neuralsymbolic integration are the Knowledge Based Artificial
Neural Networks (KBANNs)
        <xref ref-type="bibr" rid="ref15 ref24">(Towell and Shavlik 1994)</xref>
        .
KBANNs where one of the first approaches to integrate
propositional clauses with data, developed at the same time
as the closely related propositional core method
        <xref ref-type="bibr" rid="ref15 ref24">(Ho¨ lldobler
and Kalinke 1994)</xref>
        . Lifting these results towards first-order
logic, however, has been proven difficult and limited to
toysize knowledge bases
        <xref ref-type="bibr" rid="ref1 ref12 ref13 ref14">(Hitzler, Ho¨ lldobler, and Seda 2004;
Gust, Ku¨ hnberger, and Geibel 2007; Bader, Hitzler, and
Ho¨ lldobler 2008)</xref>
        .
      </p>
      <p>
        On the other hand, there are other approaches from the
Statistical Relational Learning field that do not integrate
neural networks with logic, but tackle the problem in a
symbolic manner by also combining statistical information.
Examples of this category are ProbLog
        <xref ref-type="bibr" rid="ref5">(De Raedt and Kimmig
2015)</xref>
        that is an example of probabilistic logic programming
language and Markov Logic Networks (MLNs) are a
statistical relational learning model that has been shown to be
effective on a large variety of tasks
        <xref ref-type="bibr" rid="ref17 ref20">(Richardson and
Domingos 2006; Meza-Ruiz and Riedel 2009)</xref>
        . The intuition
behind MLNs and LTNs is similar since they both base their
approach on logical languages. MLNs defines weights for
formulas and interpret the world by considering it under a
probabilistic point of view while LTNs use fuzzy logic
combined with a neural architecture to generate their inferences.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <sec id="sec-4-1">
        <title>Conclusions</title>
        <p>LTNs can be shown to obtain good results on reasoning tasks
when optimal satisfiability conditions are met. This is
often difficult to reach and using the model with a low degree
of satisfiability can generate bad inferences. Nevertheless,
LTNs show interesting capabilities and their ability to mix
logic and data might prove to be a valuable resource. LTNs
1e+02
)
itry3a 8 11 39 1.1e+02 5.1e+02 1.7e+03
(
s
e
t
ifrrcbdaeem2102 2135 5868 12..66ee++0022 71..32ee++0023 24..51ee++0033
p
o
u
N
30 34 1.3e+02 4.1e+02 2e+03 6.5e+03
4
8</p>
        <p>12
Number of constants
20
30
4500
3000
1500
fit well the data and can be used to make some simple
inferences. More complex inferences (multi-hop) are more
difficult to capture in the model.</p>
        <p>The main problem encountered in our experiments is
related to erroneous prediction generated by the LTNs and
scalability issues. We think that the former problem might
be solved with a more accurate use of logic constraints: for
example, in the ancestor experiment, adding notions about
the concepts of “siblings” might help the network to perform
better. While a more efficient use of computational resources
could help in reducing the latter problem we encountered.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Future Work</title>
        <p>While results have shown that LTNs are able to capture logic
semantics in the vector space, they should also be compared
with other statistical relational learning methods like MLNs
on similar tasks.</p>
        <p>
          Another possible next step is to apply LTNs on bigger
knowledge bases defined in the state of the art
          <xref ref-type="bibr" rid="ref3">(Bordes et
al. 2013)</xref>
          . We expect the ability to make fuzzy inferences
over the trained model to be of great help in link prediction
tasks over knowledge bases.
        </p>
        <p>
          An interesting development of this work could be
evaluating the generated groundings: constants in LTNs have
an associated vector and thus it is possible to compute the
similarity in the vector space between constants. This might
be interesting in the context of knowledge graph
embeddings
          <xref ref-type="bibr" rid="ref3">(Bordes et al. 2013)</xref>
          : vector representations of entities
and relationships of a knowledge graph.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>We thank Luciano Serafini and Artur d’Avila Garcez for
their comments and suggestions. We gratefully acknowledge
the support of NVIDIA Corporation with the donation of the
Titan Xp GPU used for this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bader</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Hitzler,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ; and Ho¨lldobler,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>Connectionist model generation: A first-order approach</article-title>
          .
          <source>Neurocomputing</source>
          <volume>71</volume>
          (
          <fpage>13</fpage>
          -15):
          <fpage>2420</fpage>
          -
          <lpage>2432</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Besold</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          R.;
          <string-name>
            <surname>d'Avila Garcez</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bader</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Bowman,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ; Domingos,
          <string-name>
            <given-names>P. M.</given-names>
            ;
            <surname>Hitzler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ; Ku¨hnberger, K.;
            <surname>Lamb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            ;
            <surname>Lowd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ; Lima, P. M. V.; de Penning, L.;
            <surname>Pinkas</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          ; Poon, H.; and Zaverucha,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Neural-symbolic learning and reasoning: A survey and interpretation</article-title>
          .
          <source>CoRR abs/1711</source>
          .03902.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Usunier</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Garcia-Duran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Weston</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Yakhnenko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Translating embeddings for modeling multi-relational data</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>2787</volume>
          -
          <fpage>2795</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S. R.</given-names>
          </string-name>
          ; Potts,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. D.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Learning distributed word representations for natural logic reasoning</article-title>
          .
          <source>In Proceedings of the Association for the Advancement of Artificial Intelligence Spring Symposium (AAAI)</source>
          ,
          <fpage>10</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>De Raedt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Kimmig</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Probabilistic (logic) programming concepts</article-title>
          .
          <source>Machine Learning</source>
          <volume>100</volume>
          (
          <issue>1</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Donadello</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>d'Avila Garcez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>Logic tensor networks for semantic image interpretation</article-title>
          .
          <source>In IJCAI</source>
          ,
          <fpage>1596</fpage>
          -
          <lpage>1602</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Garcez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Besold</surname>
            ,
            <given-names>T. R.</given-names>
          </string-name>
          ; De Raedt,
          <string-name>
            <surname>L.</surname>
          </string-name>
          ; Fo¨ldiak, P.; Hitzler,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ; Icard,
          <string-name>
            <given-names>T.</given-names>
            ; Ku¨hnberger, K.-U.;
            <surname>Lamb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            ;
            <surname>Miikkulainen</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Silver</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Neural-symbolic learning and reasoning: contributions and challenges</article-title>
          .
          <source>In Proceedings of the AAAI Spring Symposium on Knowledge Representation and Reasoning: Integrating Symbolic and Neural Approaches</source>
          , Stanford.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Garcez</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          <year>d</year>
          .;
          <string-name>
            <surname>Gabbay</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Broda</surname>
            ,
            <given-names>K. B.</given-names>
          </string-name>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Garcez</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lamb</surname>
            ,
            <given-names>L. C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Gabbay</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ; Bengio,
          <string-name>
            <given-names>Y.</given-names>
            ; and
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Deep Learning</article-title>
          . MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Gust</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <article-title>Ku¨hnberger, K.;</article-title>
          and Geibel,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2007</year>
          .
          <article-title>Learning models of predicate logical theories with neural networks based on topos theory</article-title>
          . In Hammer,
          <string-name>
            <given-names>B.</given-names>
            , and
            <surname>Hitzler</surname>
          </string-name>
          , P., eds.,
          <source>Perspectives of Neural-Symbolic Integration</source>
          , volume
          <volume>77</volume>
          <source>of Studies in Computational Intelligence</source>
          . Springer.
          <fpage>233</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Hammer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hitzler</surname>
          </string-name>
          , P., eds.
          <source>2007. Perspectives of Neural-Symbolic Integration</source>
          , volume
          <volume>77</volume>
          <source>of Studies in Computational Intelligence</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; Ho¨lldobler, S.; and
          <string-name>
            <surname>Seda</surname>
            ,
            <given-names>A. K.</given-names>
          </string-name>
          <year>2004</year>
          .
          <article-title>Logic programs and connectionist networks</article-title>
          .
          <source>J. Applied Logic</source>
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <fpage>245</fpage>
          -
          <lpage>272</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>Ho¨lldobler, S., and</article-title>
          <string-name>
            <surname>Kalinke</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Ein massiv paralleles modell fu¨r die logikprogrammierung</article-title>
          .
          <source>In WLP</source>
          ,
          <fpage>89</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Matthews</surname>
            ,
            <given-names>B. W.</given-names>
          </string-name>
          <year>1975</year>
          .
          <article-title>Comparison of the predicted and observed secondary structure of t4 phage lysozyme</article-title>
          .
          <source>Biochimica et Biophysica Acta (BBA)-Protein Structure</source>
          <volume>405</volume>
          (
          <issue>2</issue>
          ):
          <fpage>442</fpage>
          -
          <lpage>451</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Meza-Ruiz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Riedel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Jointly identifying predicates, arguments and senses using markov logic</article-title>
          .
          <source>In NAACL</source>
          ,
          <fpage>155</fpage>
          -
          <lpage>163</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Minsky</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          <year>1991</year>
          .
          <article-title>Logical versus analogical or symbolic versus connectionist or neat versus scruffy</article-title>
          .
          <source>AI</source>
          magazine
          <volume>12</volume>
          (
          <issue>2</issue>
          ):
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Petr</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>Metamathematics of fuzzy logic, vol. 4 of trends in logicstudia logica library</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Richardson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Domingos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Markov logic networks</article-title>
          .
          <source>Machine learning 62(1-2)</source>
          :
          <fpage>107</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Garcez</surname>
          </string-name>
          , A. S. d.
          <year>2016</year>
          .
          <article-title>Learning and reasoning with logic tensor networks</article-title>
          .
          <source>In Conference of the Italian Association for Artificial Intelligence</source>
          ,
          <fpage>334</fpage>
          -
          <lpage>348</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C. D.; and
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>Reasoning with neural tensor networks for knowledge base completion</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>926</volume>
          -
          <fpage>934</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Towell</surname>
            ,
            <given-names>G. G.</given-names>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>Shavlik</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Knowledgebased artificial neural networks</article-title>
          .
          <source>Artificial intelligence</source>
          <volume>70</volume>
          (1
          <issue>- 2</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>