<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scenarios Interpretation with Prior Knowledge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Daniele</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Advisor: Luciano Serafini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Florence</institution>
          ,
          <addr-line>Florence</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Statistical Relational Learning (SRL) deals with relational domains, where the samples are neither independent nor uniformly distributed. Moreover, central to SRL is the integration of logical knowledge in the learning framework. The main tasks in SRL are Collective Classification, Entity Resolution, Link Prediction and Knowledge Graph Completion. In this extended abstract we propose a new supervised learning task called Scenarios Interpretation (SI) where a sample is a Scenario, i.e. a set of (typically few) objects where each object and pair of objects have its own features. The goal is to classify objects and relationships. We propose NIoS (Neural Interpeter of Scenarios), a method for solving SI that is able to inject Prior Knowledge expressed in First Order Logic (FOL) into a neural network model. We implemented a first version and tested it on Visual Relationship Detection task (VRD) showing that NIoS outperformed state of the art systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>hsubj, rel, obji. Together with visual data we have some prior knowledge
expressed as logic formulas (e.g. W ear(x, y) → P erson(x)).</p>
      <p>
        Among the SRL approaches that exploit logical knowledge, there are Logic
Tensor Network (LTN) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Semantic Based Regularization (SBR) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Markov
Logic Network (MLN) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We propose NIoS, a method for injecting FOL clauses
inside a neural network model that can deal with the graph structure of scenarios.
The main difference with its major competitors is on the way logic formulas are
used: in NIoS they become part of the predictors instead of being used during
training. In particular, methods like LTN or SBR force the constraint satisfaction
during training making the assumption that the knowledge is in general correct.
Instead, we assume there is a relationship between clauses and correct results,
but this relationship is not known. The logical constraint are seen as a Prior
Belief rather than Prior Knowledge. More in details, NIoS has internal learnable
parameters associated to the logic formulas. In this extent, the most similar
approaches to ours are probably (Hybrid) Markov Logic Networks [
        <xref ref-type="bibr" rid="ref10 ref7">10, 7</xref>
        ] and
Probabilistic Soft Logic (PSL) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] where the formulas weights are learned.
      </p>
      <p>
        We tested NIoS for the Predicate Detection subtask of VRD on the Visual
Relationship Dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (VRD Dataset) where we outperformed state of the art
results, in particular on the Zero Shot Learning evaluation.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Scenarios Interpretation task</title>
      <p>A scenario S ∈ S is a triple composed of a set of objects O and two functions
u : O → Rk and b : O × O → Rl. An interpretation of a scenario S is a pair
I = hlo, lri</p>
      <p>
        lo : O × C → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] lr : O × R × O → [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ]
where C and R are two disjoint sets of symbols for classes and relations
respectively. The set of all interpretation is I, the set of interpretations of a particular
scenario S is IS . A constraint is a clause in First Order Logic where unary
predicates are in C and binary predicates are in R.
      </p>
      <p>Let I∗ : S → I be a function that returns a correct interpretation of a scenario
(I∗(S) ∈ IS ). Given a training set composed of scenarios and corresponding
correct interpretation S(i), I∗(S(i)) in=1 and a tuple of clauses K = hc1, . . . , cmi
representing the Prior Knowledge, the SI task is the problem of finding a function
that predicts correct interpretations of unseen scenarios. In particular, we want
to find a function I˜K parametrized by weights of clauses in K, that given a
Scenario returns an interpretation that minimize the error in the training set:
n
IK = argmin X
˜</p>
      <p>IK i=1</p>
      <p>L(IK(S(i)), I∗(S(i)))
(1)
where L : IS × IS → R+ is a function that returns a similarity score of two
Interpretations. In our first implementation we used the L2 loss function:
L(Ip, It) =</p>
      <p>X(lop(o) − lot(o))2 +</p>
      <p>(lrp(o1, r, o2) − lrt(o1, r, o2))2
oc∈∈OC</p>
      <p>X
(o1,o2)∈O2</p>
      <p>r∈R
where It and Ip are the true and predicted interpretations.
3</p>
    </sec>
    <sec id="sec-3">
      <title>NIoS: overview of the model</title>
      <p>NIoS (Neural Interpreter of Scenarios) is a method for injecting logical knowledge
into a Neural Network (NN). The original NN takes a Scenario as input and
returns an initial Interpretation that is changed by a function, called Global
Enhancer (GE), that modifies the initial predictions by enforcing the satisfaction
of the logical constraints. The function must be differentiable and it can be
seen as a new final layer for the original neural network. The entire network is
still differentiable end-to-end, making it possible to train the model with
backpropagation algorithm. GE contains additional parameters that can be learned
as well. In particular clause weights determine the strength of each clause.</p>
      <p>wA∨¬B
δAc1 δBc1 δCc1 δDc1
δc1
z</p>
      <p>CE
c1: A ∨ ¬B
y yA yB yC yD
σ
wC∨D
δAc2 δBc2 δc2 δDc12</p>
      <p>C
δc2
z</p>
      <p>CE
c2: C ∨ D
z zA zB zC zD
(a)
1</p>
      <p>GE</p>
      <p>δzc1 δAc1 δBc1 0 0</p>
      <p>CE
c1: A ∨ ¬B</p>
      <p>concat
δAc1 δc1</p>
      <p>B
1 -1
δAc1 δ¬c1B
softmax
zA z¬B
1 -1
z zA zB zC zD
(b)
the softmax function to the literals values. Intuitively, the idea is that, in order
to satisfy a clause, at least one of its literal must be true. The softmax function
act as a selector for the most promising true literal, that is the one with higher
supporting evidences (biggest preactivation). Finally there is a post-elaboration
step that works in reverse of the pre-elaboration (it sets the absent predicates
adjustments to zero and change the sign of the negated ones).
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental evaluation</title>
      <p>
        Visual Relationship Detection (VRD) is the task of finding objects in an image
and capture their interactions [
        <xref ref-type="bibr" rid="ref13 ref3 ref5">3, 5, 13</xref>
        ]. It is composed of three subtask:
Relationship Detection, Phrase Detection and Predicate Detection [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The VRD
Dataset contains 100 classes for objects and 70 types of relations. It is composed
of 4000 images for training and 1000 for testing with a total of 6672 triplets
types. Among them 1877 can be find only in the Test Set and predicting them is
the goal of the Zero Shot Learning variant of the task. For evaluating the results
we used the Recall@n (n ∈ {50, 100}) metric proposed by Lu et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that is
the percentage of times a correct relationship is found on the n predictions with
highest score.
      </p>
      <p>
        We evaluated NIoS on the Predicate Detection task using the knowledge
base of [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We implemented NIoS using TensorFlow. As original NN we used a
neural network with zero hidden layers and trained the entire network (original
NN + GE) end-to-end using RMSProp [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Results are shown in Table 1.
      </p>
      <sec id="sec-4-1">
        <title>Standard L.</title>
        <p>
          R@50 R@100
Lu et al.[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] 47.87 47.87
LTN[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] 78.63 91.88
Yu et al.[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] 85.64 94.65
NIoS 86.02 91.91
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Zero Shot L.</title>
        <p>R@50 R@100
8.45 8.45
46.28 70.15
54.20 74.65
68.95 83.83</p>
        <p>
          NIoS outperformed other methods on all the metrics except for Recall@100
where it is surpassed by Yu et Al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. The best results can be seen on the
Zero Shot Learning task, where the difference between NIoS and the second best
system is more than 10%. In Zero Shot Learning the aim is to predict previously
unseen triplets, therefore it is rather difficult to learn to predict them from the
Training Set. This confirms the ability of NIoS to use the Knowledge Base.
        </p>
        <p>
          Another interesting result is the value obtained by NIoS compared to LTN[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
In particular considering that the two works used the same Prior Knowledge.
A possible explanation is given by the ability of NIoS to learn clause weights.
Indeed, many weights results to be zero after learning. An example of a zero
weighted clause is: ¬Ride(x, y) ∨ On(x, y).
        </p>
        <p>Although the rule seems correct it is not in general satisfied on training and
test set. This is because labels have been added manually, therefore there are
plenty of missing relations. The hypothesis is that people have a tendency to
add the most informative labels making some of the clauses unsatisfied.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We proposed SI, a new SRL task where the goal is to predict an entire graph,
and we developed NIoS, a method for solving SI that can deal with learning
in presence of a FOL Prior Knowledge. We reframed the VRD task as a SI
instance and evaluated NIoS on it. With its results on VRD, NIoS showed to
be competitive against other approaches, in particular tanks to its ability to
effectively learn clauses weights.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>S.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Broecheler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Getoor</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Hinge-loss Markov random fields and probabilistic soft logic</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>18</volume>
          (
          <issue>109</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Diligenti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gori</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Sacca`,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Semantic-based regularization for learning and inference</article-title>
          .
          <source>Artif. Intell</source>
          .
          <volume>244</volume>
          ,
          <fpage>143</fpage>
          -
          <lpage>165</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Donadello</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Semantic Image Interpretation - Integration of Numerical Data and Logical Knowledge for Cognitive Vision</article-title>
          .
          <source>Ph.D. thesis</source>
          , Trento Univ.,
          <string-name>
            <surname>Italy</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hasan</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chaoji</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salem</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Link prediction using supervised learning</article-title>
          .
          <source>In: In Proc. of SDM 06 workshop on Link Analysis, Counterterrorism and Security</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krishna</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Visual relationship detection with language priors</article-title>
          .
          <source>In: ECCV (1). Lecture Notes in Computer Science</source>
          , vol.
          <volume>9905</volume>
          , pp.
          <fpage>852</fpage>
          -
          <lpage>869</lpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phung</surname>
            ,
            <given-names>D.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkatesh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Column networks for collective classification</article-title>
          .
          <source>In: AAAI</source>
          . pp.
          <fpage>2485</fpage>
          -
          <lpage>2491</lpage>
          . AAAI Press (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Richardson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Domingos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Markov logic networks</article-title>
          .
          <source>Mach. Learn</source>
          .
          <volume>62</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>107</fpage>
          -
          <lpage>136</lpage>
          (
          <year>Feb 2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Serafini</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>d'Avila Garcez</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>Logic tensor networks: Deep learning and logical reasoning from data and knowledge</article-title>
          .
          <source>CoRR abs/1606</source>
          .04422 (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Tijmen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          ,
          <source>G.: Lecture 6</source>
          .5
          <article-title>− rmsprop: Divide the gradient by a running average of its recent magnitude</article-title>
          .
          <source>COURSERA: Neural Networks for Machine Learning</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Domingos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Hybrid markov logic networks</article-title>
          .
          <source>In: Proceedings of the 23rd National Conference on Artificial Intelligence. AAAI'08</source>
          , vol.
          <volume>2</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choy</surname>
            ,
            <given-names>C.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Scene graph generation by iterative message passing</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . vol.
          <volume>2</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morariu</surname>
            ,
            <given-names>V.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>L.S.:</given-names>
          </string-name>
          <article-title>Visual relationship detection with internal and external linguistic knowledge distillation</article-title>
          .
          <source>CoRR abs/1707</source>
          .09423 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , H.,
          <string-name>
            <surname>Kyaw</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>S.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chua</surname>
          </string-name>
          , T.S.:
          <article-title>Visual translation embedding network for visual relation detection</article-title>
          .
          <source>2017 IEEE Conference on Computer Vision</source>
          and Pattern
          <string-name>
            <surname>Recognition</surname>
          </string-name>
          (CVPR) pp.
          <fpage>3107</fpage>
          -
          <lpage>3115</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>