<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Review on Intelligent Object Perception Methods Combining Knowledge-based Reasoning and Machine Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Filippos Gouidis</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>y Alexandros Vassiliades</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>y Theodore Patkos</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonis Argyros</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nick Bassiliades</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitris Plexousakis</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aristotle University of Thessaloniki</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Computer Science, Foundation for Research and Technology</institution>
          ,
          <addr-line>Hellas</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>tion for Research and Innovation (HFRI) and the General Secretariat for Research and Technology (GSRT), under grant agreement No 188. Copyright c 2020 held by the author(s). In A. Martin, K. Hinkelmann</institution>
          ,
          <addr-line>H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.)</addr-line>
          ,
          <institution>Proceedings of the AAAI 2020 Spring Symposium on Combining Machine Learning and Knowledge Engineering in Practice (AAAI-MAKE 2020). Stanford University</institution>
          ,
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Object perception is a fundamental sub-field of Computer Vision, covering a multitude of individual areas and having contributed high-impact results. While Machine Learning has been traditionally applied to address related problems, recent studies also seek ways to integrate knowledge engineering in order to expand the level of intelligence of the visual interpretation of objects, their properties and their relations with the environment. In this paper, we attempt a systematic investigation of how knowledge-based methods contribute to diverse object perception tasks. We review the latest achievements and identify prominent research directions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Despite the recent sweep of progress in Machine Learning
(ML) which is stirring public imagination about the
capabilities of future Artificial Intelligence (AI) systems, the
research community seems more composed, in part due to the
realization that current achievements are largely based on
engineering advancements, and only partly on novel
scientific progress. Undoubtedly, the accomplishments are
neither small nor temporary; in the latest “One Hundred Year
Study on AI”, a panel of renowned AI experts is
foreseeing tremendous impact of AI in a multitude of technological
and societal domains in the next decades, fueled primarily
by systems running ML algorithms
        <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Stone et al. 2016)</xref>
        . Yet,
there is still a lot of ground to cover, before we can obtain
a deep understanding of how to overcome the limitations of
data-driven approaches at a more generic level.
      </p>
      <p>Aiming at exploiting the full potential of AI, a
growing body of research is devoted to the idea of integrating
ML and knowledge-based approaches. Davies and Marcus
(2015) for instance, while discussing the multifaceted
challenges related to automating commonsense reasoning, a
crucial ability for any intelligent entity operating in real-world
conditions, underline the need to combine the strengths of
diverse AI approaches from these two fields. Others, as
for example Bengio et al. (2019) and Pearl (2018),
emphasize the inability of Deep Learning to effectively
recognize cause and effect relations. Pearl, Geffner (2018)
and recently Lenat1, suggest to seek solutions by
bridging the gap between model-free, data-intensive learners and
knowledge-based models and by building on the synergy
between heuristic level and epistemological level languages.</p>
      <p>In this paper, we review recent progress in the direction of
coupling the strengths of ML and knowledge-based
methods, focusing our attention on the topic of Object
Perception (OP), an important sub-field of Computer Vision (CV).
Tasks related to OP are at the core of a wide spectrum of
practical systems and relevant research has traditionally
relied on ML to approach the related problems. The recent
developments have significantly advanced the field, but,
interestingly, state-of-the-art studies try to integrate symbolic
methods, in order to achieve broader visual intelligence. It
seems that it is becoming less of a paradox within the CV
community that in order to build intelligent vision systems,
much of the information needed is not directly observable.</p>
      <p>
        Existing surveys on the intersection of ML and knowledge
engineering (e.g.,
        <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Nickel et al. 2016)</xref>
        ) are indeed very
informative, but they usually offer a high-level understanding of
the challenges involved. The rich literature on CV reviews,
on the other hand, adopts a more problem-specific analysis,
studying in detail the requirements of each particular CV
task (see e.g.,
        <xref ref-type="bibr" rid="ref1 ref100 ref101 ref103 ref12 ref14 ref16 ref3 ref34 ref35 ref36 ref37 ref4 ref40 ref42 ref43 ref46 ref49 ref5 ref52 ref53 ref56 ref58 ref60 ref65 ref66 ref68 ref7 ref70 ref78 ref84 ref85 ref86 ref91 ref97 ref98">(Wu et al. 2017; Herath, Harandi, and Porikli
2017; Liu et al. 2018a)</xref>
        ). Only recently was an overview
presented that shows how background knowledge can benefit
tasks, such as image understanding
        <xref ref-type="bibr" rid="ref2 ref24 ref25 ref72 ref73 ref88">(Aditya, Yang, and Baral
2019)</xref>
        ; our goal is to explore this direction on the topic of
intelligent OP, reporting state-of-the-art achievements and
showing how different facets of knowledge-based research
can contribute to addressing the rich diversity of OP tasks.
      </p>
      <p>The rest of the paper investigates state-of-the-art
literature on intelligent OP along the three pillars shown in
Fig1https://towardsdatascience.com/statistical-learning-andknowledge-engineering-all-the-way-down-1bb004040114
ure 1): i) symbolic models, further analyzed from the
perspectives of expressive representations, reasoning capacity,
and open domain, Web-based knowledge; ii) commonsense
knowledge exploitation, a key skill for any intelligent
system; and, iii) enhanced learning ability, building on hybrid
approaches. By no means should this review be considered
exhaustive; inevitably, relevant studies may not have been
included. Our intention is to capture a snapshot of the most
recent research trends, discussing studies that improve
previous results or offer new insights, hoping to provide a
starting point for following the interesting progress currently
underway, while giving pointers to directions that seem worth
investigating. In this respect, the paper concludes with a
discussion on open questions and prominent research
directions. Table 1 at the end summarizes the reviewed literature.</p>
    </sec>
    <sec id="sec-2">
      <title>Exploitation of Symbolic Models</title>
      <p>The scope of OP research ranges over a wide spectrum of
problems, from object, action and affordance detection, to
localization and recognition in images, to motion and
structure inference in videos, to scene understanding and visual
reasoning. Traditionally, OP relied on ML methodologies to
find patterns in realms of data, taking as input feature
vectors representing entities in terms of numeric or categorical
(membership to a more general class) attributes.</p>
      <p>In this section, we discuss how high-level knowledge
related to visual entities can improve the performance of OP
algorithms in manifold ways. We start by reviewing
state-ofthe-art in coupling data-driven approaches and rich
knowledge representations about aspects such as context, space
and affordances.</p>
      <sec id="sec-2-1">
        <title>Expressive Representations of Knowledge</title>
        <p>
          Although the notion of a representation is rather general,
the modeling of relational knowledge in the form of
individuals (entities) and their associated relations is
wellstudied in AI, especially in the context of symbolic
representations (see for instance
          <xref ref-type="bibr" rid="ref10 ref89">(van Harmelen et al. 2007;
Brachman and Levesque 2004)</xref>
          ). The goal is to offer the level
of abstraction needed to design a system versatile enough to
adapt to the requirements of a particular domain, yet rigid
enough to be encoded in a computer program, having clear
semantics and elegant properties
          <xref ref-type="bibr" rid="ref21">(Dean, Allen, and
Aloimonos 1995)</xref>
          . As a result, expressiveness, i.e., what can or
cannot be represented in a given model, and computation,
i.e., how fast conclusions are drawn or which statements can
be evaluated by algorithms guaranteed to terminate, are two,
often competing, aspects to be taken into consideration.
        </p>
        <p>
          The form of the representational model applied plays
a decisive role on such considerations, affecting the
richness of semantics that can be captured. While the range
of models varies significantly, even relatively shallow
representations have proven to offer improvements in the
performance of OP methodologies (see for example
          <xref ref-type="bibr" rid="ref104 ref105 ref107 ref11 ref22 ref27 ref3 ref34 ref44 ref49 ref58 ref65 ref66 ref67 ref68 ref70 ref83 ref86 ref90">(Deng et
al. 2014; Zhu, Fathi, and Fei-Fei 2014; Zhu et al. 2015b;
Redmon and Farhadi 2017)</xref>
          ). Models as simple as relational
tables or flat weighted graphs, but also more complex
multirelational graphs, often called Knowledge Graphs (KGs), or
even semantically-rich conceptualizations with formal
semantics, often called ontologies, are proposed in the relevant
literature. In the sequel, we refer to any model that offers at
least a basic structuring of data as a Knowledge Base (KB).
Utilization of Contextual Knowledge Context awareness
is the ability of a system to understand the state of the
environment, and to perceive the interplay of the entities
inhabiting it. This feature offers advantages in accomplishing
inference tasks, but also enhances a system in terms of
reusability, as contextual knowledge enables it to adapt to new
situations and environments that resemble known ones.
        </p>
        <p>
          A preliminary approach towards this direction utilized
a special semantic fusion network, which combined novel
object- and scene-based information and was able to capture
relationships between video class labels and semantic
entities (objects and scenes)
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Wu et al. 2016b)</xref>
          . This was a
Convolutional Neural Network (CNN) consisting of three layers,
with the first layer detecting low-level features, the second
layer object features and the third layer scene features.
Although no symbolic representation was applied, this
modeling of abstraction layers for the representation of the
information depicted in an image is similar to the hierarchical
structuring of knowledge used by top-down methods.
        </p>
        <p>In a similar style, Liu et al. (2018b) recently attempted to
address the problem of object detection by proposing an
algorithm that exploits jointly the context of a visual scene
and an object’s relationships. Such features are typically
taken into consideration in isolation by most object
detection methods. The algorithm uses a CNN-based framework
tailored for object detection and combines it with a
graphical model designed for the inference of object states. A
special graph is created for each image in the training set,
with nodes corresponding to objects and edges to object
relationships. The intuition behind this approach is that the
graph’s topology enables an object state to be determined
not only by its low-level characteristics (appearance details),
but also by the states of other objects interacting with it and
by the overall scene context. The experimental evaluation
conducted on different datasets underscored the importance
of knowledge stemming from local and global context.</p>
        <p>
          Currently, a collection of prominent approaches are
oriented towards Gated Graph Neural Networks (GGNNs)
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Li
et al. 2015)</xref>
          , a variation of Gated Neural Networks (GNNs)
          <xref ref-type="bibr" rid="ref81">(Scarselli et al. 2008)</xref>
          , in order to integrate contextual
knowledge in the training of a system. GNNs are a special type
of NNs tailored to the learning of information encoded in a
graph. In GGNNs, each node corresponds to a hidden state
vector that is updated in a iterative way. Two features that
differentiate GGNNs from standard Recurrent Neural
Networks (RNNs), which also use the mechanism of recurrence,
is that in the former information can move bi-directionally
and that many nodes can be updated per step.
        </p>
        <p>Chuang et al. (2018) propose a GGNN-based method
which exploits contextual information, such as the types of
objects and their spatial relations, in order to detect
actionobject affordances. In the same vein, Ye et al. (2017) utilize
a two-stage pipeline built around a CNN to detect functional
areas in indoor scenes. More recently, Sawatzky et al. (2019)
utilized a GGNN, which takes into account the global
context of a scene and infers the affordances of the contained
objects. It also proposes the most suitable object for a
specific task. The authors prove that this approach yields better
results compared to methods that rely solely on the results
of an object classification algorithm.</p>
        <p>
          In
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref49 ref78 ref85 ref97">(Li et al. 2017)</xref>
          , an approach aiming to perform
situation recognition is presented, based on the detection of
human-object interactions. The goal is to predict the most
representative verb to describe what is taking place in a
scene, capturing also relevant semantic information (roles),
such as the actor, the source and target of the action, etc. The
authors utilize a GGNN, which enables combined reasoning
about verbs and their roles through the iterative propagation
of messages along the edges.
        </p>
        <p>Spatial Contextual Knowledge A particular type of
contextual knowledge concerns the spatial properties of the
entities populating a visual scene. These may involve simple
spatial relations, such as “object x is usually part of
object y” and “object x is usually situated near the objects
x1; : : : ; xn”, but also semantically enriched statements, such
as “objects of type x are usually found inside object z (e.g.,
in the fridge), located in room y”. Due to the ubiquitous
nature of spatial data in practical domains, such relations are
often captured as a separate class of context.</p>
        <p>Semantic spatial knowledge, when fused with low-level
metric information, gives great flexibility to a system. This is
demonstrated in the study of Gemignani et al. (2016), where
a novel representation is introduced that combines the
metric information of the environment with the symbolic data
that conveys meaning to the entities inhabiting it, as well
as with topological graphs. Although delivering a generic
model with clear semantics is not their main objective, the
resulting integrated representation enables a system to
perform high-level spatial reasoning, as well as to understand
target locations and positions of objects in the environment.</p>
        <p>Generality is the aim of the model proposed by Tenorth
and Beetz (2017), which uses an OWL ontology,
combining information from OpenCyc and other Web sources that
help compile new classes while forming the environmental
map. The main reasoning mechanisms of this study is
Prolog, although probabilistic reasoners are also used to tackle
fuzzy information or uncertain relations. A CV system
annotates objects based on their shape, their distances and the
dimension of the environment, using a monotonic
Description Logic (DL), to build the environmental map. As a result,
a coherent and well-formalized representation of
environments is achieved, which can offer high-quality datasets for
training data-driven models. The use of DL can also offer
a wide spectrum of spatial reasoning capabilities with
wellspecified properties and formal semantics.</p>
        <p>
          Adopting a different approach, the KG given in
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Chen
et al. 2018)</xref>
          enables spatial reasoning both in a local and
a global context which, in turn, results in improved
performance in semantic scene understanding. The proposed
framework consists of two distinct modules, one focusing
on local regions of the image and one dealing with the
whole image. The local module is convolution-based,
analyzing small regions in the image, whereas the basic
component of the global module is a KG representing regions and
classes as nodes, capturing spatial and semantic relations.
According to the experimental evaluation performed on the
ADE and Visual Genome datasets, the network achieves
better performance over other CNN-based baselines for region
classification, by a margin which sometimes is close to 10%.
According to the ablation study, the most decisive factor for
the framework’s performance was the KG.
        </p>
        <p>Modeling Affordances Building on the geometrical
structure and physical properties of objects, such as rigidity and
hollowness, the representation of affordances helps develop
systems that can reason about how human-level tasks are
performed. While ML is invaluable for automating the
process of learning from example when data is available, rich
representations can generalize and reuse the obtained
models in situations where data-based training is not possible.</p>
        <p>
          One of the first studies that demonstrated that even very
basic semantic models can improve the performance of
recognizing human-object interaction was
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Chao et al. 2015)</xref>
          .
The authors succeeded in boosting the performance of
visual classifiers by exploiting the compositionality and
concurrency of semantic concepts contained in images.
        </p>
        <p>
          KNOWROB 2.0
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Beetz et al. 2018)</xref>
          , which is the
result of a series of research activities in the field of
Cognitive Robotics, is an excellent example of integrating
topdown knowledge engineering with bottom-up information
structuring, involving, among others, a variety of CV tasks.
A combination of KBs of different granularity helps the
KNOWROB 2.0 framework capture rich models of the
world. The representation of high-level knowledge is based
on OWL-DL ontologies, a decidable fragment of First-order
Logic (FOL), yet adequately expressive for most practical
domains. The KBs enable the system to answer questions,
such as “how to pick up the cup”, “Which body part to use”
etc. The authors provide evidence that learning human
manipulation tasks on existing methods can be boosted by using
symbolic level structured knowledge.
        </p>
        <p>
          A recently proposed novel representation model that
manages to balance between concept abstraction, uncertainty
modeling and scalability is given in
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Daruna et al. 2019)</xref>
          .
The so called RoboCSE framework encodes the abstract,
semantic knowledge of an environment, i.e., the main
concepts and their relations, such as location, material and
affordance, obtained by observations, simulations, or even from
external sources, into multi-relational embeddings. These
embeddings are used to represent the knowledge graph of
the domain in vector space, encoding vertices that
represent entities as vectors and edges that represent relations as
mappings. While the majority of similar approaches rely on
Bayesian Logic Networks and Markov Logic Networks,
suffering from well-known intractability problems, the authors
prove that their model is highly scalable, robust to
uncertainty, and generalizes learned semantics.
        </p>
        <p>
          Learning from demonstration, or imitation learning, is
a relevant, yet broader objective, which introduces
interesting opportunities and challenges to a CV system (see
          <xref ref-type="bibr" rid="ref17 ref2 ref24 ref25 ref32 ref33 ref38 ref47 ref69 ref72 ref73 ref8 ref80 ref82 ref88">(Torabi, Warnell, and Stone 2019; Ravichandar et al. 2019)</xref>
          ).
Purely ML-based methods constitute the predominant
research direction, and only few state-of-the-art studies utilize
knowledge-based methods, taking advantage of the
reusability and generalization of the learned information. A popular
choice is to deploy expressive OWL-DL
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref3 ref34 ref35 ref36 ref37 ref40 ref43 ref46 ref49 ref58 ref65 ref66 ref68 ref70 ref78 ref85 ref86 ref97">(Ramirez-Amaro,
Beetz, and Cheng 2017; Lemaignan et al. 2017)</xref>
          or pure DL
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Agostini, Torras, and Woergoetter 2017)</xref>
          representations to
capture world knowledge. The CV modules are assigned the
task to extract information about the state of the
environment, the expert agent’s pose and location, grasping areas of
objects, affordances, shapes etc. On top of these, the
coupling with knowledge-based systems assists in visual
interpretation, for example to track human motion, to
semantically annotate the movement (i.e., “how the human performs
the action”) or to understand if a task is doable in a given
setting. These studies show that such representations enable a
system to reuse the learned knowledge in diverse settings
and under different conditions, without having to re-train
classifiers from scratch. Moreover, complex queries can be
answered, a topic discussed in the next subsection.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Reasoning over Expressive KBs</title>
        <p>Encoding knowledge in a semantically structured way is
only part of the story; a rich representation model can
also offer inference capabilities to a CV system, which are
needed for accomplishing complex tasks, such as scene
understanding, or simpler tasks under realistic conditions, such
as scene analysis with occlusions, noisy or erroneous input
etc. A reasoning system can be used to connect the dots
that relate concepts together when only partial observation
is available, especially in data-scarce situations, where
annotated data are not sufficiently many. In such situations,
the compositionality of information, an inherent
characteristic of the entities encountered in visual domains, can be
exploited by applying reasoning mechanisms.</p>
        <p>
          Complex Query Answering Probably the field that
highlights more clearly the needs and challenges faced by a CV
system in answering complex queries about a visual scene is
the field of Visual Question Answering (VQA). VQA was
recently introduced as a collection of benchmark
imagebased open-domain questions that, in order to be answered,
call for a deep understanding of the visual setting. VQA goes
beyond traditional CV, since apart from image analysis, the
proposed methods apply also a repertoire of AI techniques,
such as Natural Language Processing, in order to correctly
analyze the textual form of the question, and inferencing,
in order to interpret the purpose and intentions of the
entities acting in the scene
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Krishna et al. 2017)</xref>
          . The
challenges posed by this field are complex and multifaceted, a
fact which is also demonstrated by the rather poor
performance of state-of-the-art-systems in comparison to humans.
VQA is probably the area of CV that has drawn the most
inspiration from symbolic AI approaches to date.
        </p>
        <p>
          An indicative example is the approach recently presented
by Wu et al. (2018), who introduced a VQA model
combining observations obtained from the image with
information extracted from a general KB, namely DBpedia. Given
an image-question pair, a CNN is utilized to predict a set of
attributes from the image, i.e., the most recognizable objects
in the image, in terms of clarity and size. Consequently, a
series of captions based on the attributes is generated, which
is then used to extract relevant information from DBpedia
through appropriately formulated queries. In a similar style,
in
          <xref ref-type="bibr" rid="ref39 ref61 ref74 ref92 ref93">(Narasimhan and Schwing 2018)</xref>
          an external RDF
repository is used to retrieve properties of visual concepts, such
as category, used for, created by, etc. The technique utilizes
a Graph Convolution Network (GCN), a variation of GNN,
before producing an answer. In both cases, the ablation
analysis reveals the impact of the KB in improving performance.
        </p>
        <p>Other types of questions in VQA require inferencing
about the properties of the objects depicted in an image.
For example, queries such as “How is the man going to
work?” or more complex queries, such as “When did the
plane land?”, have been the subject of the study presented
by Krishna et al. (2017), who introduced the Visual Genome
dataset and a VQA method. In fact, this is one of the first
studies to bring a model trained on an RDF-based scene
graph that had good recall results to all What, Where, When,
Who, Why, How queries. Even further, Su et al. (2018)
introduced the visual knowledge memory network (VKMN) in
order to handle questions, whose answers cannot be directly
inferred from the image visual content but require reasoning
over structured human knowledge.</p>
        <p>
          The importance of capturing the semantic knowledge in
VQA collections led also to the creation of the
RelationVQA dataset
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Lu et al. 2018)</xref>
          , which extends Visual Genome
with a special module measuring the semantic similarity of
images. In contrast to methods mining only concepts or
attributes, this model extracts relation facts related to both
concepts and attributes. The experimental evaluation
conducted on VQA and COCO dataset showed that the method
outperformed other state-of-the-art ones. Moreover, the
ablation studies show that the incorporated semantic
knowledge was crucial for the performance of the network.
        </p>
        <p>
          Despite its increasing popularity, the VQA field is still
hard to confront. The generality of existing methods is also
questioned
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Goyal et al. 2019)</xref>
          . Developing generic
solutions, less tightly coupled to specific datasets, will definitely
benefit the pursuit towards broader visual intelligence.
Visual Reasoning A task related to VQA that has gained
popularity in recent years is that of Visual Reasoning (VR).
In this case, the type of questions that have to be answered
are more complex and require a multi-step reasoning
procedure. For example, given an image containing objects of
different shapes and color, the task of recognizing the color
of an object of certain shape that lies in a certain area w.r.t.
the position of another object of certain shape and color falls
to the category of VR (in this case, first the “source” object
must be detected, then the “target” object, and, finally, its
color must be recognized). Similar to the case of VQA, a
number of VR works has drawn inspiration from symbolic
AI-based ideas.
        </p>
        <p>In general, many VR works are based on Neural
Module Networks (NMNs) which are NNs of adaptable
architecture, the topology of which is determined by the parsing of
the question that has to be answered. NMNs simplify
complex questions into simpler sub-questions (sub-tasks), which
can be more easily addressed. The modules that constitute
the MNMs are pre-defined neural networks that implement
the functions that are required for the tackling of sub-tasks,
which are assembled into a layout dynamically. Central to
many MNMs is the utilization of prior symbolic (structured)
knowledge, which facilitates the handling of the sub-tasks.</p>
        <p>Hu et al. (2017) propose End-to-End Module Networks
as a variation of NMNs. The network first uses coarse
functional expressions describing the structure of the
computation required for the answering and, then, refines it
according to the textual input in order to assemble the network. For
example, for the question “how many other objects of the
same size as the purple cube exist?”, first crude functional
expression for counting and relocating would be predicted
as relevant to the answering of the question which,
subsequently, would be refined by the parameters from text
analysis (in this case one such parameter is the color of the cube).</p>
        <p>Similarly, Johnson et al. (2017) propose a variation of
NMNs, which is based on the concept of programs.
Programs are symbolic structures of certain specification
written in a Domain-Specific Language and are defined by a
syntax and semantics. In the context of VR, programs describe
a sequence of functions that must be executed, in order for
an answer to be computed. During testing on the CLEVR
dataset the model exhibited notable performance,
generalizing better in a variety of settings, such as for new question
types and human-posed questions. Building on the notion of
programs, Yi et al. (2018) further incorporated knowledge
regarding the structural scene representation of the image.
The method achieved near-perfect accuracy, while also
providing transparency to the reasoning process.</p>
        <p>
          An alternative NN-based approach for VR is found
in
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Santoro et al. 2017)</xref>
          , where the incorporation of
Relation Networks (RNs) in CNNs and Long Sort-Term
Memory (LSTM) architectures is proposed. RNs are architectures
whose computations focus explicitly on relational reasoning
and are characterized by three important features: they can
infer relations, they are data efficient, and they operate on a
set of objects, a flexible symbolic input format that is
agnostic to the kind of inputs it receives. For example, an object
could correspond to the background, to a particular physical
object, a texture, conjunctions of physical objects etc.
        </p>
        <p>
          To conclude, it is worth indicating also a recent trend in
visual explanation approaches that couples data-driven
systems with Answer Set Programming (ASP). ASP is a
nonmonotonic logical formalism oriented towards hard search
problems. A number of studies have emerged that
combine ASP abductive or inductive reasoning for the VQA
domain, especially for cases when training data are not many
(see e.g.,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref2 ref24 ref25 ref35 ref36 ref37 ref40 ref43 ref46 ref6 ref72 ref73 ref78 ref85 ref88 ref97">(Suchan et al. 2017; Riley and Sridharan 2019;
Basu, Shakerin, and Gupta 2020)</xref>
          )
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>The Web as a Problem-Agnostic Source of Data</title>
        <p>As the recent renaissance in AI is partly due to the
availability of big volumes of training data, along with the
computational power to analyze them, it is only reasonable to expect
that data-driven approaches will turn their attention to the
Web in order to collect the data needed. Although the
benefits mentioned in the previous sections are still achievable,
the challenges faced when using a Web repository rather
than a custom-made KB are now different.</p>
        <p>The vast majority of large-scale Web repositories are not
problem-specific, containing a lot of irrelevant information
for a ML system to be trained correctly. For the time being,
ML systems are highly specific, excelling only when trained
for a particular task and tested on similar to the training
conditions. As a result, state-of-the-art approaches try to rely on
the semantics of structured KBs, in order to filter out noisy
or irrelevant knowledge, by integrating external knowledge
when visual information is not sufficiently reliable for
conclusion making.</p>
        <p>
          Exploitation of Web-based Knowledge Graphs and
Semantic Repositories There exists a multitude of
studies that use external knowledge from structured or
semistructured Web resources, in order to answer visual queries
or to perform cognitive tasks. A characteristic example is
found in
          <xref ref-type="bibr" rid="ref3 ref34 ref46 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Li, Su, and Zhu 2017)</xref>
          , where the ConceptNet KG,
a semantic repository of commonsense Linked Open Data,
is used to answer open domain questions on entities such as
“What is the dog’s favorite food?”. The approach proceeds
in a step-wise manner: first, visual objects and keywords are
extracted from an image, using a Fast-RCNN for the
objects and a LSTM for the syntactical analysis; then, queries
to ConceptNet provide properties and values for the entities
found in the image. When an answer is considered correct,
a Dynamic Memory Network, which is an embedding
vector space that contains vector representations of symbolic
knowledge triples, is renewed for future encounter of the
same query. In a rather similar style, Wu et al. (2016a)
extract properties from DBpedia, by retrieving and
performing semantic analysis on the comment boxes of relevant
Wikipedia pages. Here, a CNN performs object detection on
the image, whereas a pre-trained RNN correlates attributes
to sentence descriptions.
        </p>
        <p>
          The approach presented in
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Shah et al. 2019)</xref>
          is the first
attempt to answer a more knowledge-intensive category of
questions, such as “Who is to the left of Barack Obama?”
or ‘‘Do all the people in the image have a common
occupation?”. These questions make reference to the named
entities contained in an image, e.g., Barack Obama, White
House, France etc. and require large KBs to retrieve the
relevant information. In this case, the authors choose Wikidata,
an RDF repository. They first extract named entities and then
try to connect them with a Wikidata entity using SPARQL
queries. In addition, they extract spatial relations with other
entities shown in the image and feed them to a Bi-LSTM.
A multi-layered perceptron calculates the prediction for an
answer, taking as input the output of the LSTM, along with
the SPARQL results.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Aligning Data Obtained from Diverse Online Sources</title>
        <p>Entity resolution, also known as instance matching,
concerns the task of identifying which entities across different
KBs refer to the same individual. As the Web is growing in
size, this problem is becoming crucial, especially in
application domains that need to integrate and align knowledge
obtained from various sources. An increasing number of CV
studies face this problem, in an attempt to interpret visual
information based on commonsense, non-visual knowledge.</p>
        <p>
          Two characteristic approaches are given in
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Chernova et
al. 2017)</xref>
          and
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Young et al. 2017)</xref>
          that try to assign labels
to a visual scene using Bayesian Logic Networks (BLNs)
and relying on commonsense knowledge. In
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Chernova et al.
2017)</xref>
          , knowledge is extracted from WordNet, ConceptNet,
and Wikipedia. WordNet is utilized in order to disambiguate
seed words returned by the CV annotator with the aid of their
hypernym. ConceptNet properties, such as IsLocatedIn or
U sedF or that may point the location of an object, are also
retrieved. With this method, the system can generate a
compact semantic KB given only a small number of objects.
        </p>
        <p>
          In
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Young et al. 2017)</xref>
          , a CNN trained on ImageNet is
used to annotate objects recognized in images. The system
is capable of assigning semantic categories to specific
regions, by relying on DBpedia comment boxes to calculate
the semantic relatedness between objects. As expected, high
accuracy of such an approach is difficult to achieve, due to
the diversity of information retrieved from DBpedia;
consequently, smarter ways of identifying only the relevant part of
the comment boxes need to be devised.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Exploitation of Commonsense Knowledge</title>
      <p>
        Much of the information presented in a visual scene is not
explicitly related with the features captured at the pixel level,
but concerns observations implicitly depicted in images.
Understanding the structure and dynamics of visual entities
requires being able to interpret the semantic and
commonsense (CS) features that are relevant, in addition to the
lowlevel information obtained by photorealistic rendering
techniques
        <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Vedantam et al. 2015)</xref>
        . This is a popular conclusion
reached within the CV community in the pursue towards
achieving visual intelligence. There is a long line of studies
that attempt to address the problem of extracting
commonsense knowledge from visual scenes or, similarly, of
utilizing commonsense inferences to improve scene
understanding. In this section, we discuss state-of-the-art approaches
that advance the field in these two directions.
      </p>
      <sec id="sec-3-1">
        <title>Mining Commonsense Knowledge from Images</title>
        <p>
          Even though ML is becoming part of many systems, it is still
not able to easily capture CS knowledge from the perceived
information. Additional techniques need to be devised to
extract this valuable type of knowledge from visual scenes. A
combination of textual and visual analysis, which extracts
subject-predicate-object triples (SPO) about objects
recognized in a scene, is addressed in certain studies, e.g.,
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref19 ref27 ref44 ref51 ref67 ref77 ref83 ref90">(Vedantam et al. 2015; Lin and Parikh 2015)</xref>
          . ML classifiers for
object recognition are trained on image datasets, while
pretrained NN classifiers help extract SPO triples, by
considering both the entities identified by the classifiers and the
textual description of the images.
        </p>
        <p>
          In a different direction, in
          <xref ref-type="bibr" rid="ref19 ref51 ref77">(Sadeghi, Kumar Divvala, and
Farhadi 2015)</xref>
          the authors rely on Web images to verify the
validity of simple phrases, such as “horses eat hay”,
analyzing the spatial consistency of the relative configurations
of the entities and the relations involved. This unsupervised
method is particularly interesting, due to the leverage it
offers in automatically enriching CS repositories. In fact, the
authors show how CV-based analysis can help improve
recall in KBs, such as WordNet, Cyc and ConceptNet, offering
a complementary and orthogonal source of evidence.
        </p>
        <p>Aditya et al. (2018) address the problem of generating
linguistic descriptions of images by utilizing a special type
of graph, namely scene description graphs (SDGs). Such
graphs are built by using both low-level information
derived using perception methods and high-level features
capturing CS knowledge stemming from the image annotations
and lexical ontological knowledge from Web resources.
SDGs produce object, scene and constituent detection
tuples, accompanied by a confidence score; pre-processed
background knowledge helps remove noise contained in the
detection. A Bayesian Network is utilized, in order for the
dependencies among co-occurring entities and knowledge
regarding abstract visual concepts to be captured.
Experimental evaluations of the method on the image-sentence
alignment quality, i.e., how close the generated description
is to the image being described, on Flickr8k, 30k and COCO
datasets, showed that the method achieves comparable
performance to previous state-of-the-art methods.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Commonsense Knowledge in Addressing OP Tasks</title>
        <p>State-of-the-art CS-based methodologies improve the
performance of a CV system, mainly by taking into account
textual descriptions about the entities found in a visual scene
or by retrieving semantic information from external sources
that is relevant to the image and the task at hand.</p>
        <p>
          A combination of external Web-based knowledge, text
processing and vision analysis is at the core of the study
presented in
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref92 ref98">(Wang et al. 2018)</xref>
          . The framework annotates
objects with a Fast-RNN, trained over the MS COCO dataset.
The extracted entities are enriched with (i) knowledge
retrieved from Wikipedia, in oder to perform entity
classification; (ii) knowledge from WebChild, attempting a
comparative analysis between relevant entities; and (iii) CS
knowledge obtained from ConceptNet, to create a semantically
rich description. The enriched entity is stored in an RDF
graph and is used to address a variety of tasks. For
instance, the framework has achieved improved accuracy in
VQA benchmarks, but also it can be used to generate
explanations for its answers. Prominent recent studies, as in
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Li et
al. 2019)</xref>
          and
          <xref ref-type="bibr" rid="ref39 ref61 ref74 ref92 ref93">(Narasimhan and Schwing 2018)</xref>
          , also build on
the direction of combining textual and visual analysis with
the help of knowledge obtained from CS repositories.
        </p>
        <p>Another problem that researchers try to address with the
help of CS knowledge is the sparsity of categorical
variables in the training datasets. For example, Ramanathan et
al. (2015) utilize a neural network framework that uses
different types of cues (linguistic, visual and logical) in the
context of human actions identification. Similarly, Lu et al.
(2016) exploit language priors extracted from the semantic
features of an image, in order to facilitate the
understanding of visual relationships. The proposed model combines a
visual module tailored to the learning of visual appearance
models for objects and predicates with a language module
capable of detecting semantically related relationships.</p>
        <p>More recently, Gu et al. (2019) utilize commonsense
knowledge stemming from an external KB in the context of
scene graph generation. Namely, a special knowledge-based
feature refinement module is used, which incorporates CS
knowledge from ConceptNet for the prediction of object
labels consisting of triplets containing the top-K
corresponding relationships, the object entity and a weight
corresponding to the frequency of the triplet. This strategy, aiming to
address the long tail distribution of relationships,
differentiates the approach from the linguistic-based ones described
previously, managing to showcase improvement in
generalizability and accuracy.</p>
        <p>
          CS knowledge is also used to tackle other CV problems,
such as in understanding relevant information about
unknown objects existing in a visual scene. In
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Icarte et al.
2017)</xref>
          or
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Young et al. 2016)</xref>
          for instance, external CS
Webbased repositories are used as a source for locating
relevant information. The general idea in both approaches is to
retrieve as much information as possible about the
recognizable objects that, based on diverse metrics, are
considered semantically close to the unknown ones. RelatedT o,
IsA, U sedF or properties found in ConceptNet, or
comment boxes retrieved from DBpedia are all relevant
knowledge that can be used for developing semantic similarity
measures. Similar to some extent, is the approach presented
in
          <xref ref-type="bibr" rid="ref75 ref96">(Ruiz-Sarmiento, Galindo, and Gonzalez-Jimenez 2016)</xref>
          ,
which relies on RDF graphs with a probabilistic
distribution over relations to capture the CS knowledge, but reverts
also to a human-supervised learning approach whenever
unknown objects are encountered.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Ability to Learn New Knowledge</title>
      <p>The majority of state-of-the-art studies covered in the
previous sections exploit a loosely-coupled combination of ML
and knowledge-based methodologies. A tighter integration
of methodologies of the two fields is expected to achieve
much broader impact, especially in the process of learning.
In the sequel, we consider prominent attempts towards this
direction, originating either from a model-free standpoint or
from a more declarative, inductive-based perspective.</p>
      <sec id="sec-4-1">
        <title>Model-Free Learning</title>
        <p>Recent studies devise methods that attempt to exploit
information contained in higher-level representations, in order
to improve scalability and generalization for tasks, such as
Zero-Shot Learning (ZSL). ZSL is the problem of
recognizing objects for which no visual examples have been obtained
and is typically achieved by exploring a semantic embedding
space, e.g., attribute or semantic word vector space.</p>
        <p>For example, Fu et al. (2015) utilize a semantic class label
graph, which results in a more accurate distance metric in the
semantic embedding space and an improved performance in
ZSL. Likewise, Xian et al. (2016) address the same
problem by proposing a novel latent embedding model, which
learns a compatibility function between the image and
semantic (class) embeddings. The model utilizes image and
class-level side-information that is either collected through
human annotation or through an unsupervised way from a
Web repository of text corpora.</p>
        <p>Lee et al. (2018) propose a novel deep learning
architecture for multi-label ZSL, which relies on KGs for the
discovery of the relationships between multiple classes of objects.
The KG is built on knowledge stemming from WordNet and
contains 3 types of label relations, super-subordinate,
positive correlation, and negative correlation. The KG is coupled
to a GGNN-type module for predicting labels.</p>
        <p>In the same vein, Wang, Ye and Gupta (2018) exploit the
information contained in KGs about unseen objects, in
order to infer visual attributes that enable their detection. The
KG nodes correspond to semantic categories and the edges
to semantic relationships, whereas the input to each node
is the vector representation (semantic embedding) of each
category. A GCN is used to transfer information between
different categories. This way, by utilizing the semantic
embeddings of a novel category, the method can link categories
in the KG to familiar ones and, thus, infer its attributes. The
experimental evaluation demonstrated a significant
improvement on the ImageNet dataset, while the ablation studies
indicated that the incorporation of KGs enabled the system to
learn meaningful classifiers on top of semantic embeddings.</p>
        <p>
          In
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Marino, Salakhutdinov, and Gupta 2017)</xref>
          , the use of
structured prior knowledge led to improved performance on
the task of multi-label image classification. The KG is built
using WordNet for the concepts and Visual Genome for the
relations among them. An interesting aspect of this study is
the introduction of a novel NN architecture, Graph Search
Neural Network, as a means to efficiently incorporate large
knowledge graphs, in order to be exploited for CV tasks.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Inductive Learning</title>
        <p>
          The benefits of developing intelligent visual components
with reasoning and learning abilities are becoming evident
in broader to CV domains, such as in the field of Robotics.
This conclusion was nicely demonstrated in a recent special
issue of the AI Journal
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Rajan and Saffiotti 2017)</xref>
          , where
causality-based reasoning emerged as a key contribution. It
is, therefore, interesting to investigate how the recent trend
in combining knowledge-based representations with
modelfree models for the development of intelligent robots is
making an impact in related OP research.
        </p>
        <p>
          A highly prominent line of research for modeling
uncertainty and high-level action knowledge is focusing on
combining expressive logical probabilistic formalisms,
ontological models and ML. In
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Antanas et al. 2018a)</xref>
          for example,
the system learns probabilistic first-order rules describing
relational affordances and pre-grasp configurations from
uncertain video data. It uses the ProbFOIL+ rule learner, along
with a simple ontology capturing object categories.
        </p>
        <p>More recently, Moldovan et al. (2018) significantly
extended this approach, using the Distributional Clauses (DCs)
formalism that integrates logic programming and
probability theory. DCs can use both continuous and discrete
variables, which is highly appropriate for modeling uncertainty,
in comparison for instance to ProbLog, which is commonly
found in relevant literature. Compared to approaches that
model affordances with Bayesian Networks, this approach
scales much better, but most importantly, due to its
relational nature, structural parts of the theory, such as the
abstract action-effect rules, can be transferred to similar
domains without the need to be learned again.</p>
        <p>A similar objective is pursued by Katzouris et al. (2019),
who propose an abductive-inductive incremental algorithm
for learning and revising causal rules, in the form of Event
Calculus programs. The Event Calculus is a highly
expressive, non-monotonic formalism for capturing causal and
temporal relations in dynamic domains. The approach uses
the XHAIL system as a basis, but sacrifices completeness
due to its incremental nature. Yet, it is able to learn weighted
causal temporal rules, in the form of Markov Logic
Networks, scaling up to large volumes of sequential data with
a time-like structure.</p>
        <p>Also worth mentioning is the study of Antanas et al.
(2018b), which instead of learning how to map visual
perceptions to task-dependent grasps, it uses a probabilistic
logic module to semantically reason about the most likely
object part to be grasped, given the object properties and
task constraints. The approach models rules in Causal
Probabilistic logic, implemented in ProbLog, in order to reason
about object categories, about the most affordable tasks and
about the best semantic pre-grasps.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Open Problems and Research Questions</title>
      <p>The review of the state-of-the-art reveals prominent
solutions for various OP-related topics, as well as novel
contributions that offer new insights (Table 1). The analysis can
also help frame open questions towards combining ML and
knowledge-based approaches in the given context.</p>
      <sec id="sec-5-1">
        <title>Obtaining Human Commonsense</title>
        <p>The exploitation of CS knowledge is a characteristic
example of a still open research area. Its significance was
acknowledged more than two decades ago and the research
conducted over the years contributed methods that combine
the strengths from diverse fields of AI. At the same time, it is
evident that there is still a long way to go; just the coupling
of textual and visual embeddings, the mainstream in current
VQA related studies, has proven to be a challenging task.
Further directions need to also be explored, such as in
performing complex forms of CS inferencing or in fusing the
huge volume of general knowledge that exists on the Web,
while eliminating the bias of information found online.</p>
        <p>Progress in the field of learning from demonstration can
prove a vital contribution to CS inferencing and vice versa.
Leaving the visual challenges involved aside, this
application domain, characterized by the central role of, human
mostly, agents, offers theory building opportunities on
diverse perspectives. Interaction with human users calls for
intuitive means of communication, where high-level,
declarative languages seem to offer a natural way of capturing
human intuition. Transferring knowledge between high-level
languages and low-level models is a key area of
investigation for future symbiotic systems and a fruitful domain for
combining data-driven and symbolic approaches.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Understanding Causality</title>
        <p>
          Still, the most demanding outcomes that are expected by the
integration of knowledge-based and ML methodologies
concern the aspects of causality learning and explainability.
Existing works on harvesting causality knowledge do not yet
offer convincing models. As argued in
          <xref ref-type="bibr" rid="ref63">(Pearl 2018)</xref>
          , ML
needs to go beyond the detection of associations, in order
to exhibit explainability and counterfactual reasoning.
        </p>
        <p>
          The black-box character of ML-based methods hinders
the understanding of their behavior, and eventually the
acceptance of such systems. For example, recent studies
demonstrate the fundamental inability of neural networks
to efficiently and robustly learn visual relations, which
renders the high performance that networks of this type often
achieve worth a closer investigation
          <xref ref-type="bibr" rid="ref39 ref39 ref61 ref61 ref74 ref74 ref92 ref92 ref93 ref93">(Kim, Ricci, and Serre
2018; Rosenfeld, Zemel, and Tsotsos 2018)</xref>
          . Advancement
in exploiting CS knowledge is expected to offer a significant
leverage in understanding and reasoning with causal
relations. And, of course, transparent reasoning is vital in
understanding the abilities and constrains of existing systems.
Yet, as indicated in the current review, this latter direction is
still not pursued in a coordinated and structured way.
        </p>
      </sec>
      <sec id="sec-5-3">
        <title>Achieving a Tighter Integration</title>
        <p>
          Ultimately, unifying logical and probabilistic graphical
models seems to be at the heart of handling the majority
of real-world problems. Recent studies show that even a
loosely-coupled integration can achieve better accuracy in
classification problems with small datasets in comparison
with end-to-end deep networks and comparable accuracy
with larger datasets (see e.g.,
          <xref ref-type="bibr" rid="ref2 ref24 ref25 ref6 ref72 ref73 ref88">(Riley and Sridharan 2019;
Basu, Shakerin, and Gupta 2020)</xref>
          ). A tighter integration is
highly anticipated, as it will help build systems that learn
from data, while still being able to generalize to domains
other than the ones trained for. Existing solutions are indeed
promising, as for example approaches based on the widely
used Markov Logic, which nevertheless introduces
limitations on both the theoretical and the practical level
          <xref ref-type="bibr" rid="ref2 ref24 ref25 ref72 ref73 ref88">(Domingos and Lowd 2019)</xref>
          . Its first-order nature, for instance, often
contradicts with the non-monotonicity met in CS domains.
        </p>
        <p>
          The support for complex tasks, such as causal, temporal or
counterfactual reasoning, in a non-monotonic fashion and
over rich conceptual representations unfolds a series of
research questions worth exploring in the near future.
Indicative Recent Literature
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Chuang et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Ye et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Sawatzky et al. 2019)</xref>
          ,
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Chao et al.
2015)</xref>
          ,
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Ramanathan et al. 2015)</xref>
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Beetz et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Ramirez-Amaro,
Beetz, and Cheng 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Lemaignan et
al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Agostini, Torras, and
Woergoetter 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Moldovan et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Icarte et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Redmon and
Farhadi 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Liu et al. 2018b)</xref>
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Gemignani et al. 2016)</xref>
          ,
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Tenorth and
Beetz 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Young et al. 2016)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Beetz et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Chernova et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Young et al.
2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Aditya et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Gu et al. 2019)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref49 ref78 ref85 ref97">(Li et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Chen
et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Krishna et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref104 ref105 ref106 ref11 ref27 ref44 ref67 ref83 ref90">(Zhu et al. 2015a)</xref>
          ,
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Li et al. 2019)</xref>
          ,
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Wu et al. 2016a)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Wu et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref3 ref34 ref46 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Li, Su, and Zhu
2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref19 ref51 ref77">(Sadeghi, Kumar Divvala, and
Farhadi 2015)</xref>
          ,
          <xref ref-type="bibr" rid="ref17 ref32 ref33 ref38 ref47 ref69 ref8 ref80 ref82">(Shah et al. 2019)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Su
et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref39 ref61 ref74 ref92 ref93">(Narasimhan and Schwing
2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref92 ref98">(Wang et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Vedantam et al. 2015)</xref>
          ,
          <xref ref-type="bibr" rid="ref19 ref51 ref77">(Lin and Parikh
2015)</xref>
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Lu et al. 2018)</xref>
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Hu et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Johnson et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Yi et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Santoro et al. 2017)</xref>
          <xref ref-type="bibr" rid="ref100 ref103 ref14 ref35 ref36 ref37 ref40 ref43 ref46 ref78 ref85 ref97">(Suchan et al. 2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref2 ref24 ref25 ref72 ref73 ref88">(Riley and
Sridharan 2019)</xref>
          ,
          <xref ref-type="bibr" rid="ref6">(Basu, Shakerin, and Gupta
2020)</xref>
          <xref ref-type="bibr" rid="ref3 ref34 ref49 ref58 ref65 ref66 ref68 ref70 ref86">(Marino, Salakhutdinov, and Gupta
2017)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Lee et al. 2018)</xref>
          ,
          <xref ref-type="bibr" rid="ref39 ref61 ref74 ref91 ref92 ref93">(Wang, Ye, and
Gupta 2018)</xref>
          <xref ref-type="bibr" rid="ref104 ref105 ref11 ref27 ref44 ref67 ref83 ref90">(Fu et al. 2015)</xref>
          ,
          <xref ref-type="bibr" rid="ref102 ref30 ref55 ref62 ref83 ref94 ref95 ref99">(Xian et al. 2016)</xref>
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Antanas et al. 2018a)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Antanas et al.
2018b)</xref>
          ,
          <xref ref-type="bibr" rid="ref1 ref101 ref12 ref16 ref4 ref42 ref5 ref52 ref53 ref56 ref60 ref7 ref84 ref91 ref98">(Moldovan et al. 2018)</xref>
          ,
(Katzouris et al. 2019)
        </p>
        <p>Problem
CV
Focus
affordance
detection
affordance
detection
object detection
object detection
KB Contribution: 1:concept abstraction/reuse, 2:complex data querying, 3:spatial reasoning, 4:contextual reasoning,
5:relational reasoning, 6:temporal reasoning, 7:causal reasoning, 8:access to open-domain knowledge, 9:formal semantics</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper, we reviewed approaches that rely on both
knowledge-based and data-driven methods, in order to
offer solutions to the field of intelligent object perception. By
adopting a knowledge-driven, rather than a problem-specific
grouping, we analyzed a multitude of approaches that
attempt to unify high-level knowledge with diverse machine
learning systems. The review revealed open and prominent
directions, showing clear evidence that hybrid methods
constitute an avenue worth exploring.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Aditya et al. 2018]
          <string-name>
            <surname>Aditya</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Baral</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Aloimonos</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and Fermu¨ller,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Image Understanding using vision and reasoning through Scene Description Graph</article-title>
          .
          <source>Computer Vision and Image Understanding</source>
          <volume>173</volume>
          :
          <fpage>33</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Aditya,
          <string-name>
            <surname>Yang</surname>
            , and Baral 2019] Aditya,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Baral</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Integrating knowledge and reasoning in image understanding</article-title>
          .
          <source>In Proceedings of the TwentyEighth International Joint Conference on Artificial Intelligence, IJCAI-19</source>
          ,
          <fpage>6252</fpage>
          -
          <lpage>6259</lpage>
          .
          <source>International Joint Conferences on Artificial Intelligence Organization.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Agostini, Torras, and Woergoetter 2017]
          <string-name>
            <surname>Agostini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Torras</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Woergoetter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Efficient interactive decision-making framework for robotic applications</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>247</volume>
          :
          <fpage>187</fpage>
          -
          <lpage>212</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Antanas et al. 2018a]
          <string-name>
            <surname>Antanas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dries</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>De Raedt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2018a</year>
          .
          <article-title>Relational affordance learning for task-dependent robot grasping</article-title>
          . In Lachiche, N., and
          <string-name>
            <surname>Vrain</surname>
          </string-name>
          , C., eds.,
          <source>Inductive Logic Programming</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . Cham: Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Antanas et al. 2018b]
          <string-name>
            <surname>Antanas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Neumann</surname>
            , M.; de Figueiredo,
            <given-names>R. P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kersting</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Santos-Victor</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>De Raedt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2018b</year>
          .
          <article-title>Semantic and geometric reasoning for robotic grasping: a probabilistic logic approach</article-title>
          .
          <source>Autonomous Robots 1-26.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Basu, Shakerin, and Gupta 2020] Basu,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Shakerin</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          ; and Gupta,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2020</year>
          .
          <article-title>Aqua: Asp-based visual question answering</article-title>
          . In Komendantskaya, E., and Liu, Y. A., eds.,
          <source>Practical Aspects of Declarative Languages</source>
          ,
          <fpage>57</fpage>
          -
          <lpage>72</lpage>
          . Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Beetz et al. 2018] Beetz,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Beßler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Haidu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Pomarlan</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ;
          <article-title>Bozcuog˘lu, A. K.;</article-title>
          and Bartels,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Know rob 2.0aˆa 2nd generation knowledge processing framework for cognition-enabled robotic agents</article-title>
          .
          <source>In 2018 IEEE ICRA</source>
          ,
          <volume>512</volume>
          -
          <fpage>519</fpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Bengio et al. 2019] Bengio,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Deleu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Rahaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Ke</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Lachapelle,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Bilaniuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ;
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ; and
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>A meta-transfer objective for learning to disentangle causal mechanisms</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .10912.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[Brachman and Levesque</source>
          <year>2004</year>
          ] Brachman,
          <string-name>
            <given-names>R.</given-names>
            , and
            <surname>Levesque</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2004</year>
          .
          <article-title>Knowledge Representation and Reasoning</article-title>
          . San Francisco, CA, USA: Morgan Kaufmann Publishers Inc.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Chao et al. 2015] Chao,
          <string-name>
            <given-names>Y. W.</given-names>
            ;
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ;
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ; and
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>HICO: A benchmark for recognizing human-object interactions in images</article-title>
          .
          <source>IEEE ICCV 2015 Inter:</source>
          <fpage>1017</fpage>
          -
          <lpage>1025</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Chen et al. 2018]
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L. J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Iterative Visual Reasoning beyond Convolutions</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>IEEE CVPR 7239-7248.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Chernova et al. 2017]
          <string-name>
            <surname>Chernova</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chu</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Daruna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Garrison</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ; Hahn,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Khante</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>W.</given-names>
            ; and
            <surname>Thomaz</surname>
          </string-name>
          , A.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          2017.
          <article-title>Situated bayesian reasoning framework for robots operating in diverse everyday environments</article-title>
          .
          <source>In International Symposium on Robotics Research (ISRR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Chuang et al. 2018]
          <string-name>
            <surname>Chuang</surname>
            ,
            <given-names>C. Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Fidler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Learning to Act Properly: Predicting and Explaining Affordances from Images</article-title>
          .
          <source>IEEE CVPR 975- 983.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Daruna et al. 2019]
          <string-name>
            <surname>Daruna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>W.</given-names>
            ;
            <surname>Kira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ; and
            <surname>Chernova</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Robocse: Robot common sense embedding</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          arXiv preprint arXiv:
          <year>1903</year>
          .00412.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>[Davis and Marcus</source>
          <year>2015</year>
          ]
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Marcus</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <article-title>Commonsense reasoning and commonsense knowledge in artificial intelligence</article-title>
          .
          <source>Commun. ACM</source>
          <volume>58</volume>
          (
          <issue>9</issue>
          ):
          <fpage>92</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>[Dean</surname>
            ,
            <given-names>Allen,</given-names>
          </string-name>
          <source>and Aloimonos</source>
          <year>1995</year>
          ]
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Allen</surname>
            , J.; and Aloimonos,
            <given-names>Y.</given-names>
          </string-name>
          <year>1995</year>
          .
          <source>Artificial Intelligence: Theory and Practice</source>
          . Redwood City, CA, USA:
          <string-name>
            <surname>Benjamin-Cummings Publishing</surname>
            <given-names>Co.</given-names>
          </string-name>
          , Inc.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [Deng et al. 2014]
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Ding,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Frome</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Neven</surname>
          </string-name>
          , H.; and Adam, H.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          2014.
          <article-title>Large-scale object classification using label relation graphs</article-title>
          .
          <source>In ECCV</source>
          ,
          <fpage>48</fpage>
          -
          <lpage>64</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>[Domingos and Lowd</source>
          <year>2019</year>
          ] Domingos,
          <string-name>
            <given-names>P.</given-names>
            , and
            <surname>Lowd</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          2019.
          <article-title>Unifying logical and statistical ai with markov logic.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>Commununications of the ACM</source>
          <volume>62</volume>
          (
          <issue>7</issue>
          ):
          <fpage>74</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Fu et al. 2015]
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ; Xiang,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Kodirov</surname>
          </string-name>
          , E.; and
          <string-name>
            <surname>Gong</surname>
          </string-name>
          , S.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          2015.
          <article-title>Zero-shot object recognition by semantic manifold distance</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>2635</volume>
          -
          <fpage>2644</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [Geffner 2018] Geffner,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Model-free, model-based, and general intelligence</article-title>
          .
          <source>In Proceedings of the 27th International Joint Conference on Artificial Intelligence, IJCAI'18</source>
          ,
          <fpage>10</fpage>
          -
          <lpage>17</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [Gemignani et al. 2016] Gemignani,
          <string-name>
            <surname>G.</surname>
          </string-name>
          ; Capobianco,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Bastianelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Bloisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. D.</given-names>
            ;
            <surname>Iocchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ; and
            <surname>Nardi</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <article-title>Living with robots: Interactive environmental knowledge acquisition</article-title>
          .
          <source>Robotics and Autonomous Systems</source>
          <volume>78</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Goyal et al. 2019] Goyal,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Khot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Summers-Stay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ; and
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering</article-title>
          .
          <source>IJCV</source>
          <volume>127</volume>
          (
          <issue>4</issue>
          ):
          <fpage>398</fpage>
          -
          <lpage>414</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [Gu et al. 2019]
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cai</surname>
            , J.; and Ling,
            <given-names>M.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Scene graph generation with external knowledge and image reconstruction</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <year>1969</year>
          -
          <fpage>1978</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [Herath, Harandi, and Porikli 2017] Herath,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Harandi,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Porikli</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Going deeper into action recognition: A survey</article-title>
          .
          <source>IMAVIS</source>
          <volume>60</volume>
          :
          <fpage>4</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [Hu et al. 2017] Hu,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Andreas,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ; Rohrbach,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Darrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ; and
            <surname>Saenko</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Learning to Reason: End-to-End Module Networks for Visual Question Answering</article-title>
          .
          <source>IEEE ICCV 2017-Octob(Figure</source>
          <volume>1</volume>
          ):
          <fpage>804</fpage>
          -
          <lpage>813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [Icarte et al. 2017] Icarte, R. T.;
          <string-name>
            <surname>Baier</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ruz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Soto</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>How a general-purpose commonsense ontology can improve performance of learning-based image retrieval</article-title>
          .
          <source>arXiv preprint arXiv:1705</source>
          .
          <fpage>08844</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [Johnson et al. 2017] Johnson, J.;
          <string-name>
            <surname>Hariharan</surname>
          </string-name>
          , B.;
          <string-name>
            <surname>van der Maaten</surname>
          </string-name>
          , L.;
          <string-name>
            <surname>Hoffman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; Lawrence Zitnick,
          <string-name>
            <surname>C.</surname>
          </string-name>
          ; and Girshick,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Inferring and executing programs for visual reasoning</article-title>
          .
          <source>In IEEE ICCV</source>
          ,
          <volume>2989</volume>
          -
          <fpage>2998</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [Katzouris et al. 2019]
          <string-name>
            <surname>Katzouris</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Michelioudakis</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Artikis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and Paliouras,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Online learning of weighted relational rules for complex event recognition</article-title>
          . In Berlingerio, M.;
          <string-name>
            <surname>Bonchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ; Ga¨rtner, T.;
          <string-name>
            <surname>Hurley</surname>
          </string-name>
          , N.; and Ifrim, G., eds.,
          <source>Machine Learning and Knowledge Discovery in Databases</source>
          ,
          <volume>396</volume>
          -
          <fpage>413</fpage>
          . Cham: Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [Kim, Ricci, and Serre 2018]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Ricci,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and Serre,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Not-so-clevr: learning same-different relations strains feedforward neural networks</article-title>
          .
          <source>Interface focus 8</source>
          (
          <issue>4</issue>
          ):
          <fpage>20180011</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [Krishna et al. 2017] Krishna,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Zhu,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Groth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ;
            <surname>Johnson</surname>
          </string-name>
          , J.;
          <string-name>
            <surname>Hata</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kravitz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Kalantidis,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J.</given-names>
            ;
            <surname>Shamma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            ;
            <surname>Bernstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            ; and
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>Visual</given-names>
            <surname>Genome</surname>
          </string-name>
          <article-title>: Connecting Language and Vision Using Crowdsourced Dense Image Annotations</article-title>
          .
          <source>IJCV</source>
          <volume>123</volume>
          (
          <issue>1</issue>
          ):
          <fpage>32</fpage>
          -
          <lpage>73</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <surname>[Lee</surname>
          </string-name>
          et al.
          <year>2018</year>
          ]
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C. W.</given-names>
          </string-name>
          ; Fang,
          <string-name>
            <surname>W.</surname>
          </string-name>
          ; Yeh,
          <string-name>
            <given-names>C. K.</given-names>
            ; and
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. C. F.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Multi-label Zero-Shot Learning with Structured Knowledge Graphs</article-title>
          .
          <source>IEEE CVPR 1576-1585.</source>
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [Lemaignan et al. 2017]
          <string-name>
            <surname>Lemaignan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Warnier,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Sisbot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            ;
            <surname>Clodic</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          ; and Alami,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Artificial cognition for social human-robot interaction: An implementation</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>247</volume>
          :
          <fpage>45</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>[Li</surname>
          </string-name>
          et al.
          <year>2015</year>
          ]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tarlow</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Brockschmidt,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and Zemel,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Gated graph sequence neural networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <source>arXiv preprint arXiv:1511</source>
          .
          <fpage>05493</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <string-name>
            <surname>[Li</surname>
          </string-name>
          et al.
          <year>2017</year>
          ]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Tapaswi,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Jia,
          <string-name>
            <given-names>J.</given-names>
            ; Urtasun, R.; and
            <surname>Fidler</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Situation recognition with graph neural networks</article-title>
          .
          <source>In IEEE ICCV</source>
          ,
          <volume>4173</volume>
          -
          <fpage>4182</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <string-name>
            <surname>[Li</surname>
          </string-name>
          et al.
          <year>2019</year>
          ]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Hengel</surname>
          </string-name>
          , A.
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          v. d.
          <year>2019</year>
          .
          <article-title>Visual question answering as reading comprehension</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>6319</volume>
          -
          <fpage>6328</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [Li,
          <string-name>
            <surname>Su,</surname>
          </string-name>
          <source>and Zhu</source>
          <year>2017</year>
          ]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ; Su, H.; and
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          <article-title>Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks</article-title>
          .
          <source>arXiv preprint arXiv:1712</source>
          .
          <fpage>00733</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <source>[Lin and Parikh</source>
          <year>2015</year>
          ]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Don't just listen, use your imagination: Leveraging visual common sense for non-visual tasks</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>2984</volume>
          -
          <fpage>2993</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [Liu et al. 2018a]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ouyang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fieguth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <surname>X.</surname>
          </string-name>
          ; and Pietika¨inen,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2018a</year>
          .
          <article-title>Deep learning for generic object detection: A survey</article-title>
          . arXiv preprint arXiv:
          <year>1809</year>
          .02165.
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [Liu et al. 2018b]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Shan,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          <year>2018b</year>
          .
          <article-title>Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level Relationships</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          <source>IEEE CVPR 6985-6994.</source>
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [Lu et al. 2016]
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Krishna</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Bernstein,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and FeiFei,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Visual relationship detection with language priors</article-title>
          .
          <source>Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) 9905 LNCS(Figure</source>
          <volume>2</volume>
          ):
          <fpage>852</fpage>
          -
          <lpage>869</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          [Lu et al. 2018] Lu,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          ; Zhang, W.; Duan,
          <string-name>
            <given-names>N.</given-names>
            ;
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2018</year>
          . R-VQA:
          <article-title>Learning visual relation facts with semantic attention for visual question answering</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          <source>Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          <year>1880</year>
          -
          <year>1889</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          [Marino, Salakhutdinov, and Gupta 2017] Marino,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Salakhutdinov</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>The more you know: using knowledge graphs for image classification</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          <string-name>
            <surname>IEEE CVPR 2017-Janua</surname>
          </string-name>
          :
          <fpage>20</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref60">
        <mixed-citation>
          [Moldovan et al. 2018]
          <string-name>
            <surname>Moldovan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nitti</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Santos-Victor</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>De Raedt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Relational affordances for multiple-object manipulation</article-title>
          .
          <source>Autonomous Robots</source>
          <volume>42</volume>
          (
          <issue>1</issue>
          ):
          <fpage>19</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref61">
        <mixed-citation>
          <source>[Narasimhan and Schwing</source>
          <year>2018</year>
          ] Narasimhan,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Schwing</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Straight to the facts: Learning knowledge base retrieval for factual visual question answering</article-title>
          .
          <source>In Proceedings of the ECCV (ECCV)</source>
          ,
          <fpage>451</fpage>
          -
          <lpage>468</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref62">
        <mixed-citation>
          [Nickel et al. 2016] Nickel,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Tresp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ; and
            <surname>Gabrilovich</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>A review of relational machine learning for knowledge graphs</article-title>
          .
          <source>Proceedings of the IEEE</source>
          <volume>104</volume>
          (
          <issue>1</issue>
          ):
          <fpage>11</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref63">
        <mixed-citation>
          <source>[Pearl</source>
          <year>2018</year>
          ]
          <string-name>
            <surname>Pearl</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Theoretical impediments to machine learning with seven sparks from the causal revolution</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref64">
        <mixed-citation>
          arXiv preprint arXiv:
          <year>1801</year>
          .04016.
        </mixed-citation>
      </ref>
      <ref id="ref65">
        <mixed-citation>
          <source>[Rajan and Saffiotti</source>
          <year>2017</year>
          ]
          <string-name>
            <surname>Rajan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Saffiotti</surname>
          </string-name>
          , A., eds.
        </mixed-citation>
      </ref>
      <ref id="ref66">
        <mixed-citation>
          2017.
          <article-title>Special Issue on AI and Robotics</article-title>
          , volume
          <volume>247</volume>
          . Elsevier.
          <volume>1</volume>
          -
          <fpage>440</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref67">
        <mixed-citation>
          [Ramanathan et al. 2015]
          <string-name>
            <surname>Ramanathan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Deng</surname>
            , J.; and Han,
            <given-names>W.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Learning semantic relationships for better action retrieval in images ( Supplementary )</article-title>
          .
          <source>Computer Vision and Pattern Recognition 1-4.</source>
        </mixed-citation>
      </ref>
      <ref id="ref68">
        <mixed-citation>
          [
          <string-name>
            <surname>Ramirez-Amaro</surname>
          </string-name>
          , Beetz, and Cheng 2017]
          <article-title>Ramirez-</article-title>
          <string-name>
            <surname>Amaro</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Beetz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and Cheng, G.
          <year>2017</year>
          .
          <article-title>Transferring skills to humanoid robots by extracting semantic representations from observations of human activities</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>247</volume>
          :
          <fpage>95</fpage>
          -
          <lpage>118</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref69">
        <mixed-citation>
          [Ravichandar et al. 2019] Ravichandar,
          <string-name>
            <given-names>H.</given-names>
            ;
            <surname>Polydoros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            ;
            <surname>Chernova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Billard</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Robot learning from demonstration: A review of recent advances</article-title>
          .
          <source>Annual Review of Control</source>
          , Robotics, and Autonomous Systems In Press.
        </mixed-citation>
      </ref>
      <ref id="ref70">
        <mixed-citation>
          <source>[Redmon and Farhadi</source>
          <year>2017</year>
          ] Redmon,
          <string-name>
            <given-names>J.</given-names>
            , and
            <surname>Farhadi</surname>
          </string-name>
          , A.
        </mixed-citation>
      </ref>
      <ref id="ref71">
        <mixed-citation>
          2017.
          <article-title>Yolo9000: better, faster, stronger</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>7263</volume>
          -
          <fpage>7271</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref72">
        <mixed-citation>
          <source>[Riley and Sridharan</source>
          <year>2019</year>
          ] Riley,
          <string-name>
            <given-names>H.</given-names>
            , and
            <surname>Sridharan</surname>
          </string-name>
          , M.
        </mixed-citation>
      </ref>
      <ref id="ref73">
        <mixed-citation>
          2019.
          <article-title>Integrating non-monotonic logical reasoning and inductive learning with deep learning for explainable visual question answering</article-title>
          .
          <source>Frontiers in Robotics and AI</source>
          <volume>6</volume>
          :
          <fpage>125</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref74">
        <mixed-citation>
          [Rosenfeld, Zemel, and Tsotsos 2018]
          <string-name>
            <surname>Rosenfeld</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zemel</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Tsotsos</surname>
            ,
            <given-names>J. K.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>The elephant in the room</article-title>
          . arXiv preprint arXiv:
          <year>1808</year>
          .03305.
        </mixed-citation>
      </ref>
      <ref id="ref75">
        <mixed-citation>
          [
          <string-name>
            <surname>Ruiz-Sarmiento</surname>
          </string-name>
          , Galindo, and Gonzalez-Jimenez 2016]
          <article-title>Ruiz-</article-title>
          <string-name>
            <surname>Sarmiento</surname>
            , J.-R.; Galindo,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Gonzalez-Jimenez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Probability and common-sense: Tandem towards robust robotic object recognition in ambient assisted living</article-title>
          .
          <source>In Ubiquitous Computing and Ambient Intelligence.</source>
        </mixed-citation>
      </ref>
      <ref id="ref76">
        <mixed-citation>
          Springer. 3-
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref77">
        <mixed-citation>
          [Sadeghi,
          <string-name>
            <given-names>Kumar</given-names>
            <surname>Divvala</surname>
          </string-name>
          , and Farhadi 2015]
          <string-name>
            <surname>Sadeghi</surname>
            ,
            <given-names>F.; Kumar</given-names>
          </string-name>
          <string-name>
            <surname>Divvala</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Farhadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Viske: Visual knowledge extraction and question answering by visual verification of relation phrases</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>1456</volume>
          -
          <fpage>1464</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref78">
        <mixed-citation>
          [Santoro et al. 2017]
          <string-name>
            <surname>Santoro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Raposo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Barrett</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref79">
        <mixed-citation>
          <string-name>
            <given-names>G. T.</given-names>
            ;
            <surname>Malinowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Pascanu</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Battaglia,
          <string-name>
            <surname>P.</surname>
          </string-name>
          ; and Lillicrap,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>A simple neural network module for relational reasoning</article-title>
          .
          <source>(Nips).</source>
        </mixed-citation>
      </ref>
      <ref id="ref80">
        <mixed-citation>
          [Sawatzky et al. 2019]
          <string-name>
            <surname>Sawatzky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Souri,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Grund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Gall</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2019</year>
          .
          <string-name>
            <given-names>What</given-names>
            <surname>Object Should I Use</surname>
          </string-name>
          ?
          <article-title>- Task Driven Object Detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref81">
        <mixed-citation>
          [Scarselli et al. 2008]
          <string-name>
            <surname>Scarselli</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Gori</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tsoi</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hagenbuchner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and Monfardini,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>The graph neural network model</article-title>
          .
          <source>IEEE Transactions on NN 20</source>
          (
          <issue>1</issue>
          ):
          <fpage>61</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref82">
        <mixed-citation>
          [Shah et al. 2019]
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mishra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yadati</surname>
            , N.; and Talukdar,
            <given-names>P. P.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Kvqa: Knowledge-aware visual question answering</article-title>
          .
          <source>AAAI.</source>
        </mixed-citation>
      </ref>
      <ref id="ref83">
        <mixed-citation>
          [Stone et al. 2016] Stone,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Brooks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Brynjolfsson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Calo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ;
            <surname>Hager</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ;
            <surname>Hirschberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ; Kalyanakrishnan,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Kamar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ;
            <surname>Kraus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ;
            <surname>Leyton-Brown</surname>
          </string-name>
          , K.;
          <string-name>
            <surname>Parkes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; Press, W.;
          <string-name>
            <surname>Saxenian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Tambe,
          <string-name>
            <given-names>M.</given-names>
            ; ; and
            <surname>Teller</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Artificial intelligence and life in 2030</article-title>
          .
          <source>One Hundred Year Study on Artificial Intelligence: Report of the 2015-2016 Study Panel.</source>
        </mixed-citation>
      </ref>
      <ref id="ref84">
        <mixed-citation>
          [Su et al. 2018]
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Learning Visual Knowledge Memory Networks for Visual Question Answering</article-title>
          .
          <source>IEEE CVPR 7736- 7745.</source>
        </mixed-citation>
      </ref>
      <ref id="ref85">
        <mixed-citation>
          [Suchan et al. 2017]
          <string-name>
            <surname>Suchan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Bhatt,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Walega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            ; and
            <surname>Schultz</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. P. L.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Visual explanation by highlevel abduction: On answer-set programming driven reasoning about moving objects</article-title>
          .
          <source>CoRR abs/1712</source>
          .00840.
        </mixed-citation>
      </ref>
      <ref id="ref86">
        <mixed-citation>
          <source>[Tenorth and Beetz</source>
          <year>2017</year>
          ] Tenorth,
          <string-name>
            <given-names>M.</given-names>
            , and
            <surname>Beetz</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref87">
        <mixed-citation>
          <article-title>Representations for robot knowledge in the knowrob framework</article-title>
          .
          <source>Artificial Intelligence</source>
          <volume>247</volume>
          :
          <fpage>151</fpage>
          -
          <lpage>169</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref88">
        <mixed-citation>
          [Torabi, Warnell, and Stone 2019]
          <string-name>
            <surname>Torabi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Warnell</surname>
            , G.; and Stone,
            <given-names>P.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>Recent advances in imitation learning from observation</article-title>
          .
          <source>In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19</source>
          ,
          <fpage>6325</fpage>
          -
          <lpage>6331</lpage>
          .
          <source>International Joint Conferences on Artificial Intelligence Organization.</source>
        </mixed-citation>
      </ref>
      <ref id="ref89">
        <mixed-citation>
          <string-name>
            <surname>[van Harmelen</surname>
          </string-name>
          et al. 2007
          <string-name>
            <surname>] van Harmelen</surname>
          </string-name>
          , F.;
          <string-name>
            <surname>van Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lifschitz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Porter</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2007</year>
          .
          <article-title>Handbook of Knowledge Representation</article-title>
          . San Diego, USA: Elsevier Science.
        </mixed-citation>
      </ref>
      <ref id="ref90">
        <mixed-citation>
          [Vedantam et al. 2015] Vedantam,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ;
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          ; Lawrence Zitnick,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Learning common sense through visual abstraction</article-title>
          .
          <source>In IEEE ICCV</source>
          ,
          <volume>2542</volume>
          -
          <fpage>2550</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref91">
        <mixed-citation>
          <string-name>
            <surname>[Wang</surname>
          </string-name>
          et al. 2018]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dick</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and van den Hengel,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Fvqa: Fact-based visual question answering</article-title>
          .
          <source>IEEE Trans. on PAMI</source>
          <volume>40</volume>
          (
          <issue>10</issue>
          ):
          <fpage>2413</fpage>
          -
          <lpage>2427</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref92">
        <mixed-citation>
          [Wang, Ye, and Gupta 2018]
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Gupta</surname>
          </string-name>
          , A.
        </mixed-citation>
      </ref>
      <ref id="ref93">
        <mixed-citation>
          2018.
          <article-title>Zero-Shot Recognition via Semantic Embeddings and Knowledge Graphs</article-title>
          .
          <source>IEEE CVPR 6857-6866.</source>
        </mixed-citation>
      </ref>
      <ref id="ref94">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al. 2016a]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dick</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and van den Hengel,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2016a</year>
          .
          <article-title>Ask me anything: Free-form visual question answering based on knowledge from external sources</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>4622</volume>
          -
          <fpage>4630</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref95">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al. 2016b]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Jiang</surname>
          </string-name>
          , Y.-G.; and
          <string-name>
            <surname>Sigal</surname>
          </string-name>
          , L.
        </mixed-citation>
      </ref>
      <ref id="ref96">
        <mixed-citation>
          2016b.
          <article-title>Harnessing object and scene semantics for largescale video understanding</article-title>
          .
          <source>In TheIEEE CVPR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref97">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al.
          <year>2017</year>
          ]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Teney</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dick</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and van den Hengel,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Visual question answering: A survey of methods and datasets</article-title>
          .
          <source>CVIU</source>
          <volume>163</volume>
          :
          <fpage>21</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref98">
        <mixed-citation>
          <string-name>
            <surname>[Wu</surname>
          </string-name>
          et al.
          <year>2018</year>
          ]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Dick</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Van Den Hengel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>Image Captioning and Visual Question Answering Based on Attributes and External Knowledge</article-title>
          .
          <source>IEEE Trans. on PAMI 40</source>
          (
          <issue>6</issue>
          ):
          <fpage>1367</fpage>
          -
          <lpage>1381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref99">
        <mixed-citation>
          [Xian et al. 2016] Xian,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Akata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ;
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ;
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            ;
            <surname>Hein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ; and
            <surname>Schiele</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Latent embeddings for zero-shot classification</article-title>
          .
          <source>In IEEE CVPR</source>
          ,
          <volume>69</volume>
          -
          <fpage>77</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref100">
        <mixed-citation>
          [Ye et al. 2017]
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Fermuller,
          <string-name>
            <surname>C.</surname>
          </string-name>
          ; and Aloimonos,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>What can i do around here? Deep functional scene understanding for cognitive robots</article-title>
          .
          <source>IEEE ICRA 4604-4611.</source>
        </mixed-citation>
      </ref>
      <ref id="ref101">
        <mixed-citation>
          [Yi et al. 2018]
          <string-name>
            <surname>Yi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Torralba</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Kohli,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Tenenbaum</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. B.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Neural-symbolic VQA: Disentangling reasoning from vision and language understanding</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          2018-December(NeurIPS):
          <fpage>1031</fpage>
          -
          <lpage>1042</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref102">
        <mixed-citation>
          [Young et al. 2016]
          <string-name>
            <surname>Young</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kunze</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cabrio</surname>
          </string-name>
          , E.; and
          <string-name>
            <surname>Hawes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Towards lifelong object learning by integrating situated robot perception and semantic web mining</article-title>
          .
          <source>In Proceedings of the Twenty-second European Conference on Artificial Intelligence</source>
          ,
          <fpage>1458</fpage>
          -
          <lpage>1466</lpage>
          . IOS Press.
        </mixed-citation>
      </ref>
      <ref id="ref103">
        <mixed-citation>
          [Young et al. 2017]
          <string-name>
            <surname>Young</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Suchi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kunze</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hawes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vincze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Making sense of indoor spaces using semantic web mining and situated robot perception</article-title>
          .
          <source>In European Semantic Web Conference</source>
          ,
          <volume>299</volume>
          -
          <fpage>313</fpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref104">
        <mixed-citation>
          [Zhu et al. 2015a]
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Groth</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and FeiFei,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2015a</year>
          .
          <article-title>Visual7W: Grounded Question Answering in Images.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref105">
        <mixed-citation>
          [Zhu et al. 2015b]
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; Zhang,
          <string-name>
            <surname>C.</surname>
          </string-name>
          ; Re´,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref106">
        <mixed-citation>
          2015b.
          <article-title>Building a large-scale multimodal knowledge base for visual question answering</article-title>
          .
          <source>CoRR abs/1507</source>
          .05670.
        </mixed-citation>
      </ref>
      <ref id="ref107">
        <mixed-citation>
          [Zhu, Fathi, and
          <string-name>
            <surname>Fei-Fei 2014] Zhu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fathi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and FeiFei,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Reasoning about object affordances in a knowledge base representation</article-title>
          . In Fleet, D.; Pajdla,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Schiele</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          ; and Tuytelaars, T., eds.,
          <source>ECCV</source>
          ,
          <fpage>408</fpage>
          -
          <lpage>424</lpage>
          . Cham: Springer International Publishing.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>