<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Corresponding author.
m.sridharan@bham.ac.uk (M. Sridharan)
{ https://www.cs.bham.ac.uk/~sridharm/ (M. Sridharan)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Cognitive Adequacy: Insights from Developing Robot Architectures⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohan Sridharan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Intelligent Robotics Lab, School of Computer Science, University of Birmingham</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper discusses cognitive adequacy for robots collaborating with and assisting humans. We share insights from the development of robot architectures that use knowledge-driven and data-driven methods to jointly address challenges in transparent knowledge representation, reasoning, and learning in robotics. Consider a robot delivering objects to particular places or stacking objects in desired configurations. Such robots have to reason with different descriptions of incomplete domain knowledge and uncertainty. These descriptions include commonsense knowledge, e.g., relations between some domain objects and default statements such as “textbooks are usually in the library” that hold true in all but a few exceptional circumstances. At the same time, information extracted from noisy sensor inputs is often associated with quantitative measures of uncertainty, e.g., “I am 90% certain I saw the robotics book in the office”. Also, any robot in a practical domain will have to revise its existing knowledge over time, often using data-driven methods. In addition, for effective collaboration with humans, the robot may need to describe (or justify) its decisions and beliefs. In state of the art architectures that combine knowledge-based reasoning (e.g., for planning) and data-driven learning (e.g., for object recognition) for such integrated robot systems, cognitive adequacy, i.e., the ability to support the desired behavior, thus poses open problems in knowledge representation, reasoning, and learning. This paper builds on expertise in designing robot architectures [1, 2, 3] to identify some key underlying principles for cognitive adequacy.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Non-monotonic logical reasoning</kwd>
        <kwd>Probabilistic reasoning</kwd>
        <kwd>Interactive learning</kwd>
        <kwd>Robotics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Motivation</title>
    </sec>
    <sec id="sec-2">
      <title>2. Architecture and Insights</title>
      <p>Figure 1(left) is an overview of the architecture that encodes the principle of stepwise iterative
refinement . It is based on tightly-coupled transition diagrams at different resolutions, and may be
viewed as a logician, statistician, and an explorer working together. These diagrams are described
using an action language, which has a sorted signature with statics, (Boolean, non-Boolean)
lfuents and actions, and supports (deterministic, non-deterministic) causal laws, state constraints,</p>
      <p>Coarser−resolution</p>
      <p>Representation
(Resolution 1)
(Logician) Representation
(Explorer) ianbtsetnrtaicotnal (Resolution i)</p>
      <p>Interactive transition
Learning
observed
outcomes</p>
      <p>Representation
(Statistician) (Resolution i+1)
Finer−resolution</p>
      <p>Representation
(Resolution N)</p>
      <p>Commonsense
knowledge + theories
of cognition, learning</p>
      <p>Non−monotonic
Logical reasoning
Probabilistic
Execution
Probabilistic models
of uncertainty
Labels
(training phase)</p>
      <p>Features
extraction
Deincdisuiocntiotnree Current state</p>
      <p>New axioms</p>
      <p>Answer set
Classification
block</p>
      <p>ASP
program</p>
      <p>Real scenes</p>
      <p>Baxter
Plan</p>
      <p>Goal
Answer set,
domain
knowledge</p>
      <p>Text/Audio
processing
Raliextlieeorvmaalnss,t Ptreoxctessed</p>
      <p>Program
analyzer
Outputs:</p>
      <p>Output labels
(occlusion, stability)</p>
      <p>
        Explanations
(relational description)
and executability conditions. The domain’s history includes observations, action executions, and
prioritized defaults. For any given task, the robot plans and executes actions at two resolutions,
but constructs on-demand relational descriptions of decisions and beliefs at other resolutions.
Knowledge representation and reasoning: With two resolutions, the robot represents and
reasons with commonsense domain knowledge, including cognitive theories, in the coarse-resolution.
For example, a robot fetching objects in an ofcfie building reasons about the knowledge it has
about some attributes and default room locations of objects. It also has a adaptive theory of
intentions encoding principles of non-procrastination and persistence. The fine-resolution transition
diagram is defined as a refinement of the coarse-resolution transition diagram, introducing a theory
of observations that models the robot’s ability to sense the values of domain fluents. A robot in an
office building would (for example) now consider grid cells in rooms and object parts, attributes
that were previously abstracted away by the designer. The definition of refinement guarantees
that for any given coarse-resolution transition, there exists a path in the fine-resolution diagram
between states that are refinements of the coarse-resolution states. Also, the refined diagram is
randomized to model non-determinism in action outcomes. For any given goal, non-monotonic
logical reasoning at the coarse-resolution provides a plan of intentional abstract actions; this is
achieved using Answer Set Prolog, a declarative programming paradigm [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The robot
implements each abstract transition as a sequence of concrete actions by automatically identifying (i.e.,
zooming to) and reasoning with the relevant part of the fine-resolution diagram. Execution in the
ifne-resolution uses probabilistic models of the uncertainty (e.g., in perception, actuation), with
the outcomes added to the coarse-resolution history for subsequent reasoning [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
Interactive learning and transparency: Reasoning with incomplete knowledge can result in
incorrect or suboptimal outcomes. State of the art machine learning methods (e.g., using deep
networks) require many labeled examples and considerable computational resources that are
often not available in practical robot domains. The architecture supports three strategies for
incremental, efcfiient acquisition of knowledge of previously unknown action capabilities and
axioms: (i) verbal descriptions of observed behavior; (ii) exploration of new transitions; and
(iii) reactive exploration of unexpected transitions. These strategies are formulated as suitable
interactive (e.g., inductive, reinforcement) learning problems. Reasoning and learning guide each
other, enabling the automatic identification and use of only the relevant information to construct
mathematical models for the different formulations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. For example, to estimate the stability
of objects in a scene, the robot first attempts to reason with domain knowledge and information
(e.g., object category, spatial relations) extracted from input images. Relevant regions of interest
are automatically extracted from images for which reasoning is unable to make a decision (or
makes an incorrect decision), and used to train a deep network. Information from these regions
is also used to induce axioms used for subsequent reasoning—Figure 1(right). This approach
substantially improves reliability and efficiency in comparison with deep network methods [
        <xref ref-type="bibr" rid="ref3 ref6">3, 6</xref>
        ].
      </p>
      <p>
        The architecture supports transparent reasoning and learning, i.e., explainable agency, by
encoding a theory of explanations comprising: (i) claims about representing, reasoning with, and
learning knowledge to support relational descriptions of decisions and beliefs; (ii) a
characterization of explanations based on representational abstraction, and explanation specificity and
verbosity; and (iii) a methodology for constructing explanations. This theory is implemented
in conjunction with the components summarized above—see Figure 1(right). The robot then
provides on-demand relational descriptions of decisions and beliefs in response to different types
of questions (e.g., descriptive, contrastive, counterfactual) posed by a human. The human is able
to interactively obtain descriptions at the desired abstraction, specificity, and verbosity, with the
robot posing disambiguation questions to the human as needed [
        <xref ref-type="bibr" rid="ref3 ref7">3, 7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Execution Trace</title>
      <p>The following execution traces demonstrate the working of the architecture.</p>
      <p>Execution Example 1. [Planning and Learning]
Consider a robot in a  that is asked to fetch a cup.</p>
      <p>• The plan of abstract actions: (1, ℎ), (1, ), (1, ),
(1, ), is based on the default knowledge that cups are usually in the ℎ.
• For each abstract transition, the relevant (zoomed) fine-resolution description is identified,
e.g., only cells in the  and the ℎ are relevant to the first , and used to
obtain a probabilistic policy that is invoked repeatedly to execute a sequence of concrete
actions, e.g., robot is in a cell in the ℎ after first .
• Suppose the robot’s attempt to pick up a cup in the kitchen fails. Using the knowledge
that the cup is ℎ and its its arm is ℎ, the robot learns the axiom: impossible
(1, ) if (1, ℎ), _ℎ(, ℎ),
• When asked to provide a detailed description after plan execution, the robot revises the
abstraction level to use the fine-resolution description.</p>
      <p>Human: “Please describe the executed plan in detail.”
Robot: “I moved to cell 2 in the ℎ. I picked the large cup by its handle from the
counter [...] I moved to cell 4 of the . I put the cup down on the red table.”
Execution Example 2. [Explanation and Disambiguation]
In the simulated scenario in Figure 2, the human asks the robot to “Move the yellow object on the
green cube.”. The reference to yellow object is ambiguous, and the robot asks for clarification.
• Robot: “Should I move yellow duck on the green cube?”</p>
      <p>Human: “No. Move yellow cylinder on the green cube.”
• The robot computes a plan: pick up green mug; put green mug on table; pick up red cube;
put red cube on table; pick up yellow cube; put yellow cube on table; pick up yellow
cylinder; put yellow cylinder on green cube.
• The robot traces beliefs and axioms to answer questions after plan execution.</p>
      <p>
        Human: “Why did you not pick up red cube at step1?”
Robot: “Because the red cube was below the green mug.”
Human: “Why did you move yellow cube to the table?”
Robot: “I had to put the yellow cylinder on top of the green cube. The green cube was
below the yellow cube.”
Summary: Implementing principles such as stepwise iterative
refinement and relevance, and exploiting the interplay between
representation, reasoning, and learning, are key steps towards achieving cognitive
adequacy in architectures for robots. Such an architecture that
combines knowledge-based reasoning and data-driven learning has
providing promising results in simulation and on physical robots [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref5">1, 2, 3, 5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This work is the result of research threads pursued in collaboration with Ben Meadows, Tiago
Mota, Heather Riley, Rocio Gomez, Michael Gelfond, Shiqi Zhang, and Jeremy Wyatt. This
work was supported in part by the U.S. ONR Awards N00014-13-1-0766, N00014-17-1-2434 and
N00014-20-1-2390, AOARD award FA2386-16-1-4071, and U.K. EPSRC award EP/S032487/1.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sridharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Riley</surname>
          </string-name>
          ,
          <article-title>What do you really want to do? Towards a Theory of Intentions for Human-Robot Collaboration</article-title>
          ,
          <source>Annals of Mathematics and Artificial Intelligence</source>
          ,
          <source>special issue on commonsense reasoning 89</source>
          (
          <year>2021</year>
          )
          <fpage>179</fpage>
          -
          <lpage>208</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sridharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gelfond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Wyatt,
          <article-title>REBA: A Refinement-Based Architecture for Knowledge Representation and Reasoning in Robotics</article-title>
          ,
          <source>Journal of Artificial Intelligence Research</source>
          <volume>65</volume>
          (
          <year>2019</year>
          )
          <fpage>87</fpage>
          -
          <lpage>180</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sridharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leonardis</surname>
          </string-name>
          ,
          <article-title>Integrated Commonsense Reasoning and Deep Learning for Transparent Decision Making in Robotics</article-title>
          , Springer Nature CS
          <volume>2</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gebser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kaminski</surname>
          </string-name>
          , B. Kaufmann, T. Schaub, Answer Set Solving in Practice,
          <source>Synthesis Lectures on Artificial Intelligence and Machine Learning</source>
          , Morgan Claypool Publishers,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sridharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Meadows</surname>
          </string-name>
          ,
          <article-title>Knowledge Representation and Interactive Learning of Domain Knowledge for Human-Robot Collaboration</article-title>
          ,
          <source>Advances in Cognitive Systems</source>
          <volume>7</volume>
          (
          <year>2018</year>
          )
          <fpage>77</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Riley</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Sridharan, Integrating Non-monotonic Logical Reasoning and Inductive Learning With Deep Learning for Explainable Visual Question Answering, Frontiers in Robotics and AI, special issue on Combining Symbolic Reasoning and Data-Driven Learning for Decision-Making 6 (</article-title>
          <year>2019</year>
          )
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sridharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Meadows</surname>
          </string-name>
          ,
          <article-title>Towards a Theory of Explanations for Human-Robot Collaboration</article-title>
          ,
          <source>Kunstliche Intelligenz</source>
          <volume>33</volume>
          (
          <year>2019</year>
          )
          <fpage>331</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>