<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>K. Kipper, A. Korhonen, N. Ryant, M. Palmer, A
large-scale classification of english verbs, Language
Resources and Evaluation</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Natural Language Question Answering with Goal-directed Answer Set Programming</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kinjal Basu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gopal Gupta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The University of Texas at Dallas</institution>
          ,
          <addr-line>Richardson, Texas</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <volume>42</volume>
      <issue>2008</issue>
      <abstract>
        <p>Understanding the meaning of a text is a fundamental challenge of natural language understanding (NLU) research. An ideal NLU system should process a language in a way that is not exclusive to a single task or a dataset. To do so, knowledge driven generalized semantic representation for English text is utmost important for any NLU applications. Ideally, for any realistic (human like) NLU system, commonsense reasoning must be an integral part of it and goal directed answer-setprogramming (ASP) is indispensable to do commonsense reasoning. Keeping all of these in mind, we have developed various NLU application ranging from visual question answering to a conversational agent. In contrast to existing purely machine learning-based methods for the same tasks, we have shown, our applications not only maintain high accuracy but also provides explanation for the answer it computes.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Answer Set Programming</kwd>
        <kwd>Natural Language Understanding</kwd>
        <kwd>Question Answering Conversational Agent</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>ented conversation, a human remembers all the details
given in the past and most of the time performs
nonThe long term goal of natural language understanding monotonic reasoning to accomplish the assigned task.
(NLU) research is to make applications, e.g., chatbots We believe that an automated QA system or a goal
oriand visual/textual question answering (QA) systems, that ented closed domain chatbot should work in a similar
act exactly like a human assistant. A human assistant way.
will understand the user’s intent and fulfill the task. The If we want to build AI systems that emulate humans,
task can be answering questions about a story or an im- then understanding natural language sentences is the
age, giving directions to a place, or reserving a table foremost priority for any NLU application. In an ideal
in a restaurant by knowing user’s preferences. Human scenario, an NLU application should map the sentence
level understanding of natural language is needed for to the knowledge (semantics) it represents, augment it
an NLU application that aspires to act exactly like a hu- with commonsense knowledge related to the concepts
man. To understand the meaning of a natural language involved–just as humans do—then use the combined
sentence, humans first process the syntactic structure of knowledge to do the required reasoning. In this paper, we
the sentence and then infer its meaning. Also, humans introduce to one of our algorithm [1] for automatically
use commonsense knowledge to understand the often generating the semantics corresponding to each English
complex and ambiguous meaning of natural language sentence using the comprehensive verb-lexicon for
Ensentences. Humans interpret a passage as a sequence of glish verbs - VerbNet [2]. For each English verb, VerbNet
sentences and will normally process the events in the gives the syntactic and semantic patterns. The algorithm
story in the same order as the sentences. Once humans employs partial syntactic matching between parse-tree
understand the meaning of a passage, they can answer of a sentence and a verb’s frame syntax from VerbNet
questions posed, along with an explanation for the an- to obtain the meaning of the sentence in terms of
Verbswer. Similarly, for visual question answering, an image Net’s primitive predicates. This matching is motivated by
should be represented in human’s mind, then it is able denotational semantics of programming languages and
to answer natural language questions by understanding can be thought of as mapping parse-trees of sentences to
the intent. Moreover, by using commonsense, a human knowledge that is constructed out of semantics provided
assistant understands the user’s intended task and asks by VerbNet. The VerbNet semantics is expressed using a
questions to the user about the required information to set of primitive predicates that can be thought of as the
successfully carry-out the task. Also, to hold a goal ori- semantic algebra of the denotational semantics.
Answering questions about a given picture, or Visual
ICLP’21: International Conference on Logic Programming, September, Question Answering (VQA) can be processed similar to
2021 the textual QA. To answer questions about a picture,
hu" Kinjal.Basu@utdallas.edu (K. Basu); gupta@utdallas.edu mans generally first recognize the objects in the picture,
(G. Gupta©) 2021 Copyright for this paper by its authors. Use permitted under Creative then they reason with the questions asked using their
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org) commonsense knowledge. To be efective, we believe
a VQA system should work in a similar way. Thus, to
perceive a picture, ideally, a system should have
intuitive abilities like object and attribute recognition and
understanding of spatial-relationships. To answer
questions, it must use reasoning. Natural language questions
are complex and ambiguous by nature, and also require
commonsense knowledge for their interpretation. Most
importantly, reasoning skills such as counting, inference,
comparison, etc., are needed to answer these questions.</p>
      <p>Here, we present out VQA work — AQuA (ASP-based
Visual Question Answering), that closely simulates the
above described way of an ideal VQA [3].</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <p>Answer Set Programming (ASP): An answer set
program is a collection of rules of the form
0 ←
1, ... , ,  +1, ... ,  .</p>
      <p>NP V NP
Example “She grabbed the rail”</p>
      <sec id="sec-2-1">
        <title>Syntax Agent V Theme</title>
      </sec>
      <sec id="sec-2-2">
        <title>Semantics Continue(E,Theme),Cause(Agent,E)</title>
      </sec>
      <sec id="sec-2-3">
        <title>Contact(During(E),Agent,Theme)</title>
        <p>2. Semantic Algebra: these are the basic domains
along with the associated operations; meaning
of a program is expressed in terms of these basic
operations applied to the elements in the domain.</p>
        <sec id="sec-2-3-1">
          <title>3. Valuation Function: these are mappings from</title>
          <p>abstract syntax trees (and possibly the semantic
algebra) to values in the semantic algebra.</p>
        </sec>
        <sec id="sec-2-3-2">
          <title>Given a program P written in language L, P’s denotation</title>
          <p>(meaning), expressed in terms of the semantic algebra,
is obtained by applying the valuation function of L to
program P’s syntax tree. Details can be found elsewhere
[8].</p>
          <p>Classical logic denotes each  is a literal [4]. In an ASP
rule, the left hand side is called the head and the
righthand side is the body. Constraints are ASP rules without VerbNet: Inspired by Beth Levin’s classification of verbs
head, whereas facts are without body. The variables start and their syntactic alternations [9], VerbNet [2] is the
with an uppercase letter, while the predicates and the largest online network of English verbs. A verb class in
constants begin with a lowercase. We will follow this VerbNet is mainly expressed by syntactic frames, thematic
convention throughout the paper. The semantics of ASP roles, and semantic representation. The VerbNet lexicon
is based on the stable model semantics of logic program- identifies thematic roles and syntactic patterns of each
ming [5]. ASP supports negation as failure [4], allowing verb class and infers the common syntactic structure and
it to elegantly model common sense reasoning, default semantic relations for all the member verbs. Figure 1
rules with exceptions, etc., and serves as the secret sauce shows an example of a VerbNet frame of the verb class
for AQuA’s sophistication. grab.
s(CASP) System: s(CASP) [6] is a query-driven,
goaldirected implementation of ASP that includes constraint
solving over reals. Goal-directed execution of s(CASP) is
indispensable for automating commonsense reasoning,
as traditional grounding and SAT-solver based
implementations of ASP may not be scalable. There are three major
advantages of using the s(CASP) system: (i) s(CASP) does
not ground the program, which makes our framework
scalable, (ii) it only explores the parts of the knowledge
base that are needed to answer a query, and (iii) it
provides natural language justification (proof tree) for an
answer [7].</p>
          <p>Denotational Semantics: In programming language
research, denotational semantics is a widely used approach
to formalize the meaning of a programming language in
terms of mathematical objects (called domains, such as
integers, truth-values, tuple of values, and, mathematical
functions) [8]. Denotational semantics of a programming
language has three components [8]:</p>
        </sec>
        <sec id="sec-2-3-3">
          <title>1. Syntax: specified as abstract syntax trees.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Commonsense Reasoning with</title>
    </sec>
    <sec id="sec-4">
      <title>Default Theories</title>
      <p>As mentioned earlier, a realistic socialbot should be able
to understand and reason like a human. In human to
human conversations, we do not always tell every detail,
we expect the listener to fill gaps through their
commonsense knowledge and commonsense reasoning. Thus, to
obtain a conversational bot, we need to automate
commonsense reasoning, i.e., automate the human thought
process. The human thought process is flexible and
nonmonotonic in nature, which means “what we believe
today may become false in the future with new knowledge”.</p>
      <p>We can model commonsense reasoning with (i) default
rules, (ii) exceptions to defaults, (iii) preferences over
multiple defaults [5], and (iv) modeling multiple worlds
[4, 10].</p>
      <p>Much of human knowledge consists of default rules,
for example, the rule: Normally, birds fly . However, there
are exceptions to defaults, for example, penguins are
exceptional birds that do not fly . Reasoning with default
rules is non-monotonic, as a conclusion drawn using a tions and does not need any annotation or generation
default rule may have to be withdrawn if more knowl- of function units such as what is employed by several
edge becomes available and the exceptional case applies. approaches proposed for the CLEVR dataset [12, 13, 14].
For example, if we are told that Tweety is a bird, we will Also, instead of predicting an answer, AQuA augments
conclude it flies. Later, knowing that Tweety is a penguin the parsed question with commonsense knowledge to
will cause us to withdraw our earlier conclusion. truly understand it and to compute the correct answer</p>
      <p>Humans often make inferences in the absence of com- (e.g., it understands that block means cube, or shiny object
plete information. Such an inference may be revised later means metal object).
as more information becomes available. This
humanstyle reasoning is elegantly captured by default rules and 4.1. Technical Approach
exceptions. Preferences are needed when there are
multiple default rules, in which case additional information AQuA represents knowledge using ASP paradigm and it
gleaned from the context is used to resolve which rule is made up of five modules that perform the following
is applicable. One could argue that expert knowledge tasks: (i) object detection and feature extraction using the
amounts to learning defaults, exceptions and preferences YOLO algorithm [11], (ii) preprocessing of the natural
in the field that a person is an expert in. language question, (iii) semantic relation extraction from</p>
      <p>Also, humans can naturally deal with multiple worlds. the question, (iv) Query generation based on semantic
These worlds may be consistent with each other in some analysis, and (v) commonsense knowledge
representaparts, but inconsistent in other parts. For example, ani- tion. AQuA runs on the query-driven, scalable s(CASP)
mals don’t talk like humans in the real world, however, in [6] answer set programming system that can provide a
the cartoon world, animals do talk like humans. So, a fish proof tree as a justification for a query being processed.
called Nemo, may be able to swim in both the real world Figure 2 shows AQuA’s architecture. The five modules
and the cartoon world, but can only talk in the cartoon are labeled, respectively, YOLO, Preprocessor, Semantic
world. Humans have no trouble separating cartoon world Relation Extractor (SRE), Query Generator, and
Commonfrom real world and switching between the two as the sit- sense Knowledge.
uation demands. Default reasoning augmented with the Preprocessor module extracts information from the
ability to operate in multiple worlds, allows one to closely question by using Stanford CoreNLP parts-of-speech
represent the human thought process. Default rules with (POS) tagger and dependency graph generator. The
outexceptions and preferences and multiple worlds can be put of the Preprocessing module will be consumed by
elegantly realized with answer set programming [4, 10] the Query Generator and the Semantic Relation
Extracand the s(CASP) system [6]. tion (SRE) modules. AQuA transforms natural language
questions to a logical representation before feeding it
to the ASP engine. The logical representation module
4. Visual Question Answering is inspired by Neo-Davidsonian formalism [15], where
every event is recognized with a unique identifier. Next,
the semantic relation labeling is the process of assigning
relationship labels to two diferent phrases in a sentence
Our work — AQuA (ASP-based Question Answering) is
an Answer Set Programming (ASP) based visual question
answering framework that truly “understands” an input
picture and answers natural language questions about
that picture [3]. This framework achieves 93.7%
accuracy on CLEVR dataset, which exceeds human baseline
performance. What is significant is that AQuA
translates a question into an ASP query without requiring
any training. AQuA replicates a human’s VQA behavior
by incorporating commonsense knowledge and using
ASP for reasoning. VQA in the AQuA framework
employs the following sources of knowledge: (i) knowledge
about objects extracted using the YOLO algorithm [11],
(ii) semantic relations extracted from the question, (iii)
query generated from the question, and (iv)
commonsense knowledge. AQuA runs on the query-driven,
scalable s(CASP) [6] answer set programming system that
can provide a proof tree as a justification for the query
being processed.</p>
      <p>AQuA processes and reasons over raw textual
ques</p>
    </sec>
    <sec id="sec-5">
      <title>5. Textual Question Answering</title>
      <p>Unlike programming languages, the denotation of a
natural language can be quite ambiguous. English is no
exception and the meaning of a word or sentence may
depend on the context. The generation of correct
knowledge from a sentence, hence, is quite hard. We have
developed a VerbNet based algorithm for semantic
generation of English text. In this section, we present a novel
approach to automatically map parse trees of simple
English sentences to their denotations, i.e., knowledge they
represent [17]. We applied this approach to construct
two NLU applications that we present here: SQuARE
(Semantic-based Question Answering and Reasoning
Engine) and StaCACK (Stateful Conversational Agent using
Commonsense Knowledge).
based on the context. To understand the CLEVR dataset
questions, AQuA requires two types of semantic rela- 5.1. Semantics-driven ASP Code
tions (i.e., quantification and property) to be extracted (if Generation
they exists) from the questions. Based on the knowledge
from a question, AQuA generates a list of ASP clauses
with the query, which runs on the s(CASP) engine to
ifnd the answer. In general, questions with one-word
answer are categorized into: (i) yes/no questions, and
(ii) attribute/value questions. Similar to a human, AQuA
requires commonsense knowledge to correctly compute
answers to questions. For the CLEVR dataset questions,
AQuA needs to have commonsense knowledge about
properties (e.g., color, size, material), directions (e.g., left,
front), and shapes (e.g., cube, sphere). AQuA will not
be able to understand question phrases such as ’... red
metal cube ...’, unless it knows red is a color, metal is a
material, and cube is a shape. Finally, the ASP engine
is the brain of our system. All the knowledge (image
representation,commonsense knowledge, semantic
relations) and the query in ASP syntax are executed using
the query-driven s(CASP) system
semantic definition in ASP. Our goal is to find the partial
matching between the sentence parse tree and the
VerbNet frame syntax and ground the thematic-role variables
so that we can get the semantics of the sentence from the
frame semantics and represent it in ASP.</p>
      <p>The illustration of the process of semantic knowledge
generation from a sentence is described in the Figure 3.</p>
      <p>We have used Stanford’s CoreNLP parser [18] to generate
the parse tree, , of an English sentence. The semantic
generator component consists of the valuation function
to map the  to its meaning. To accomplish this, we
have introduced Semantic Knowledge Generation
algorithm (Algorithm 1). First, the algorithm collects the list
of verbs mentioned in the sentence and for each verb
it accumulates all the syntactic (frame syntax) and
corresponding semantic information (thematic roles and</p>
      <p>Sentence
(John grabbed
the apple there)</p>
      <p>Stanford CoreNLP</p>
      <p>Parser
contact(during(grab),agent(john),theme(the_apple)).
continue(event(grab),theme(the_apple)).
transfer(during(grab),theme(the_apple)).
cause(agent(john),event(grab)).
...
...</p>
      <p>Sentence Semantics
Represented in ASP</p>
      <p>Parse Tree
Semantic Generator
(Valuation Function</p>
      <p>Verb
(grab)</p>
      <p>VerbNet</p>
      <p>Frames</p>
      <p>VerbNet
predicates) from VerbNet using the verb-class of the verb.</p>
      <p>The algorithm finds the grounded thematic-role variables
by doing a partial tree matching (described in Algorithm
2) between each gathered frame syntax and . From
the verb node of , the partial tree matching algorithm
performs a bottom-up search and, at each level through
a depth-first traversal, it tries to match the skeletal parse
tree of the frame syntax. If the algorithm finds an exact
or a partial match (by skipping words, e.g., prepositions),
it returns the thematic roles to the parent Algorithm 1.</p>
      <p>Finally, Algorithm 1 grounds the pre-defined predicate
with the values of thematic roles and generates ASP code.</p>
      <p>The ASP code generated by the above mentioned
approach represents the meaning of a sentence comprised
of an action verb. Since VerbNet does not cover the
semantics of the ‘be’ verbs (i.e., am, is, are, have, etc.), for
sentences containing ‘be’ verbs, the semantic generator
uses pre-defined handcrafted mapping of the parsed
information (i.e., syntactic parse tree, dependency graph,
etc.) to its semantics. Also, this semantics is represented
as ASP code. The generated ASP code can now be used
in various applications, such as natural language QA,
summarization, information extraction, Conversational
Agents (CA), etc.
5.2. SQuARE</p>
      <sec id="sec-5-1">
        <title>Question answering system for reading comprehension</title>
        <p>is a challenging task for the NLU research community.</p>
        <p>In recent times with the advancement of ML applied
to NLU, researchers have created more advanced QA
systems that show outstanding performance in QA for
reading-comprehension tasks. However, for these high
performing neural-networks based agents, the question
rises whether they really “understand” the text or not.</p>
        <p>These systems are outstanding in learning data patterns
and then predicting the answers that require shallow
Syntactic
Parse Tree</p>
        <p>(Text)
Semantic
Generator</p>
        <p>Semantic
Knowledge in</p>
        <p>ASP
Natural Language</p>
        <p>Processor
(CoreNLP &amp; spaCy)</p>
        <p>Valuation Function
Commonsense Knowledge
s(CASP)
Engine</p>
        <p>Answer
or no reasoning capabilities. Moreover, for some QA 1 contact(t3,during(grab),agent(john),
tasks, if a system claims that it performs equal or bet- theme(the_apple)).
ter than a human in terms of accuracy, then the system 2 cause(t3,agent(john),event(grab)).
must also show human level intelligence in explaining 3 transfer(t3,during(grab),
its answers. Taking all this into account, we have created theme(the_apple)).
our SQuARE QA system that uses ML based parser to Question and ASP Query: For the question - “How
generate the syntax tree and uses Algorithm 1 to trans- many objects is John carrying?”, the ASP query generator
late a sentence into its knowledge in ASP. By using the generates a generic query-rule and the specific ASP query
ASP-coded knowledge along with pre-defined generic (it uses the process template for counting).
commonsense knowledge, SQuARE outperforms other count_object(T,Per,Count)
:ML based systems by achieving 100% accuracy in 18 tasks findall(O,property(possession,T,Per,O),Os),
(99.9% accuracy in all 20 tasks) of the bAbI QA dataset set(Os,Objects),list_length(Objects,Count).
(note that the 0.01% inaccuracy is due to the dataset’s flaw, ?- count_object(t6,john,Count).
not of our system). SQuARE is also capable of generating
English justification of its answers. Answer: The s(CASP) system finds the correct answer</p>
        <p>SQuARE is composed of two main sub systems: the 1.
semantic generator and the ASP query generator. Both Justification: The s(CASP) system generated
justificasubsystems inside the SQuARE architecture (illustrated tion for this answer is shown in Figure 5.
in Figure 4) share the common valuation function.</p>
        <p>Example: To demonstrate the power of the SQuARE 5.3. StaCACK
system, we next discuss a full-fledged example showing Conversational AI has been an active area of research,
tShtoerdya:taA-flocwusatnodmtihzeedinsteegrmmeenditaotef aresstuolrtys. from the bAbI sPtAarRtRinYg[f2r0o]m, taortuhlee-rbeacseendt soypsetnemd,osmuachina,sdEatLaI-ZdAriv[1e9n]CaAnds
QA dataset about counting objects (Task-7) is taken. like Amazon’s Alexa, Google Assistant, or Apple’s Siri.
1 John moved to the bedroom. Early rule-based bots were based on just syntax analysis,
2 John got the football there. while the main challenge of modern ML based chat-bots
3 John grabbed the apple there. is the lack of “understanding” of the conversation. A
re4 John picked up the milk there. alistic socialbot should be able to understand and reason
5 John gave the apple to Mary. like a human. In human to human conversations, we
6 John left the football. do not always tell every detail, we expect the listener to
Parsed Output: CoreNLP and spaCy parsers parse each ifll gaps through their commonsense knowledge. Also,
sentence of the story and pass the parsed information our thinking process is flexible and non-monotonic in
to the semantic generator. Details are omitted due to nature, which means “what we believe today may become
lack of space, however, parsing can be easily done at false in the future with new knowledge”. We can model
https://corenlp.run/. this human thinking process with (i) default rules, (ii)
Semantics: From the parsed information, the semantic exceptions to defaults, and (iii) preferences over multiple
generator generates the semantic knowledge in ASP. We defaults [4].</p>
        <p>Start exceptions and preferences in ASP.</p>
        <p>Task-specific CAs follow a certain scheme in their
in</p>
        <p>Understand uInsteerntintent Yes TquhieryFSthMatiscailnlubsetrmatoeddeilnedFaigsuarfinei6te. Hstoatweemvearc,htihneet(aFsSkMs )i.n
IInnfcoormmpaletoen Ask preferences based on the intent each state transition are not simple as in every level it
Has Updates Verify and updCaotmeplqetue eInrfyorma on HUaspPdraetfeesrence reqSutairCesAdCiKfereancthtieyvpeess o10f0(c%omacmcuornasceynsoen) rtehaesoFnaicnegb. ook
MUoserer Udentsaailssfied ComplePtEreoxtevaicdsuketearenqsRNUduueoseseglUurtrilpv(Styssdae)atsdefiseedtaidNlNesotoamiRloserseults ThNAaosnkkfoyytohNEuenordgNNRdrooeee-Usetpuadilnaltstsgess WvrcbpaaoArenoncbptafaIoberncrduesusiltwaaserlrareooyvirngnadqitandiuagogateenats(nastdsedtieosetitsantta)lhsiosloaustgtfhiatasafivryrte[ees2Mgtt1deaiL]evmsske(ci.insghnnaIccinnetlrbudeafodaofdttiolsndelrodigctawaifoOnsoinnpnrO,geoaVcStsit:reficwaecoCtistatuitAhsoatko-nuCosus-Kf)t-.
(e.g., restaurant reservation).</p>
        <p>Figure 6: FSM for StaCACK framework Example: StaCACK is able to hold the conversation in
a more natural way by using commonsense knowledge,
which may not be possible with a rule-based system based
on a monotonic logic. Following example shows how
Sta</p>
        <p>Following the discussion above, we have created Sta- CACK can understand the cuisine preference of a user,
CACK, a general closed-domain chatbot framework. Sta- just by performing reasoning over commonsense
inforCACK is a stateful framework that maintains states by mation about a cuisine (that curry is predominant in
remembering every past dialog between the user and Indian and Thai cuisine).
itself. The main diference between StaCACK and the
other stateful or stateless chatbot models is the use of
commonsense knowledge for understanding user
utterances and generating responses. Moreover, it is capable User:
of doing non-monotonic reasoning by using defaults with</p>
        <p>Tasks</p>
        <p>Model</p>
        <sec id="sec-5-1-1">
          <title>Single Supporting Fact</title>
        </sec>
        <sec id="sec-5-1-2">
          <title>Two Supporting Facts</title>
        </sec>
        <sec id="sec-5-1-3">
          <title>Three Supporting Facts</title>
        </sec>
        <sec id="sec-5-1-4">
          <title>Two Arg. Relation</title>
        </sec>
        <sec id="sec-5-1-5">
          <title>Three Arg. Relation</title>
        </sec>
        <sec id="sec-5-1-6">
          <title>Yes/No Questions</title>
        </sec>
        <sec id="sec-5-1-7">
          <title>Counting</title>
        </sec>
        <sec id="sec-5-1-8">
          <title>Lists/Sets</title>
        </sec>
        <sec id="sec-5-1-9">
          <title>Simple Negation</title>
        </sec>
        <sec id="sec-5-1-10">
          <title>Indefinite Knowledge</title>
        </sec>
        <sec id="sec-5-1-11">
          <title>Basic Coreference</title>
        </sec>
        <sec id="sec-5-1-12">
          <title>Conjunction</title>
        </sec>
        <sec id="sec-5-1-13">
          <title>Compound Coreference</title>
        </sec>
        <sec id="sec-5-1-14">
          <title>Time Reasoning</title>
        </sec>
        <sec id="sec-5-1-15">
          <title>Basic Deduction</title>
        </sec>
        <sec id="sec-5-1-16">
          <title>Basic Induction</title>
        </sec>
        <sec id="sec-5-1-17">
          <title>Positional Reasoning</title>
        </sec>
        <sec id="sec-5-1-18">
          <title>Size Reasoning</title>
        </sec>
        <sec id="sec-5-1-19">
          <title>Path Finding</title>
        </sec>
        <sec id="sec-5-1-20">
          <title>Agent’s Motivations</title>
        </sec>
        <sec id="sec-5-1-21">
          <title>MEAN ACCURACY</title>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>The SQuARE and the StaCACK system have been tested</title>
        <p>on the bAbI QA [22] and the bAbI dialog dataset
respectively [21]. With the aim of improving NLU research,
Facebook researchers have created the bAbI datasets suit
comprising of diferent NLU application-oriented
simple task-based datasets. The datasets are designed in
such a way that it becomes easy for human to reason
and reach an answer with proper justification whereas
dificult for machines due to the lack of understanding
about the language. In the SQuARE system, the accuracy
has been calculated by matching the generated answer
with the actual answer given in the bAbI QA dataset.
Whereas, StaCACK’s accuracy is calculated on the basis
of per-response as well as per-dialog. Table 2 and table 3
compares our results in terms of accuracy with the
ex</p>
        <sec id="sec-5-2-1">
          <title>MemNN (AM+ NG+ NL)</title>
          <p>Mitra
SQuet al. ARE
believe that intelligent systems that emulate human
ability should follow this approach, especially, if we desire
true understanding and explainability.</p>
          <p>CASPR’s conversation planning is centered around a
loop in which it moves from topic to topic, and within a
topic, it moves from one attribute of that topic to another.
Thus, CASPR has an outer conversation loop to hold the
conversation at the topmost level and an inner loop in
which it moves from attribute to attribute of a topic. The
logic of these loops is slightly involved, as a user may
return to a topic or an attribute at any time, and CASPR
must remember where the user left of in that topic or
attribute. For the inner loops, CASPR uses a template,
called conversational knowledge template (CKT), that
can be used to automatically generate code that loops
over the attributes of a topic, or loops through various
dialogs (mini-CKT) that need to be spoken by CASPR for
a given topic.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Conclusion and Future Work</title>
      <p>6. Social-Bot In this paper, we discussed about our ASP based
approaches to overcome the challenges of NLU. In the
proUsing the similar technology of the StaCACK system, We cess of that we presented a visual question answering
have designed and developed the CASPR system, a social- framework — AQuA. In the textual QA domain, we
inbot designed to compete in the Amazon Alexa Socialbot troduced to our novel semantics-driven English text to
Challenge 4. CASPR’s distinguishing characteristic is answer set program generator. Also, we showed how
that it will use automated commonsense reasoning to commonsense reasoning coded in ASP can be leveraged
truly “understand” dialogs, allowing it to converse like a to develop advanced NLU applications, such as SQuARE
human. Three main requirements of a socialbot are that it and StaCACK. We make use of the s(CASP) engine, a
should be able to “understand” users’ utterances, possess query-driven implementation of ASP, to perform
reasona strategy for holding a conversation, and be able to learn ing while generating a natural language explanation for
new knowledge. We developed techniques such as con- any computed answer. At the end, we discussed about the
versational knowledge template (CKT) to approximate design philosophy behind our social-bot CASPR and how
commonsense reasoning needed to hold a conversation we have qualified to participate in the Amazon Alexa
on specific topics. Socialbot Challenge 4. As part of future work, we plan</p>
      <p>Our philosophy is to design a socialbot that emulates, to extend the SQuARE system to handle more complex
as much as possible, the way humans conduct social con- sentences and eventually handle complex stories. Our
versations. Humans employ both learned-pattern match- goal is also to develop an open-domain conversational
ing (e.g., recognizing user sentiments) and commonsense AI chatbot based on automated commonsense
reasonreasoning (e.g., if a user starts talking about having seen ing that can “converse” with a human based on “truly
the Eifel Tower, we infer that they must have traveled to understanding” that person’s utterances.
France in the past) during a conversation. Thus, ideally,
a socialbot should make use of both machine learning
as well as commonsense reasoning technologies. Our References
goal is to use the appropriate technology for a task, i.e.,
use machine learning and commonsense reasoning for
respective tasks that they are good at. Machine learning
is good for tasks such as parsing, topic modeling, and
sentiment detection while commonsense reasoning is
good for tasks such as generating a response to an
utterance. In a nutshell, we should use machine learning
for modeling System 1 thinking and commonsense
reasoning for modeling System 2 thinking [23]. We strongly
[3] K. Basu, F. Shakerin, G. Gupta, Aqua: Asp-based 55–60. doi:10.3115/v1/P14-5010.
visual question answering, in: International Sympo- [19] J. Weizenbaum, ELIZA—a computer program for
sium on Practical Aspects of Declarative Languages, the study of natural language communication
beSpringer, 2020, pp. 57–72. tween man and machine, CACM 9 (1966) 36–45.
[4] M. Gelfond, Y. Kahl, Knowledge representation, [20] K. M. Colby, S. Weber, F. D. Hilf, Artificial paranoia,
reasoning, and the design of intelligent agents: Artificial Intelligence 2 (1971) 1–25.
The answer-set programming approach, Cambridge [21] A. Bordes, Y.-L. Boureau, J. Weston, Learning
University Press, 2014. end-to-end goal-oriented dialog, arXiv preprint
[5] M. Gelfond, V. Lifschitz, The stable model semantics arXiv:1605.07683 (2016).</p>
      <p>for logic programming., in: ICLP/SLP, volume 88, [22] J. Weston, et al., Towards AI-Complete Question
1988, pp. 1070–1080. Answering: A Set of Prerequisite Toy Tasks, arXiv
[6] J. Arias, M. Carro, E. Salazar, K. Marple, G. Gupta, preprint arXiv:1502.05698 (2015).</p>
      <p>Constraint answer set programming without [23] D. Kahneman, Thinking, fast and slow, Macmillan,
grounding, TPLP 18 (2018) 337–354. doi:10.1017/ 2011.</p>
      <p>S1471068418000285.
[7] J. Arias, M. Carro, Z. Chen, G. Gupta, Justifications
for goal-directed constraint answer set
programming, arXiv preprint arXiv:2009.10238 (2020).
[8] D. A. Schmidt, Denotational semantics: a
methodology for language development, William C, Brown</p>
      <p>Publishers, Dubuque, IA, USA, 1986.
[9] B. Levin, English verb classes and alternations: A
preliminary investigation, U. Chicago Press, 1993.</p>
      <p>doi:10.1075/fol.2.1.16noe.
[10] C. Baral, Knowledge representation, reasoning and
declarative problem solving, Cambridge Uni. Press,
2003.
[11] J. Redmon, A. Farhadi, Yolov3: An incremental
im</p>
      <p>provement, arXiv preprint arXiv:1804.02767 (2018).
[12] J. Johnson, et al., Inferring and executing programs
for visual reasoning, in: Proceedings of the IEEE
International Conference on Computer Vision, 2017,
pp. 2989–2998.
[13] J. Suarez, J. Johnson, F.-F. Li, Ddrprog: A clevr
differentiable dynamic reasoning programmer, arXiv
preprint arXiv:1803.11361 (2018).
[14] K. Yi, et al., Neural-symbolic VQA: Disentangling
reasoning from vision and language understanding,
in: NIPS’18, 2018, pp. 1031–1042.
[15] D. Davidson, Inquiries into truth and interpretation:</p>
      <p>Philosophical essays, volume 2, Oxford University</p>
      <p>Press, 2001.
[16] J. Johnson, B. Hariharan, L. van der Maaten, L.
Fei</p>
      <p>Fei, C. Lawrence Zitnick, R. Girshick, Clevr: A
diagnostic dataset for compositional language and
elementary visual reasoning, in: IEEE CVPR’17,
2017, pp. 2901–2910.
[17] K. Basu, S. Varanasi, F. Shakerin, J. Arias, G. Gupta,</p>
      <p>Knowledge-driven natural language understanding
of english text and its applications, in: Proceedings
of the AAAI Conference on Artificial Intelligence,
volume 35, 2021, pp. 12554–12563.
[18] C. D. Manning, M. Surdeanu, J. Bauer, J. Finkel, S. J.</p>
      <p>Bethard, D. McClosky, The Stanford CoreNLP NLP
toolkit, in: ACL System Demonstrations, 2014, pp.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>