<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Associative Reasoning for Com monsense Knowledge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Schon</string-name>
          <email>schon@uni-koblenz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Web Science and Technologies, Universität Koblenz</institution>
          ,
          <addr-line>Universitätsstraße 1, 56070 Koblenz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Associative reasoning refers to the human ability to focus on knowledge that is relevant to a particular problem. In this process, the meaning of symbol names plays an important role: when humans focus on relevant knowledge about the symbol ice, similar symbols like snow also come into focus. In this paper, we model this associative reasoning by introducing a selection strategy that extracts relevant parts from large commonsense knowledge sources. This selection strategy is based on word similarities from word embeddings and is therefore able to take the meaning of symbol names into account. We demonstrate the usefulness of this selection strategy with a case study from creativity testing.</p>
      </abstract>
      <kwd-group>
        <kwd>selection strategies</kwd>
        <kwd>commonsense knowledge</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        According to Kahneman [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], humans rely on two diferent systems for reasoning. System 1
is fast, emotional and less accurate, system 2 is slow, more deliberate and logical. Reasoning
with system 2 is much more dificult and exhausting for humans. Therefore, system 1 is usually
used first to solve a task and system 2 is only used for demanding tasks. Examples of tasks that
system 1 does are solving simple math problems like 3 + 3, driving a car in an empty street,
or associating a certain profession with the description like a quiet, shy person who prefers to
deal with numbers rather than people. Especially, the associative linking of information with
each other falls within the scope of system 1. In contrast, we typically use system 2 for tasks
that require our full concentration such as driving in a crowded downtown area or following
complex logical reasoning.
      </p>
      <p>Humans have vast amounts of background knowledge that they skillfully use in reasoning. In
doing so, they are able to focus on knowledge that is relevant for a specific problem. Associative
thinking and priming play an important role in this process. These are things handled by system
1. The human ability of focusing on relevant knowledge is strongly dependent on the meaning
of symbol names. When people focus on relevant background knowledge for a statement like
The pond froze over for the winter., similarities of symbols play an important role. For this
statement, a human will certainly not only focus on background knowledge that relates exactly
to the terms pond, froze, and winter, but also knowledge about similar term such as ice and snow.
We refer to the process of focusing on relevant knowledge as associative reasoning.
nEvelop-O
LGOBE</p>
      <p>If we want to model the versatility of human reasoning, it is necessary to model not only
diferent types of reasoning such as deductive, abductive, and inductive reasoning, but also the
ability to focus on relevant background knowledge using associative reasoning.</p>
      <p>
        This aspect of human reasoning, focusing on knowledge relevant to a problem or situation, is
what we model in this paper. For this purpose, we develop a selection strategy that extracts
relevant parts from a large knowledge base containing background knowledge. We use background
knowledge formalized in knowledge graphs like ConceptNet [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], ontologies like Adimen SUMO
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and Cyc [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], or knowledge bases. A nice property of commonsense knowledge sources is
that the used symbol names are often based on natural language words. For example in Adimen
SUMO you find symbol names like c__SecondarySchool. To model the associative nature of
human focusing, we exploit this nice property and use word similarities from word embeddings.
      </p>
      <p>In word embeddings, large amounts of text are used to learn a vector representation of words.
These so called word vectors have the nice property that similar words are represented by
similar vectors. We propose a representation of background knowledge in terms of vectors such
that similar statements in the background knowledge are represented by similar vectors.</p>
      <p>Based on this vector representation of background knowledge, we present a new selection
strategy, the vector-based selection that pays attention to the meaning of symbol names and thus
models associative reasoning as it is done by humans. The main contributions of this paper are:
• The introduction of the vector-based selection strategy, a statistical selection technique
for commonsense knowledge which is based on word embeddings.
• A case study using benchmarks for creativity testing in humans which demonstrates that
the vector-based selection allows to model associative reasoning and selects commonsense
knowledge in a very focused way.</p>
      <p>The paper is structured as follows: after discussing related work in Sec. 2 and preliminaries
in Sec. 3, we briefly revise SInE, a selection strategy for first-order logic reasoning with large
theories in Sec. 4. Next, we turn to the integration of statistical information into selection
strategies in Sec. 5 where, after revising distributional semantics , we introduce the vector-based
selection strategy. In Sec. 6 we present experimental results. Finally, we discuss future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>Selecting knowledge that is relevant to a specific problem is also an important task in automatic
theorem proving. In this area, often a large set of axioms called a knowledge base is given as
background knowledge, together with a much smaller set of axioms  1, … ,   and a query  .
The reasoning task of interest is to show that the knowledge base together with the axioms
 1, … ,   implies the query  . This corresponds to showing that  1 ∧ … ∧   →  is entailed
by the knowledge base.  1 ∧ … ∧   →  is usually referred to as goal. As soon as the size
of the knowledge base forbids to use the entire knowledge base to show that  follows from
the knowledge base using an automated theorem prover, it is necessary to select the axioms
from the knowledge base that are necessary for this reasoning task. However, identifying these
axioms is not trivial, so common selection strategies are based on heuristics and are usually
incomplete. This means that it is not always possible to solve the reasoning task with the
The pond froze over for the winter. What happened as a result?
1. People brought boats to the pond.</p>
      <p>2. People skated on the pond.
selected axioms: If too few axioms have been selected, the prover cannot find a proof. If too
many have been selected, the reasoner may be overwhelmed with the set of axioms and run
into a timeout.</p>
      <p>
        Most strategies for axiom selection are purely syntactic like the SInE selection [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], lightweight
relevance filtering [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and axiom relevance ordering [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A semantic strategy for axiom selection
is SRASS [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] which is a model-based approach. This strategy is based on the computation of
models for subsets of the axioms and consecutively extends these sets. Another interesting
direction of research is the development of metrics for the evaluation of selection techniques
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] which allow to measure the quality of selection strategies without having to actually run
the automated theorem prover on the selected axioms and the conjecture at hand Another
approach to axiom selection is the use of formula metrics [10] which measure the dissimilarity
of diferent formulae and lead to selection strategies which allows to select the  axioms from a
knowledge base most similar to a given problem. None of the selection methods mentioned so
far in this section take the meaning of symbol names into account.
      </p>
      <p>An area where the meaningfulness of symbol names was evaluated is the semantic web
[11]. The authors come to the conclusion that the semantics encoded in the names of IRIs
(Internationalized Resource Identifiers) carry a kind of social semantics which coincides with
the formal meaning of the denoted resource.</p>
      <p>Similarity SInE [12] is an extension of SInE selection which uses a word embedding to
take similarity of symbols into account. By this mixture of syntactic and statistical methods,
Similarity SInE represents a hybrid selection approach. In contrast, the vector-based selection
presented in this paper is a purely statistical approach.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Preliminaries and Task Description</title>
      <p>Numerous sets of benchmarks exist for the area of commonsense reasoning. Typically these
problems are multiple choice questions about everyday situations which are given in natural
language. Fig. 1 shows a commonsense reasoning problem from the choice of plausible
alternative challenge (COPA) [13]. Usually, for these commonsense reasoning problems it is not
the case that one of the answer alternatives can actually be logically inferred. Often only one
of the answer alternatives is more plausible than the others. To solve these problems, a broad
background knowledge is necessary. For the example given in Fig. 1, knowledge about winter,
ice, frozen surfaces and boats is necessary. In humans, system 1 with associative reasoning is
responsible to focus on relevant background knowledge for a specific problem.</p>
      <p>In this paper, we aim at modeling the human ability to focus on background knowledge
relevant for a specific task. We introduce a selection strategy based on word embeddings to
achieve this. For this, we assume that the background knowledge is given in first-order logic.
One reason for this assumption is that this allows to use already existing automated theorem
provers for modeling human reasoning in further steps. Furthermore, this allows to easily
compare our approach to selection strategies for first-order logic theorem proving like SInE.
Moreover, this assumption is not a limitation, since knowledge given in other formes like for
example in the form of a knowledge graph can be easily transformed into first-order logic [ 14].</p>
      <p>We furthermore assume that the description of the commonsense reasoning problem is given
as a first-order logic formula. Again, this is not a limitation since, for example, the KnEWS [ 15]
system can convert natural language into first-order logic formulas. Following terminology from
ifrst-order logic reasoning, we refer to the formula for the commonsense reasoning problem
as goal. Referring to the example from Fig. 1, we would denote the first-order logic formula
for the statement The pond froze over for the winter. as  , the formula for People brought boats
to the pond. as  1, and the formula for People skated on the pond. as  2. This leads to the two
goals  →  1 and  →  2 for the commonsense reasoning problem from Fig. 1. For these goals,
we could now select from knowledge bases with background knowledge using first-order logic
axiom selection techniques.</p>
      <p>Axiom selection for a given goal in first-order logic as described at the beginning of Sect. 2 is
very similar to the problem of selecting background knowledge relevant for a specific problem
in commonsense reasoning. Both problems have in common that large amounts of background
knowledge are given that is too large to be considered completely. The main diference is the
fact that in commonsense reasoning we cannot necessarily assume that a proof for a certain
goal can be found. Therefore, drawn inferences are also interesting in this domain. In both
cases, the task is to select knowledge that is relevant for the given goal formula.</p>
      <p>In the case study in Sect. 6, we will compare the vector-based selection strategy presented in
Sect. 5 with the syntax-based SInE selection strategy which is broadly used in first-order logic
theorem proving. Therefore, we briefly introduce the SInE selection in the next section.</p>
      <p>In the following we denote the set of all predicate and function symbols occurring in a
formula  by sym(F ). We slightly exploit notation and use sym(KB) for the set of all predicate
and function symbols occurring in a knowledge base KB.</p>
    </sec>
    <sec id="sec-4">
      <title>4. SInE: a Syntax-Based Selection Strategy</title>
      <p>
        In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] the SInE selection strategy is introduced which is successfully used by many automated
theorem provers. Since this selection strategy does not consider the meaning of symbol names,
we classify this strategy as a syntax-based selection. The basic idea of SInE is to determine a set
of symbols for each axiom in the knowledge base which is allowed to trigger the selection of
this axiom. For this a trigger relation is defined as follows:
Definition 4.1 (Trigger relation for the SInE selection [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] ). Let KB be a knowledge base,  be
an axiom in KB and  ∈ sym(A) be a symbol. Let furthermore occ(s, KB) denote the number of
axioms in which  occurs in KB and  ∈ ℝ ,  ≥ 1 . Then the triggers relation is defined as
triggers(, ) if for all symbols  ′ occurring in  we have (,
KB) ≤  ⋅ (
′, KB) (1)
      </p>
      <p>Note that an axiom can only be triggered by symbols occurring in the axiom. Parameter 
specifies how strict we are in selecting the symbols that are allowed to trigger an axiom. For
 = 1 (the default setting of SInE), a symbol  may only trigger an axiom  if there is no symbol
 ′ in  that occurs less frequently in the knowledge base than  . This prevents frequently
occurring symbols such as subClass and instanceOf from being allowed to trigger all axioms
they occur in.</p>
      <p>The triggers relation is then used to select axioms for a given goal. The basic idea is that
starting from the symbols occurring in the goal, the symbols occurring in the goal are considered
to be relevant and an axiom  is selected if  is triggered by some symbol occurring in the
set of relevant symbols. The symbols occurring in the selected axioms are added to the set of
relevant symbols and if desired, the selection can be repeated.</p>
      <p>
        Definition 4.2 (Trigger-based selection [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] ). Let KB be a knowledge base,  be an axiom in
KB and  ∈ sym(KB). Let furthermore  be a goal to be proven from KB.
      </p>
      <sec id="sec-4-1">
        <title>1. If  is a symbol occurring in the goal  , then  is 0-step triggered. 2. If  is  -step triggered and  triggers  (triggers(s, A)), then  is  + 1 -step triggered. 3. If  is  -step triggered and  occurs in  , then  is  -step triggered, too. An axiom or a symbol is called triggered if it is  -step triggered for some  ≥ 0 .</title>
        <p>For a given knowledge base, goal  and some  ∈ ℕ SInE selects all axioms which are  -step
triggered. In the following the SInE selection selecting all  -step triggered axioms is called SInE
with recursion depth  .</p>
        <p>SInE selection can also be used in commonsense reasoning to select background knowledge
relevant to a statement: to do this, we just need to convert this statement into a first-order logic
formula, and use the formula as a goal and select with SInE for it.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Use of Statistical Information for the Selection of Axioms</title>
      <p>SInE selection completely ignores the meaning of symbol names. For SInE it makes no diference
whether a predicate is called  or dog. If we consider knowledge bases with commonsense
knowledge, the meaning of symbol names provides information that can be exploited by a
selection strategy. For example, the symbol dog is more similar to the symbol puppy than to
the symbol car. If a goal containing the symbol dog is given, it is more reasonable to select
axioms containing the symbol puppy than axioms containing the symbol car. This corresponds
to human associative reasoning, which also takes into account the meaning of symbol names
and similarities.</p>
      <sec id="sec-5-1">
        <title>5.1. Distributional Semantics</title>
        <p>To determine the semantic similarity of symbol names, we rely on distributional semantics of
natural language, which is used in natural language processing. The basic idea of distributional
semantics is best explained by a quote from Firth, one of the founders of this approach:</p>
        <sec id="sec-5-1-1">
          <title>You shall know a word by the company it keeps. [16]</title>
          <p>The basis of distributional semantics is the distributional hypothesis [17], according to which
words with similar distributional properties on large texts also have similar meaning. In other
words: Words that occur in a similar context are similar.</p>
          <p>An approach used in many domains which is based on the distributional hypothesis are
word embeddings [18, 19]. Word embeddings map the words of a vocabulary to vectors in ℝ .
Typically, word embeddings are learned using neural networks on very large text sets. WeSince
we use existing word embeddings in the following, we do not go into the details of creating
word embeddings. An interesting property of word embeddings is that semantic similarity of
words corresponds to the relative similarities of the vector representations of those words. To
determine the similarity of two vector representations the cosine similarity is usually used.
Definition 5.1 (Cosine similarity of two vectors). Let ,  ∈ ℝ  , both non-zero. The cosine
similarity of  and  is defined as:
cos_sim(,  ) =</p>
          <p>⋅ 
|||| || ||</p>
          <p>
            The cosine similarity of two vectors  and  takes values between -1 and 1. For exactly
opposite vectors the value is -1, for orthogonal vectors the value is 0 and for equal vectors the
value 1. The more similar two vectors are, the greater is their cosine similarity. For example in
the ConceptNet Numberbatch word embedding [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ], the cosine similarity of dog and puppy is
0.84140545 which is much larger than the cosine similarity of dog and car that is 0.13056317.
Based on these similarities, word embeddings can furthermore be used to determine the  words
in the vocabulary most similar to a given word for some  ∈ ℕ .
          </p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Vector-Based Selection: A Statistical Selection Strategy</title>
        <p>Word embeddings represent words as vectors in such a way that words that are frequently used
in a similar context are mapped to similar vectors. Vector-based selection aims to represent the
axioms of a knowledge base as vectors in such a way that similar axioms are mapped to similar
vectors. Where we consider two axioms of a knowledge base to be similar if they represent
similar knowledge.</p>
        <p>Fig. 2 gives an overview of the vector-based selection strategy. In a preprocessing step, vector
representations are computed for all axioms of the knowledge base using an existing word
embedding. This preprocessing step has to be performed only once. Given a goal  for which we
want to check if it is entailed by the knowledge base, we transform  into a vector representation
using the same word embedding as for the vector transformation of the knowledge base. Next,
vector-based selection determines the  vectors in the vector representation of the knowledge
base most similar to the vector representation of goal  . The corresponding  axioms form the
result of the selection. Various metrics can be used for determining the  vectors that are most
similar to the vector representation of  . We use cosine similarity, which is also widely used in
word embeddings.</p>
        <p>One way to represent an axiom as a vector is to look up the vectors of all the symbols occurring
in the axiom in the word embedding and represent the axiom by the average of these vectors.
However this treats all symbols occurring in an axiom equally. This is not always useful, as the
axiom in Fig. 3 from Adimen SUMO illustrates for which it seems desirable that the symbols
instance, agent and patient contribute less to the computation of the vector representation
than the symbols carnivore, eating and animal. The reason for this lies in the frequency of the
symbols in the knowledge base which are given in the Table in Fig. 3. Symbols carnivore, eating
and animal occur much less frequently in Adimen SUMO than instance, agent and patient. This
suggests that carnivore, eating and animal are more important for the statement of the axiom.
This is similar to the idea in SInE that only the least common symbol in an axiom is allowed to
trigger the axiom. We implement this idea in the computation of the vector representation of
axioms by weighting the influence of a symbol using inverse document frequency (idf). In the
area of information retrieval, for the task of rating the importance of word  to a document
 in a set of documents  , idf is often used to diminish the weight of a word that occurs very
frequently in the set of documents. Assuming that there is at least one document in  , in which
 occurs, idf ( , ) is defined as:
idf ( , ) =
log
|{ ∈  ∣ 
||
occurs in }|
If  occurs in all documents in  , the fraction is equal to 1 and idf ( , ) = 0 . If  occurs in
only one of the documents in  , the fraction is equal to || and idf ( , ) &gt; 0 . The higher the
proportion of documents in which  occurs, the lower idf ( , ) .</p>
        <p>We transfer this idea to knowledge bases by interpreting a knowledge base as a set of
documents and each axiom in this knowledge base as a document. The resulting computation
of idf for a symbol in a knowledge base is given in Def. 5.2. For the often used tf-idf (term
frequency - inverse document frequency) the idf value is multiplied by the term frequency
of a term in a certain document. However since the number of occurrences of a symbol in a
single axiom does not necessarily correspond to its importance to the axiom (as illustrated by
the axiom given in Fig. 3), we omit this multiplication and use idf for the weighting instead.
Multiplying the idf value of a symbol with its tf value in a formula could even increase the
∀ ,  , 
((( ,  ) ∧ ( , )
∧ ( ,  ) ∧ ( ,  )
) → ( , )
)
Symbol Name:
Frequency:
instance
4237
agent
140
patient
183
carnivore
5
eating
6
animal
63
influence of frequent symbols like instance, since they often appear more than once in a formula.</p>
        <p>For simplicity, we assume that sym(F ) is a subset of the vocabulary of the word embedding
in the following definition.</p>
        <p>Definition 5.2 (idf-based vector representation of an axiom, a knowlege base). Let KB =
{ 1, … ,   },  ∈ ℕ be a knowledge base,  ∈ KB be an axiom.  be a vocabulary and  ∶  → ℝ 
a word embedding. Let furthermore sym(F ) ⊆  . The idf value for a symbol  ∈ sym( ) w.r.t.
KB is defined as |KB|
idf (, KB) = log</p>
        <p>|{ ′ ∈ KB ∣  ∈ sym(F ′)}|
The idf-based vector representation of  is defined as</p>
        <p>∑∈ sym(F ) idf (, KB) ⋅  ()
 idf ( ) =</p>
        <p>∑∈ sym(F ) idf (, KB)
Furthermore,  idf (KB) = { idf ( 1), … ,  idf (  )} denotes the idf-based vector representation of
KB.</p>
        <p>Note that this definition completely ignores the structure of axioms resulting that axiom
∀ ( animal( ) ∧ flufy ( )) is represented by the same vector as ∀ ( anima( ) ∨ flufy ( )) .
However, this is not a disadvantage, since our goal is a selection of axioms that matches the
topic of a goal, and therefore we need the vector representation of an axiom to represent only
the topic and not the exact statement of the axiom.</p>
        <p>Given a goal  and a knowledge base, we can use the vector representations of the
knowledge base and  to select the  axioms from the knowledge base most similar to the vector
representation of  for some  ∈ ℕ (see Fig. 2).</p>
        <p>Definition 5.3 (Vector-based selection). Let KB be a knowledge base,  be a goal with sym() ⊆
sym( ) and  ∶  → ℝ  a word embedding. Let furthermore  KB be a vector representation
of KB and   a vector representation for  both constructed using  . For  ∈ ℕ ,  ≤ | KB| the 
axioms in KB most similar to  are given as
mostsimilar (KB, , ) = {</p>
        <p>1, … ,   ∣ { 1, … ,   } ⊆ KB and
∀ ′ ∈ KB ⧵ { 1, … ,   }
cos_sim(  ′,   ) ≤ min cos_sim(   ,   )}.</p>
        <p>=1,…,
For KB,  and  ∈ ℕ given as above described, vector-based selection selects mostsimilar (KB, , ) .</p>
        <p>Def. 5.3 is intentionally very general and allows other vector representations besides idf-based
vector representation. Furthermore, the similarity measure cos_sim can be easily replaced by
some other measure like euclidean distance.</p>
        <p>
          In the previous section we assumed the set of symbols in a knowledge base to be a subset of
the vocabulary of the used word embedding. However in practice this is not always the case
and in many cases it might be necessary to construct a mapping for this. Each combination of
knowledge base and word embedding requires a specific mapping. As an example we describe
in [20] how we generated diferent mappings to relate the symbols in knowledge base Adimen
SUMO [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] to the vocabulary of the ConceptNet Numberbatch word embedding. For the case
study we present in the next section such a mapping is not necessary which is why we refrain
from presenting it.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Evaluation: A Case Study on Commonsense Knowledge</title>
      <p>In areas where commonsense knowledge is used as background knowledge, automated theorem
provers can be used not only for finding proofs, but also as inference engines. One reason
for this is that even if there are large ontologies and knowledge bases with commonsense
knowledge, this knowledge is still incomplete. Therefore, it is likely that not all the information
needed for a proof is represented. Nevertheless, automated theorem provers can be very helpful
on commonsense knowledge, because the inferences that a prover can draw from a problem
description and selected background knowledge provide valuable information. How well these
inferences fit the problem description depends strongly on the selected background knowledge.
Here it is very important that the selected background knowledge is broad enough but still
focused.</p>
      <sec id="sec-6-1">
        <title>6.1. Functional Remote Association Tasks</title>
        <p>The benchmark problems we use to evaluate the vector-based selection introduced in this paper
are the functional Remote Association Tasks (fRAT) [21] which were developed to measure
human creativity. In fRAT, three words like tulip, daisy and vase are given and the task is to
ifnd a fourth connecting word, called target word (here flower ). The words are chosen in such a
way that a functional connection must be found between the three words and the target word.
To solve these problems, a broad background knowledge is necessary. The solution of the above
fRAT task requires the background knowledge that tulips and daisies are flowers and that a
vase is a container in which flowers are kept.</p>
        <p>The dataset [22] used for this evaluation consists of 48 fRAT tasks. Tab. 1 gives some examples
for tasks in the dataset.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Experimental Results</title>
        <p>
          For an fRAT task consisting of the words  1,  2,  3 and the target word   , we first generate a
simple goal
 1( 1) ∧  2( 2) ∧  3( 3)
(2)
tulip, daisy, vase
sensitive, sob, weep
algebra, calculus, trigonometry
duck, sardine, sinker
finger, glove, palm
flower
cry
math
swim
hand
using the query words of the tasks as predicate and constant symbols and then select for this goal
using diferent selection strategies. Then we check whether the word   occurs in the selected
axioms. Since we only want to evaluate selection strategies on commonsense knowledge, we do
not use a reasoner in the following experiments and leave that to future work. As background
knowledge we use ConceptNet [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] which is a knowledge graph containing broad commonsense
knowledge in the form of triples. For this evaluation, we use a first-order logic translation [ 14]
of around 125,000 of the English triples of ConceptNet as knowledge base.
        </p>
        <p>We use both vector-based selection and SInE to select axioms for a goal created for an fRAT
task and then check if the target word   occurs in the selected axioms. Tab. 2 shows the results
for vector-based selection, Tab. 3 for SInE. Note that for vector-based selection the  parameter
naturally determines the number of axioms contained in the result of the selection. Since the
selected axioms are sorted in descending order with respect to the similarity to the goal in
vector-based selection, Tab. 2 furthermore provides the average position of the target word in
the selected axioms.</p>
        <p>The results for SInE in Tab. 3 show that even for recursion depth 6, were SInE selected 2045.88
axioms on average for an fRAT task, in only 37.5% of the tasks the target word occurred in the
selection. Compared to that, the result of vector-based selection of only three axioms already
contains the target word in 50% of the tasks. As soon as the vector-based selection selects more
than 235 axioms, the target word is contained in the selection for all of the tasks. The Fig. 4
illustrates the relationship between the number of axioms selected and the percentage of target
SInE on fRAT
rec.
depth
words found for the two selection strategies.</p>
        <p>Although SInE selects significantly more axioms than vector-based selection, axioms
containing the target word are often not selected. In contrast, vector-based selection is much more
focused and even small sets of selected axioms contain axioms mentioning the target word.</p>
        <p>The experiments revealed another problem specific for the task of selecting background
knowledge from commonsense knowledge bases: Since knowledge bases in this area usually are
extremely large, it is reasonable to assume that a user looking for background knowledge for a
set of keywords is not aware of the exact symbol names used in the knowledge base. Therefore
it can easily happen that a user looks for background knowledge for a set of words which do not
coincide with the symbol names used in the knowledge base. For example none of the query
words tulip, daisy and vase corresponds to a symbol name in our first-order logic translation of
ConceptNet. Therefore a selection using SInE with the goal created from these query words
results in an empty selection. In contrast to that, vector-based selection constructs a query
vector from the symbol names occurring in the goal (idf-based selection can assume the average
idf value for unknown symbols) and selects the  most similar axioms even though the query
words from the fRAT task do not occur as symbol names in the knowledge base. As long as the
query words occur in the vocabulary of the used word embedding or can be mapped to this
vocabulary, it is possible to construct the query vector and select axioms.</p>
        <p>The experiments show that vector-based selection is a promising approach for selection
on commonsense knowledge. Experiments using reasoners on the selected axioms will be
considered in future work.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion and Future Work</title>
      <p>Although humans possess large amounts of background knowledge, it is easy for them to focus
on the knowledge relevant to a specific problem. Associative reasoning plays an important
role in this process. The vector-based selection presented in this paper uses word similarities
from word embeddings to model associative reasoning. Our experiments on benchmarks for
testing human creativity show that vector-based selection is able to select in a very focused way
on commonsense knowledge. In future work, we want to use deductive as well as abductive
reasoning on the result of these selections.</p>
      <p>In another line of future work, we want to evaluate the usefulness of vector-based selection
for the task of solving benchmarks from the commonsense reasoning area like COPA [23].
[10] Q. Liu, Y. Xu, Axiom selection over large theory based on new first-order formula metrics,
Appl. Intell. 52 (2022) 1793–1807. URL: https://doi.org/10.1007/s10489-021-02469-1. doi:10.
1007/s10489- 021- 02469- 1.
[11] S. de Rooij, W. Beek, P. Bloem, F. van Harmelen, S. Schlobach, Are names meaningful?
quantifying social meaning on the semantic web, in: ISWC (1), volume 9981 of Lecture
Notes in Computer Science, 2016, pp. 184–199.
[12] U. Furbach, T. Krämer, C. Schon, Names are not just sound and smoke: Word embeddings
for axiom selection, in: CADE, volume 11716 of Lecture Notes in Computer Science, Springer,
2019, pp. 250–268.
[13] N. Maslan, M. Roemmele, A. S. Gordon, One hundred challenge problems for logical
formalizations of commonsense psychology, in: Twelfth International Symposium on
Logical Formalizations of Commonsense Reasoning, Stanford, CA, 2015.
[14] C. Schon, S. Siebert, F. Stolzenburg, Using conceptnet to teach common sense to an
automated theorem prover, in: ARCADE@CADE, volume 311 of EPTCS, 2019, pp. 19–24.
[15] V. Basile, E. Cabrio, C. Schon, KNEWS: Using Logical and Lexical Semantics to Extract
Knowledge from Natural Language, in: Proceedings of the European Conference on
Artificial Intelligence (ECAI) 2016 conference, 2016.
[16] J. R. Firth, Papers in Linguistics 1934 - 1951: Rep, Oxford University Press, 1991.
[17] G. A. Miller, W. G. Charles, Contextual correlates of semantic similarity, Language and
Cognitive Processes 6 (1991) 1–28. URL: http://eric.ed.gov/ERICWebPortal/recordDetail?
accno=EJ431389.
[18] T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, J. Dean, Distributed representations of
words and phrases and their compositionality, in: NIPS, 2013, pp. 3111–3119.
[19] T. Mikolov, K. Chen, G. Corrado, J. Dean, Eficient estimation of word
representations in vector space, CoRR abs/1301.3781 (2013). URL: http://arxiv.org/abs/1301.3781.
arXiv:1301.3781.
[20] C. Schon, Selection strategies for commonsense knowledge, 2022. URL: https://arxiv.org/
abs/2202.09163. doi:10.48550/ARXIV.2202.09163.
[21] P. M. C. Blaine R. Worthen, Toward an improved measure of remote associational ability,</p>
      <p>Journal of Educational Measurement 8 (1971) 113–123.
[22] A. Olteteanu, M. Schöttner, S. Schuberth, Computationally resurrecting the functional
remote associates test using cognitive word associates and principles from a computational
solver, Knowl. Based Syst. 168 (2019) 1–9. URL: https://doi.org/10.1016/j.knosys.2018.12.023.
doi:10.1016/j.knosys.2018.12.023.
[23] M. Roemmele, C. A. Bejan, A. S. Gordon, Choice of plausible alternatives: An evaluation
of commonsense causal reasoning, in: AAAI Spring Symposium: Logical Formalizations
of Commonsense Reasoning, AAAI, 2011.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kahneman</surname>
          </string-name>
          , Thinking, Fast and Slow, Macmillan,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Speer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Havasi</surname>
          </string-name>
          ,
          <article-title>Conceptnet 5.5: An open multilingual graph of general knowledge</article-title>
          , in: AAAI, AAAI Press,
          <year>2017</year>
          , pp.
          <fpage>4444</fpage>
          -
          <lpage>4451</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Álvez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lucio</surname>
          </string-name>
          , G. Rigau,
          <article-title>Adimen-sumo: Reengineering an ontology for first-order reasoning</article-title>
          ,
          <source>Int. J. Semantic Web Inf. Syst</source>
          .
          <volume>8</volume>
          (
          <year>2012</year>
          )
          <fpage>80</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Lenat</surname>
          </string-name>
          ,
          <article-title>Cyc: A large-scale investment in knowledge infrastructure</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>38</volume>
          (
          <year>1995</year>
          )
          <fpage>33</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hoder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Voronkov</surname>
          </string-name>
          ,
          <article-title>Sine qua non for large theory reasoning</article-title>
          , in: CADE, volume
          <volume>6803</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2011</year>
          , pp.
          <fpage>299</fpage>
          -
          <lpage>314</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Paulson</surname>
          </string-name>
          ,
          <article-title>Lightweight relevance filtering for machine-generated resolution problems</article-title>
          ,
          <source>J. Applied Logic</source>
          <volume>7</volume>
          (
          <year>2009</year>
          )
          <fpage>41</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Roederer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Puzis</surname>
          </string-name>
          , G. Sutclife,
          <article-title>Divvy: An ATP meta-system based on axiom relevance ordering</article-title>
          ,
          <source>in: CADE</source>
          , volume
          <volume>5663</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2009</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sutclife</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Puzis</surname>
          </string-name>
          , SRASS
          <article-title>- A semantic relevance axiom selection system</article-title>
          ,
          <source>in: CADE</source>
          , volume
          <volume>4603</volume>
          of Lecture Notes in Computer Science, Springer,
          <year>2007</year>
          , pp.
          <fpage>295</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Sutclife, Evaluation of axiom selection techniques</article-title>
          ,
          <source>in: PAAR+SC2@IJCAR</source>
          , volume
          <volume>2752</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>