<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Query Processing: Estimating Relational Purity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan-Christoph Kalo</string-name>
          <email>kalo@ifis.cs.tu-bs.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Lofi</string-name>
          <email>c.lofi@tudelft.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>René Pascal Maseli</string-name>
          <email>r.maseli@tu-bs.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wolf-Tilo Balke</string-name>
          <email>balke@ifis.cs.tu-bs.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institut für Informationssysteme</institution>
          ,
          <addr-line>TU Braunschweig</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Web Information Systems Group, Delft University of Technology</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>The use of semantic information found in structured knowledge bases has become an integral part of the processing pipeline of modern intelligent information systems. However, such semantic information is frequently insufficient to capture the rich semantics demanded by the applications, and thus corpus-based methods employing natural language processing techniques are often used conjointly to provide additional information. However, the semantic expressiveness and interaction of these data sources with respect to query processing result quality is often not clear. Therefore, in this paper, we introduce the notion of relational purity which represents how well the explicitly modelled relationships between two entities in a structured knowledge base capture the implicit (and usually more diverse) semantics found in corpus-based word embeddings. The purity score gives valuable insights into the completeness of a knowledge base, but also into the expected quality of complex semantic queries relying on reasoning over relationships, as for example analogy queries.</p>
      </abstract>
      <kwd-group>
        <kwd>Semantics of Relationships</kwd>
        <kwd>LOD</kwd>
        <kwd>Structured Knowledge Repositories</kwd>
        <kwd>Word Embeddings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        To provide an intuitive and efficient user experience, future information systems
need to offer powerful query capabilities with an awareness of the semantics of the
query, as for example question answering systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or intelligent digital assistants
like MS Cortana, Google Now, or Apple Siri. Here, entities referred to in queries and
relationships between those entities play a central role in the query processing process.
Structured Knowledge Repositories like the Google Knowledge Vault [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or WordNet
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Linked Open Data sources like DBpedia [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Yago [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], usually serve as a
premier source of such semantic information. However, due to their nature as structured
knowledge bases, they are not always sufficient to implement some concepts required
for intuitive queries: for example, information on human perception (like similarity or
relatedness) or information on less clear attributes or fuzzy relationships are often
omitted. As an example, consider the entities Brad Pitt and Angelina Jolie. In addition to
their only (and outdated) DBpedia relationship spouse, their semantic relationship of
course is much more complex. There exists a plethora of additional relationships
between them, which is quite hard to model: as for example, they co-acted in the same
movies together, had many joint public appearances, and publicly split up again.
      </p>
      <p>
        Here, word embeddings (such as recent skip-n-gram, neural embeddings [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]) have
been shown to provide an interesting additional source of semantic information on
entities and their relationships: word embeddings learn a vector representation of words
used in a large natural language corpus by exploiting the distributional hypothesis [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
thus promising to encode semantics based on the actual human perception and everyday
use, implicitly provided by the language structure of the natural text corpus used for
training (like news articles or encyclopedias). For example, it has been shown that the
similarity of the resulting word vectors closely correlates with the perceived
attributional similarity of their respective real-world entities [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which allows for similarity
queries but also for mapping between diverging user and knowledge base vocabulary
for query processing in question answering [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] (i.e., when querying for “husband”, but
the knowledge base only has information on “spouses”, similarity can help to suggest
using spouse instead of husband).
      </p>
      <p>
        An interesting use case of powerful semantic queries which were (re-)popularized
by word embeddings are analogy queries. Using analogies in natural speech allows
communicating dense information easily and naturally by implying that the “essence”
of two concepts is similar or at least perceived similarly. Thus, analogies can be used
to map factual and behavioral properties from one (usually better-known concept, the
source) to another (usually less well known, the target) concept by exploiting both
attributional and relational similarity. This is particularly effective for natural querying
and explaining when only vague domain knowledge is available (e.g., “Okinawa is to
Japan as is Hawaii to the US”). It has been argued [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] that such semantics can be
expressed by simple vector arithmetics within the word embeddings space, however, it
has also been shown that the performance of this type of analogy processing varies
greatly with the type of relationships involved [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>In this paper, based on contemporary computational analogy processing theory, we
investigate the semantic relationship between vector arithmetics of word embeddings,
the relationships found in structured knowledge repositories, and analogy semantics.
The resulting contributions can be summarized as follows:
 We introduce the concept of purity scores for relationships in knowledge bases by
investigating the vectors associated with each relationship in a word embeddings
space. As word embeddings are based on rich language semantics of a large text
corpus, we assume that they contain richer (or at least different) semantic
information than structured knowledge bases, but only in implicit form. The purity score
of a relationship represents the degree to which the knowledge base covers the
implicit semantics of embeddings.
 We provide an extensive overview and examples of different relationships, their
associated vectors, as well as their related source texts which have been involved
in creating those vectors to clarify the concept of purity scores
 We show that the popular analogy reasoning technique using word embedding
vector arithmetics work well for pure relationships, while it does not work well for
impure ones.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Concepts</title>
      <p>
        In the following section, we revisit and summarize some of the core concepts
underpinning our findings. This especially covers the general semantics of 4-term analogies,
common word embeddings, and the offset method for analogical reasoning using vector
arithmetics (as we already discussed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>
        Analogies and Relational Similarity. The semantics of analogies have been
researched in depth in the fields of philosophy, linguistics, logics, and in cognitive
sciences, such as [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13–15</xref>
        ]. However, those models are rather complex and hard to grasp
computationally, and thus most recent works on computational analogy processing rely
on the simple 4-term analogy model, which is given by two sets of word pairs (the
socalled analogons), with one pair being the source and one pair being the target. A
4term analogy holds true if there is a high degree of relational similarity between those
two pairs. This is denoted by [ 1,  2] ∷ [ 1,  2], where one relationship between  1 and
 2 is similar to a relationship between  1 and  2, as for example in [US Dollar,
USA]∷[Euro, Germany]. This model has several limitations, as is discussed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]: the
semantics of “a high degree of relational similarity'' from an ontological point of view
is unclear as there can be plethora of relationships between the concepts of an analogon,
but only some of them are of relevance for valid analogy semantics.
      </p>
      <p>
        Therefore, we rely on an improved interpretation of the 4-term analogy model [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ],
and assume that there can be multiple relationships between the concepts of an
analogon, some of them being relevant for the semantics of an analogy (the defining
relationships), and some of them not. An analogy holds true if the sets of defining
relationships of both analogons show a high degree of relational similarity. For illustrating the
difference and importance of this change in semantics, consider the analogy statement:
[Tokyo, Japan]∷[Braunschweig, Germany]. Tokyo is a city in Japan, and Braunschweig
is a city in Germany, therefore both analogons contain the same ''city is located in
country'' relationship, and this could be considered a valid analogy with respect to the simple
4-term analogy model. Still, this is a poor analogy statement from a human perspective
because Braunschweig is not like Tokyo at all (therefore, this statement does neither
describe the essence of Tokyo nor the essence of Braunschweig particularly well): the
defining traits (relationships) of Tokyo in Japan should at least cover that Tokyo is the
single largest city in Japan, and its capital. There are many other cities, which are also
located in Japan, but only Tokyo has these two defining traits. Braunschweig, however,
is just a smaller city in Germany, which might stand out for either its technical
university or its historic city center (therefore, the defining relationships of both word pairs
are not very similar). The closest match to a city like Tokyo in Germany should
therefore be Berlin, which is also the largest city and the capital city.
      </p>
      <p>Understanding which relationships define the essence of an analogon as perceived
by humans is a very challenging problem, but this understanding is crucial for judging
the usefulness and value of an analogy statement. Furthermore, the degree in which
relationships are defining an analogon may vary with different contexts (e.g., the role
of Berlin in Germany in a political discussion vs. the role of Berlin in Germany in a
discussion about nightlife).</p>
      <p>
        Word Embeddings, Relational Similarity, and Analogy Processing. Word
embeddings represent each word in a predefined vocabulary with a real-valued vector, i.e.
words are embedded in a vector space (usually with 50-600 dimensions). Most word
embeddings will directly or indirectly rely on the distributional hypothesis [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] (i.e.
words frequently appearing in similar linguistic contexts will also have similar
realworld semantics), and are thus particularly well-suited to measure semantic similarity
and relatedness between words (which is one of the foundation of the 4-term analogy
definition), e.g., see [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In recent years, especially word embeddings relying on neural
networks have become popular, with the skip-gram negative sampling approach
(SGNS) [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] being one of the best known examples .
      </p>
      <p>
        The straight-forward application of word embeddings is computing similarity
between two given words [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] by measuring the cosine similarity. However, many (but not
all) word embeddings show some very interesting and surprising additional property: it
seems that not only the cosine distance between vectors represents a measure for
similarity and relatedness of the embedded words, but that also the difference vectors
between word pairs implicitly represent the relationships between two entities, and thus
carry analogy semantics [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. For example, the difference between the vector for “man”
and “king” seems to represent the concept/relationship of being a ruler, and the closest
word vector to “woman” plus the “ruler” concept vector will be “queen” (see Fig. 1;
this method is also sometimes called the offset method).
      </p>
      <p>To a certain extent, these semantics can be attributed to the distributional hypothesis:
in natural speech, concepts carrying similar semantics will frequently co-occur in
similar context. Therefore, the difference vector should implicitly encode the defining
relationships between two concepts as discussed in the previous section (i.e.:
Toyko/Japan and Berlin/Germany will likely occur in similar contexts in natural speech, while
Braunschweig/Germany will likely appear in different context and will thus have a
different difference vector). The ability of word embeddings to perform this analogical
reasoning process has been evaluated using several standardized test sets (see next
section), but is still not well understood and can fail quite often, which is related to our
introduced purity score.</p>
      <p>
        In a more formal fashion, a word embedding can be used to solve analogy
completion queries as follows [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]: Given the query [ 1,  2] ∷ [ 1, ? ], the word embedding
provides the respective word vectors ⃗⃗⃗⃗1, ⃗⃗⃗⃗2, and ⃗⃗⃗1. Then, the vector ⃗⃗⃗2 representing
the query’s solution can be determined by finding the word vector in the trained vector
      </p>
      <sec id="sec-2-1">
        <title>Tokyo</title>
      </sec>
      <sec id="sec-2-2">
        <title>Japan</title>
        <p>≈</p>
      </sec>
      <sec id="sec-2-3">
        <title>Germany</title>
      </sec>
      <sec id="sec-2-4">
        <title>Berlin</title>
      </sec>
      <sec id="sec-2-5">
        <title>Braunschweig</title>
        <p>⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗ − ⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗ + ⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗ ≈ ⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗⃗
⃗⃗⃗2 = arg</p>
        <p>max
 ∈ , ≠⃗⃗⃗⃗⃗2,
≠⃗⃗⃗⃗1
(⃗⃗⃗⃗2 − ⃗⃗⃗⃗1 + ⃗⃗⃗1)  .</p>
        <p>space  which is closest to ⃗⃗⃗⃗2 − ⃗⃗⃗⃗1 + ⃗⃗⃗1 with respect to the cosine vector distance, i.e.</p>
        <p>
          Relational Benchmarking. The Mikolov Benchmark set [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] is one of the most
popular benchmark sets for testing the analogy reasoning capabilities of word
embeddings, covering 19,558 4-term analogies. However, only 14 distinct relationships are
covered, and most of them (9) focus on grammatical properties the relationship “is
plural for a noun”, e.g., [mouse,mice]∷[dollar,dollars] or “is superlative”. Five
relationships
are
of a
semantic
nature
(i.e. “is
capital
city
for country”
[Athens,Greece]∷[Oslo,Norway], “is currency of country”, “city in state”, “male-female
version” (including the often cited [king,queen]∷[man,woman]). The test set is
generated by collecting pairs of entities which are members of the selected relationship either
manually or from Wikipedia and DBpedia, and then combining these pairs into 4-term
analogy tuples. For example, for the “city in state” relationship, 68 word pairs like
[Dallas,Texas] or [Miami,Florida] are collected, and then combined by a cross product.
The related Wordrep dataset [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] extends the Mikolov set by adding more challenges,
and expanding to 25 different relationships.
        </p>
        <p>
          For the Mikolov dataset, the authors showed that skip n-gram word embeddings [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]
using the offset
        </p>
        <p>
          method could solve analogy completion queries (i.e.
[Athens,Greece]∷[Oslo,?]) with an accuracy of 53.3% overall, 50% for semantic
relationships (like ‘capital of’), and 55.9% for syntactic ones (like ‘plural of’). No deeper
analysis of the relationships for which this technique performs well was provided. However,
this was analyzed in more detailed in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] using subjective feedback on relational
similarity from human users. Here, the authors identified as the core problem which hinders
the offset methods for analogy reasoning the presence of multiple relationships between
two entities which influence human perception, as some relationships are perceived
more dominant than others, e.g., quite often, the relationship intended in the Mikolov
analogy challenge (e.g., “city in state”) was not perceived as dominant as some other
relationship perceived by human subjects (e.g., “home of best football team in state”).
While this argument is formulated slightly differently, those experimental results
strongly support the hypothesis of defining relationships [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] discussed in the previous
sections. Also, different perception of relationships introduces problems with symmetry
or transitivity of relational similarity not holding from a user’s perspective [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. A
similar result is also supported by experiments in [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>Based on the intuition obtained in those experimental results, in the following, we
define the concept of relational purity approximating in how far a relationship given in
an analogy challenge is indeed perceived as the relevant or defining relationship, which
gives us insights both into the its suitability for analogy query processing, but can also
serve as an indirect and implicit measure for the semantic completeness of a knowledge
base with respect to that relationship.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Purity Score for Relationships</title>
      <p>In the following section, we further explain our idea of relationship purity and provide
a formal definition of the concept. We generalize the idea to relationships between
entities, and support our findings with an analysis of DBpedia relationships.</p>
      <p>Motivation. As we motivated, the semantic relationships modelled in triple format
(consisting of subject, predicate, object triples) in state-of-the-art knowledge bases are
often a stark simplification of the relationships between the respective entities as
perceived by humans. As a result, many relationships are left out (e.g., Angelina Jolie and
Brad Pitt having a public fight about their children), or several related but still
perceptually different relationships are generalized and grouped into a single relationship. In
Figure 2, we have visualized a 2-dimensional scaling of a GloVe embedding trained on
Wikipedia for the “spouse” relationship (i.e., the difference vectors between two
spouses). The more parallel those relationship vectors are, the more similar their
representation is in the embedding space, and therefore, the more similar their captured
semantics should be (as the texts in which those relationships are discussed share similar
contexts). In the case of the spouse relationship, we observe that the relationship vectors
of different couples are quite diverse: However, we can observe that the vectors of the
two actor couples (Jolie/Pitt and Heard/Depp), are more similar, since they are linked
by more than one single relationship.</p>
      <p>If, in DBpedia, we look at the entities Tolkien and The Hobbit, we observe that they
are connected by only a single relationship instance: Tolkien is the author of The
Hobbit. In the text found in Wikipedia, a whole paragraph is used to describe the
relationship between the respective entities: how he initially wrote The Hobbit for his children,
how he never planned to publish it, how his friends liked it and pressured him towards
publication, etc. Similarly, the same “is author of” relationship is used to link Max
Frisch to his novel Homo Faber, whereas the Wikipedia text offers a much deeper
insight into their relationship than the structured knowledge base does (see Fig. 2 for a
visualization).</p>
      <p>However, as defined by the DBpedia ontology, the recommended use of the “is
author of” relationship is much more general and it can connect any type of author with
their work of any kind (this can be novels, scientific papers, screenplays, music,
paintings and even software programs). Thus, it generalizes the semantics of a larger amount
of diverse subrelationships to one single relationship, not capturing their perceived
semantic differences anymore (i.e., Pablo Picasso authoring the painting “Les
Demoiselles d'Avignon” is perceived quite different than Max Frisch authoring “Homo Faber”
by most). However, we believe that this loss of semantics can be modelled and captured
by analyzing rich textual corpora. Here, for the “is author of” relationship, we claim
that the rich diversity of different types of author relationship leads to a low purity
score, while other relationships like “is currency of country” have high purity scores
(while there might be diverse relationships between countries and their currencies, this
diversity is comparably low). This purity, i.e., diversity of usage of a relationship can
be observed in word embeddings.</p>
      <p>As a further example relationship in Fig. 2, consider the “is in country” relationship
for cities: the relationship between Braunschweig and Germany for example (i.e., there</p>
      <p>Barack Obama</p>
      <p>Brad Pitt
Johnny
Depp</p>
      <p>Amber Heard</p>
      <p>Angelina Jolie</p>
      <p>Michelle Obama
Elsa Einstein</p>
      <p>UK
Braunschweig</p>
      <p>Germany</p>
      <p>London</p>
      <p>Paris
France</p>
      <p>Rome</p>
      <p>Italy</p>
      <p>Albert Einstein</p>
      <p>Pablo Picasso
Claude
Monet</p>
      <p>Chicago</p>
      <p>Picasso
ISmupnrreissesion,
Les Demoiselles Max Frisch
d’Avignon</p>
      <p>The Fire Raisers</p>
      <p>The Fellowship
of the Ring</p>
      <p>J.R.R. Tolkien</p>
      <p>The Silmarillion
The Hobbit
is nothing particularly unique about Braunschweig), is quite different from the
relationship from Paris to France. Paris is not only a city located within France, but also the
country’s largest city, the location of the French government and its capital. The
relationship vector of Rome to Italy and London to UK look very similar.</p>
      <p>Please note that all those vectors have 300-dimensions, and that thus such simple
visual analytics are quite crude due to the loss of semantics when mapped to the
2Dplane using Principal Component Analysis. Hence, we introduce a formal definition of
relational purity in the next section.</p>
      <p>Computing the Purity Score. Word embeddings were shown to provide linear
substructures that can represent implicit relational similarity of all relationships between
two entities as having similar (i.e., having a high cosine similarity) difference vectors
between their word vectors. This characteristic is mainly used for analogical query
processing using the offset method. However, we can adopt a similar notion to compute
the purity score of relationships. Given a set subject entities  and object entities  ,
which are connected by a relationships  in a structured knowledge base, we compute
the purity score of the relationship</p>
      <p>⊆  ×  as the standard deviation from the
average difference vector of the entities. Given a triple ( ,  ,  ), its relationship vector in a
word embedding is defined by the respective entities difference vector:  =  −  .
We define the average vector for relationship  , given the set of relation vectors  ⃗ as
⃗⃗⃗⃗ = ∑⃗ ∈⃗ 
| ⃗ |</p>
      <p>. Now, we define the purity score of a relationship as the standard deviation
of the cosine distance from every relationship instance vector to the average vector ⃗⃗⃗⃗ .
The cosine distance between two vectors is defined as cos( ,  ) = 1 − ⃗⃗⃗⃗⃗⃗⃗‖⃗⃗⃗2⋅⃗‖⃗⃗⃗⃗⃗⃗‖⃗⃗⃗2
‖
. Note
that negative similarity leads to distances between 1 and 2 in case the vectors are
directed in opposite directions. The purity of a relationship R is now defined as:
⃗ ⋅⃗
A high variance in the directions of the relationship vectors (represented by cosine
distances to the average vector), results in impure relational information in the embedding,
i.e. low purity scores. Similarly, a low variance in the directions (so parallel relationship
vectors) lead to a high purity score.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>In this section, we introduce our experimental setup, and specifically focus on how we
represented DBpedia entities and relationships in a corpus-based word embedding.
Afterwards, we compute the purity scores for DBpedia relationships and visualize and
discuss some examples to obtain a better intuition of the results.</p>
      <p>
        Experimental Setup. We extracted relationships between entities from the largest
Linked Open Data (LOD) data store DBpedia [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a knowledge graph that is built by
extracting knowledge from Wikipedia Infoboxes. As result, our dataset covers 18
million unique relationship instances between entities from around 1,200 relationships as
defined by the DBpedia ontology, ranging from capital city relationships, over causal
relationships, to biological relationships between living organisms.
      </p>
      <p>
        As a training corpus for our word embedding, we downloaded a dump of the English
version of Wikipedia from 01/2017. For linking DBpedia entities and relationships to
the embedding, as a first step, we performed named entity recognition and
disambiguation on the text corpus using DBpedia Spotlight [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and replace the recognized
entities by their respective DBpedia URI. For creating the word embedding, we use GloVe
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and train multiple models with varied window size as described in the original
GloVe paper between 8 and 15. Similar to their results [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we found out that semantic
relationships are better represented in the embedding when we use the larger window
size. Furthermore, we ensured that the sliding window does not reach over different
articles and sentence boundaries, preventing words appearing in wrong contexts. We
varied the number of embedding dimensions between 50, 100 and 300, finding that
DBpedia relationships are represented best by choosing 300 dimensions. Thus, the
following results all are based on a GloVe embedding with window size 15, 300
dimensions and 100 training iterations. The minimum word frequency is set to 8. Our corpus
and embeddings are available on request.
      </p>
      <p>Result Visualization. For visualizing relationship vectors and their purity, we
selected 9 DBpedia relationships from very pure to extremely impure relationships from
a 300-dimensional GloVe embedding. We used principal component analysis to scale
the results to two dimensions. The results are visualized in Figure 3. Subject entities are
visualized by a red dot, object entities by a blue triangle and the difference vector,
representing the relationship between the entities, is visualized by a green line. The
relational similarity between the relationships given by cosine distance of the difference
vectors. Hence, parallel vectors have cosine similarity 0, whereas orthogonal vectors
have similarity 1. (Due to the two-dimensional scaling of the original vectors, the cosine
similarity is not perfectly represented in the visualization.) Pure relationships (top left)
have highly parallel relationship vectors, whereas impure relationships (bottom right)
are very diverse. Since impure relationships have nearly no parallel relationship
vectors, the offset method does not lead to meaningful results. Hence, most analogy queries
based on this method return incorrect results.</p>
      <p>Purity of DBpedia Relationships. We have evaluated the purity of all DBpedia
ontology relationships for which we could find at least two relationship instances in our
embedding. This resulted in more than 400 different relationship embeddings with
different purity scores. In Table 1, we show an excerpt covering the complete purity
spectrum. Particularly pure are relationships with only a few different subjects or objects,
as for example, the biological kingdom relationship from DBpedia which connects
species to one of six different biological kingdoms. The country relationship linking cities
to their countries or states also has a high purity score of 0.81: most of the relationships
instances are parallel, having some exceptions as shown in Figure 3. The author
relationship, as already discussed in Section 3, has a purity of 0.36. Since its domain
comprises entities of very different type, the resulting relationship vectors show only few
similarities. However, we can see several different clusters of similar (parallel)
relationships, indicating that the relationship is impure (i.e., each cluster represents a
subrelationship which is not modelled directly in the knowledge base).</p>
      <p>The spouse relationship connecting two married persons is the most diverse (and
thus impurest) relationship in our dataset, having a purity score of only 0.01. This
diversity has several reasons: On the one hand, this relationship is symmetric (which is
not covered by the default similarity measurement using cosine distance), therefore the
vectors for man and his wife is directed in the opposite direction to the vector that
connects a woman to her husband. Furthermore, the persons being married and their
relationships to each other are quite different from couple to couple (see our introductory
examples) which is well represented in text but not in a knowledge base.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Summary and Future Work</title>
      <p>In this paper, we investigated the semantic interplay between explicit relationships
modelled in structured knowledge bases like DBpedia, and their representation in
corpus-based word embeddings. While word embeddings do not explicitly represent
individual relationship instances between two entities, the implicit representation of all
relationship instances as a difference vectors between two word/entity vectors can
potentially cover a much wider range of (perceptual) semantics based on the corpus used for
training. This notion is usually exploited for embedding-based relational similarity and
analogy queries. However, it can also be used to shed some light on the nature of a
specific relationship and how well it is represented in a knowledge base. To this end,
we introduced the concept of relational purity, which implicitly represents how uniform
the usage of a give relationship is in natural text. This results on several interesting
observations: some relationships (like “is spouse of”) carry much richer semantics in
their textual representations (e.g., while DBpedia just contains that Brad Pitt and
Angelina Jolie are/used to be married, texts talking about both indicate a large number of
additional potential and currently not covered relationships) indicated by a low purity
score, while hierarchical relationships from the Linné taxonomy for animal and plant
life are very pure – e.g., indicating that there is exactly one specific semantic
relationship between two entities in the taxonomy (for example, squirrels are rodents – there
are no significant other relationships mentioned in natural text connecting “squirrel”
with “rodent”).</p>
      <p>For the future, we plan to exploit our insights obtained in this work for improving
complex semantic query processing (by e.g., being able to give an assessment of the
potential reliably of an answer based on reasoning over relationships), but also for
designing processes for uncovering potential semantic gaps in knowledge bases and
mining for missing information in a targeted fashion.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ferrucci</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Chu-Carroll</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gondek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalyanpur</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lally</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murdock</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyberg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prager</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlaefer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Building Watson: An Overview of the DeepQA Project</article-title>
          .
          <source>AI Magazine</source>
          .
          <volume>31</volume>
          ,
          <fpage>59</fpage>
          -
          <lpage>79</lpage>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabrilovich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heitz</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horn</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lao</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strohmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Zhang, W.:
          <article-title>Knowledge vault: A web-scale approach to probabilistic knowledge fusion</article-title>
          .
          <source>In: Proceedings of the 20th International Conference on Knowledge Discovery and Data Mining (SIGKDD)</source>
          . pp.
          <fpage>601</fpage>
          -
          <lpage>610</lpage>
          . ACM Press, New York, New York, USA (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , G.:
          <article-title>WordNet: a lexical database for English</article-title>
          .
          <source>Communications of the ACM</source>
          .
          <volume>38</volume>
          ,
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>In: Proceedings of the 6th International Semantic Web Conference (ISWC )</source>
          . pp.
          <fpage>722</fpage>
          -
          <lpage>735</lpage>
          . Springer Berlin Heidelberg (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago: a core of semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 16th international conference on World Wide Web - WWW '07</source>
          . p.
          <fpage>697</fpage>
          . ACM Press, New York, New York, USA (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweig</surname>
          </string-name>
          , G.:
          <article-title>Linguistic Regularities in Continuous Space Word Representations</article-title>
          . In:
          <article-title>Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language (NAACL-HLT)</article-title>
          . pp.
          <fpage>746</fpage>
          -
          <lpage>751</lpage>
          . , Atlanta, USA (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>GloVe: Global Vectors for Word Representation</article-title>
          .
          <source>In: Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.: Distributional</given-names>
          </string-name>
          <string-name>
            <surname>Structure</surname>
          </string-name>
          .
          <source>Word</source>
          .
          <volume>10</volume>
          ,
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          (
          <year>1954</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lofi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Measuring Semantic Similarity and Relatedness with Distributional and Knowledge-based Approaches</article-title>
          .
          <source>Database Society of Japan (DBSJ) Journal. 14</source>
          ,
          <issue>1</issue>
          -
          <fpage>9</fpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faria</surname>
          </string-name>
          , F.F. de,
          <string-name>
            <surname>Seán O'Riain</surname>
          </string-name>
          ,
          <string-name>
            <surname>Curry</surname>
          </string-name>
          , E.:
          <article-title>Answering natural language queries over linked data graphs: a distributional semantics approach</article-title>
          .
          <source>In: Proceedings of the International Conference on Conference on Research and Development in Information Retrieval (SIGIR)</source>
          . , Dublin, Ireland (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peterson</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffiths</surname>
            ,
            <given-names>T.L.</given-names>
          </string-name>
          :
          <article-title>Evaluating vector-space models of analogy</article-title>
          .
          <source>In: Proceedings of the Conference of the Cognitive Science Society</source>
          . , London, UK (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lofi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahamed</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kulkarni</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thakkar</surname>
          </string-name>
          , R.:
          <article-title>Benchmarking Semantic Capabilities of Analogy Querying Algorithms</article-title>
          .
          <source>In: Proceedings of the International Conference on Database Systems for Advanced Applications (DASFAA)</source>
          . pp.
          <fpage>463</fpage>
          -
          <lpage>478</lpage>
          . , Dallas, TX, USA (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Dedre</surname>
            <given-names>Gentner</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Keith J.</given-names>
            <surname>Holyoak</surname>
          </string-name>
          , Boicho N. Kokinov eds:
          <article-title>The analogical mind: perspectives from cognitive science</article-title>
          . MIT Press (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Shelley</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Multiple Analogies In Science And Philosophy</article-title>
          . John Benjamins Pub. (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Gentner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Structure-mapping: A theoretical framework for analogy</article-title>
          .
          <source>Cognitive science. 7</source>
          ,
          <fpage>155</fpage>
          -
          <lpage>170</lpage>
          (
          <year>1983</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lofi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nieke</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Modeling Analogies for Human-Centered Information Systems</article-title>
          . In: Jatowt,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.-P.</given-names>
            ,
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Miura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Tezuka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Dias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Tanaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Flanagin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            , and
            <surname>Dai</surname>
          </string-name>
          , B.T. (eds.)
          <article-title>Social Informatics (SocInfo)</article-title>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Collobert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavukcuoglu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuksa</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Natural language processing (almost) from scratch</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          .
          <volume>2493</volume>
          -
          <fpage>2537</fpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient Estimation of Word Representations in Vector Space</article-title>
          .
          <source>CoRR. abs/1301</source>
          .3, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bian</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu, T.-Y.:
          <article-title>WordRep: A Benchmark for Research on Learning Word Representations</article-title>
          . In: ICML Workshop on
          <article-title>Knowledge-Powered Deep Learning for Text Mining</article-title>
          . , Beijing, China (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Linzen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Issues in evaluating semantic spaces using word analogies</article-title>
          .
          <source>In: Workshop on Evaluating Vector Space Representations for NLP.</source>
          , Berlin, Germany (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>García-Silva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : DBpedia Spotlight:
          <article-title>Shedding Light on the Web of Documents</article-title>
          .
          <source>In: Proceedings of the 7th International Conference on Semantic Systems - I-Semantics '11</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . ACM Press, New York, New York, USA (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>