<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>More is not Always Better: The Negative Impact of A-box Materialization on RDF2vec Knowledge Graph Embeddings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreea Iana</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiko Paulheim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science Group, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>RDF2vec is an embedding technique for representing knowledge graph entities in a continuous vector space. In this paper, we investigate the efect of materializing implicit A-box axioms induced by subproperties, as well as symmetric and transitive properties. While it might be a reasonable assumption that such a materialization before computing embeddings might lead to better embeddings, we conduct a set of experiments on DBpedia which demonstrate that the materialization actually has a negative efect on the performance of RDF2vec. In our analysis, we argue that despite the huge body of work devoted on completing missing information in knowledge graphs, such missing implicit information is actually a signal, not a defect, and we show examples illustrating that assumption.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;RDF2Vec</kwd>
        <kwd>Embedding</kwd>
        <kwd>Reasoning</kwd>
        <kwd>Knowledge Graph Completion</kwd>
        <kwd>A-box Materialization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>A straightforward assumption is that completing miss</title>
        <p>
          ing knowledge in a knowledge graph before computing
RDFvec [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] was originally conceived for exploiting knowl- node representations will lead to better results. However,
edge graphs in data mining. Since most popular data in this paper, we show that the opposite actually holds:
mining tools require a feature vector representation of completing the knowledge graph before computing an
records, various techniques have been proposed for cre- RDF2vec embedding actually leads to worse results in
ating vector space representations from subgraphs, in- downstream tasks.
cluding adding datatype properties as features or
creating binary features for types [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Given the increasing
popularity of the word2vec family of word embedding 2. Related Work
techniques [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], which learns feature vectors for words
based on the context in which they appear, this approach The base algorithm of RDF2vec uses random walks on
has been proposed to be transferred to graphs as well. the knowledge graph to produce sequences of nodes and
Since word2vec operates on (word) sequences, several edges. Those sequences are then fed into a word2vec
approaches have been proposed which first turn a graph embedding learner, i.e., using either the CBOW or the
into sequences by performing random walks, before ap- Skip-Gram method.
plying the idea of word2vec to those sequences. Such Since its original publication in 2016, several
improveapproaches include node2vec [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], DeepWalk [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], and the ments for RDF2vec have been proposed. The main
famaforementioned RDF2vec. ily of approaches for improving RDF2vec is to use
al
        </p>
        <p>
          There is a plethora of work addressing the completion ternatives for completely random walks to generate
seof knowledge graphs [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], i.e., the addition of missing quences. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] explores 12 variants of biased walks, i.e.,
knowledge. Since some knowledge graphs come with ex- random walks which follow non-uniform probability
dispressive schemas [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] or exploit upper ontologies [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], one tributions when choosing an edge to follow in a walk.
such approach is the exploitation of explicit ontological Heuristics explored include, e.g., preferring successors
knowledge. For example, if a property  is known to be with a high or low PageRank, preferring frequent or
insymmetric, a reverse edge (, ) can be added to the frequent edges, etc.
knowledge graph for each edge (, ) found. In [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], the authors explore the automatic
identification of a relevant subset of edge types for a given class
of entities. They show that restricting the graph for a
class of entities at hand (e.g., movies) can outperform the
results of pure RDF2vec.
        </p>
        <p>While those works exploit merely knowledge graph
internal signals (e.g., by computing PageRank over the
graph), other works include external signals as well. For
Proceedings of the CIKM 2020 Workshops, October 19-20, 2020,
Galway, Ireland
email: andreea@informatik.uni-mannheim.de (A. Iana);
heiko@informatik.uni-mannheim.de (H. Paulheim)
orcid: 0000-0002-7248-7503 (A. Iana); 0000-0003-4386-8195
(H. Paulheim)</p>
        <p>
          © 2020 Copyright for this paper by its authors. Use permitted under Creative
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)
3. Experiments
example, [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] shows that exploiting an external measure equivalent properties in Wikidata are defined as inverse
for the importance of an edge can lead to improved re- of one another6.
sults over other biasing strategies. The authors utilize
page transition probabilities obtained from server log 3.1.2. Enrichment using DL-Learner
ifles in Wikipedia to compute a probability distribution
for creating the random walks.
        </p>
        <p>
          A work that explores a similar direction to the one
proposed in this paper is presented in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. The authors
analyze the information content of statements in a
knowledge graph by computing how easily a statement can be
predicted from the other statements in the knowledge
graph. They show that translational embeddings can
benefit from being tuned towards focusing on statements
with a high information content.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>The second strategy is applying DL-Learner [16] to learn</title>
        <p>additional symmetry, transitivity, and inverse axioms for
enriching the ontology. After inspecting the results of
DL-Learner, and to avoid false T-box axioms, we used
thresholds of 0.53 for symmetric properties, and 0.45
for transitive properties. Since the list of pairs of inverse
properties generated by DL-Learner contained quite a few
false positives (e.g., dbo:isPartOf being the inverse
of dbo:countySeat as the highest scoring result), we
manually filtered the top results and kept 14 T-box axioms
which we rated as correct.</p>
        <sec id="sec-1-2-1">
          <title>3.1.3. Materializing the Enriched Graphs</title>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>To evaluate the efect of knowledge graph materializa</title>
        <p>
          tion on the quality of RDF2vec embeddings, we repeat In both cases, we identify a number of inverse,
transithe experiments on entity classification and regression, tive, and symmetric properties, as shown in Table 1. The
entity relatedness and similarity and document similar- symmetric properties identified by the two approaches
ity introduced in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], and compare the results on the highly overlap, while the inverse and transitive
propermaterialized and unmaterialized graphs.1 ties identified difer a lot.
        </p>
        <p>
          With the enriched ontology, we infer additional A-box
3.1. Experiment Setup axioms on DBpedia. We use two settings, i.e., all
subproperties plus (a) all inverse, transitive, and symmetric
For our experiments, we use the 2016-10 dump of DB- properties found using mappings to Wikidata, and (b)
pedia, which was the latest oficial release during the all plus all inverse, transitive, and symmetric properties
time at which the experiments were conducted. For cre- found with DL-Learner.
ating RDF2Vec embeddings, we use KGvec2go [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] for The inferring of additional A-box axioms was done
computing the random walks, and the fast Python reim- in iterations. In each iteration, additional A-box axioms
plementation of the original RDF2Vec code2 for training were created for symmetric, transitive, inverse, and
subthe RDF2Vec models3. properties. Using this iterative approach, chains of
prop
        </p>
        <p>Since the original DBpedia ontology provides informa- erties could also be respected. For example, from the
tion about subproperties, but does not define any sym- axioms
metric, transitive, and inverse properties, we first had to
enrich the ontology with such axioms.</p>
        <sec id="sec-1-3-1">
          <title>3.1.1. Enrichment using Wikidata</title>
        </sec>
      </sec>
      <sec id="sec-1-4">
        <title>The first strategy is utilizing owl:equivalentProperty</title>
        <p>
          links to Wikidata [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. We mark a property  in DBpedia
as symmetric if its Wikidata equivalent has a symmetric
constraint in Wikidata4, and we mark it as transitive if its
Wikidata equivalent is an instance of the Wikidata class
transitive property5. For a pair of properties  and  in
DBpedia, we mark them as inverse if their respective
        </p>
      </sec>
      <sec id="sec-1-5">
        <title>1Please note that the results on the unmaterialized graphs difer</title>
        <p>
          from those reported in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], since we use a more recent version of
DBpedia in our experiments.
        </p>
        <p>2https://github.com/IBCNServices/pyRDF2Vec
3https://github.com/andreeaiana/rdf2vec-materialization
4https://www.wikidata.org/wiki/Q21510862
5https://www.wikidata.org/wiki/Q18647515
Cerebellar_tonsil</p>
        <p>isPartOfAnatomicalStructure Cerebellum .
Cerebellum isPartOfAnatomicalStructure</p>
        <p>Hindbrain .
and the two identified T-box axioms
isPartOf a owl:TransitiveProperty .
isPartOfAnatomicalStructure</p>
        <p>rdfs:subPropertyOf isPartOf .
the first iteration adds</p>
        <p>Cerebellar_tonsil isPartOf Cerebellum .</p>
        <p>Cerebellum isPartOf Hindbrain .
whereas the second iteration adds</p>
        <p>Cerebellar_tonsil isPartOf Hindbrain .</p>
      </sec>
      <sec id="sec-1-6">
        <title>6https://www.wikidata.org/wiki/Property:P1696</title>
        <p>3. Entity relatedness and entity similarity, based on
the KORE50 dataset; and
4. Document similarity, based on the LP50 dataset,
where the similarity of two documents is
computed from the pairwise similarities of entities
identified in the texts.
3.2. Training RDF2vec Embeddings</p>
      </sec>
      <sec id="sec-1-7">
        <title>The experimental protocol in the framework used for</title>
        <p>
          On all three graphs (Original, Enriched Wikidata, and En- evaluation is defined as follows [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]:
riched DL-Learner), experiments were conducted in the For regression and classification, three (linear
regressame fashion as in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The RDF2vec approach extracts sion, k-NN, M5 rules) resp. four (Naive Bayes, C4.5
decisequences of nodes and properties by performing ran- sion tree, k-NN, Support Vector Machine) are used and
dom walks from each node. Following [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], we started 500 evaluated using 10-fold cross validation. k-NN is used
random graph walks of depth 4 and 8 from each node. with k=3; for SVM, the parameter C is varied between
        </p>
        <p>
          The resulting sequences are then used as input to 10− 3, 10− 2, 0.1, 1, 10, 102, 103, and the best value is
choword2vec. Here, two variants exist, i.e., CBOW and Skip- sen. All other algorithms are run in their respective
stanGram (SG), where SG consistently yielded better results dard configurations. 8,9
in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], so we used the SG to compute embeddings vectors For entity relatedness and similarity, the task is to rank
with a dimensionality of 200 and 500. Following [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the a list of entities w.r.t. a main entity. Here, the entities are
parameters chosen for word2vec were window size = 5, ranked by cosine similarity between the main entity’s
no. of iterations = 10, and negative sampling with no. of and the candidate entities’ RDF2vec vectors.10
samples = 25. The code and data used for the experiments For the document similarity task, the similarity of two
are available online.7 documents 1 and 2 is computed by comparing all
entities mentioned in 1 to all entities mentioned in 2 using
3.2.1. Experiments Conducted on the Enriched the metric above. For each entity in each document, the
Graphs maximum similarity to an entity in the other document is
considered, and the similarity of 1 and 2 is computed
as the average of those maxima.11
This results in 12 diferent embeddings to be compared
against each other. For evaluation, we use the evaluation
framework provided in [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The tasks to evaluate were
1. Regression: five regression datasets where an
external variable not contained in DBpedia is to be
predicted for a set of entities (cities, universities,
companies, movies, and albums);
2. Classification: five classification datasets derived
from the aforementioned regression dataset by
discretizing the target variable;
        </p>
      </sec>
      <sec id="sec-1-8">
        <title>7https://github.com/andreeaiana/rdf2vec-materialization</title>
        <p>3.3. Results on Diferent Tasks</p>
      </sec>
      <sec id="sec-1-9">
        <title>The first step of experiments are regression and classifi</title>
        <p>cation, with the results depicted in Tables 2 and 3. For the</p>
      </sec>
      <sec id="sec-1-10">
        <title>8https://github.com/mariaangelapellegrino/</title>
        <p>Evaluation-Framework/blob/master/doc/Classification.md</p>
        <p>9https://github.com/mariaangelapellegrino/
Evaluation-Framework/blob/master/doc/Regression.md</p>
        <p>10https://github.com/mariaangelapellegrino/
Evaluation-Framework/blob/master/doc/EntityRelatedness.md</p>
        <p>11https://github.com/mariaangelapellegrino/
Evaluation-Framework/blob/master/doc/DocumentSimilarity.md
regression task, we can observe that the best result for 3.4. A Closer Look at the Generated
each combination of a task and RDF2vec configuration Walks
(depth of walks, and dimensionality) is achieved on the
unmaterialized graph in 15 out of 20 cases, with linear In order to analyze the findings above, we first tried
regression or KNN delivering the best results. If we con- to correlate the findings with the actual change on the
sider all combinations of a task, an embedding, and a entities in the respective test sets. However, there is
learner, the unmaterialized graph yields better results in no clear trend which can be identified. For example,
39 out of 60 cases. in the classification and regression cases, the dataset</p>
        <p>The observations for classification are similar. For 19 which is most negatively impacted by materialization,
out of 20 combinations of a task and an RDF2vec con- i.e., the Metacritic Albums dataset, has the lowest change
ifguration, the best results are obtained on the original, in its instances’ degree (the avg. degree of the instances
unmaterialized graphs, most often with an SVM. If we changes by 0.003% and 0.007% with the Wikidata and
consider all combinations of a task, an embedding, and a the DL-Learner enrichment, respectively). On the other
learner, the unmaterialized graph yields better results in hand, the increase in the degree of the instances on the
60 out of 80 cases. cities dataset is much stronger (1.03% and 1.04%), while</p>
        <p>Moreover, if we look at how much the results degrade the decrease of the predictive models on that dataset is
for the materialized graphs, we can observe that the vari- comparatively low.
ation is much stronger for the longer walks of depth 8 We also took a closer look at the generated random
than the shorter walks of depth 4. walks on the diferent graphs. To that end, we computed</p>
        <p>The observations on the other tasks are similar. For distributions of all properties occurring in the random
entity similarity, we see that better results are achieved graph walks, for both strategies and for both depths of 4
on the unmaterialized graphs in 16 out of 20 cases, and and 8, which are depicted in Fig. 1.
in all of the four overall considerations. As far as entity From those figures, we can observe that the
distriburelatedness is concerned, the results on the unmaterial- tion of properties in the walks extracted from the
enized graphs are better in 13 out of 20 cases, as well as in riched graphs is drastically diferent from those on the
all four overall considerations. It is noteworthy that only original graphs; the Pearson correlation of the
distribuin three out of ten cases – enriching the IT companies tion in the enriched and original case is 0.44 in the case
test set with DL-Learner and Wikidata, and enriching of walks of depth 4, and only 0.21 in the case of walks
the Hollywood celebrities test set with Wikidata – the of depth 8. The property distributions among the two
degree of the entities at hand changes. This hints at the enrichment strategies, on the other hand, is very
simiefects (both positive and negative) being mainly caused lar, with the respective distributions exposing a Pearson
by information being added to the entities connected to correlation of more than 0.99.
the entities at hand (e.g., the company producing a video Another observation from the graphs is that the
distrigame), which is ultimately reflected in the walks. bution is much more uneven for the walks extracted from</p>
        <p>Finally, for document similarity, we see a diferent pic- the enriched graphs, with the most frequent properties
ture. Here, the results on the unmaterialized graphs are being present in the walks at a rate of 14-18%, whereas
always outperformed by those obtained on the material- the most frequent property has a rate of about 3% in the
ized graphs, regardless of whether the embeddings were original walks. The three most prominent properties in
computed on the shorter or longer walks. The exact rea- the enriched case – location, country, and
locationcounson for this observation is not known. One observation, try – altogether occur in about 20% of the walks in the
however, is that the entities in the LP50 dataset have by depth 4 setup, and even 30% of the walks in the depth 8
far the largest average degree (2,088, as opposed to only setup. This means that information related to locations
18 and 19 for the MetacriticMovies and MetacriticAlbums is over-represented in walks extracted from the enriched
dataset, respectively). Due to the already pretty large de- graphs. As a consequence, the embeddings tend to focus
gree, it is less likely that the materialization skews the on location-related information much more. This
obserdistributions in the random walks too much, and, instead, vation might be a possible explanation for the
degradaactually adds meaningful information. Another possible tion in results on the music and movies datasets being
reason is that the entities in LP50 are very diverse (as more drastic than, e.g., on the cities dataset.
opposed to a uniform set of cities, movies, or albums), Finally, we also looked into the correctness of the
Aand that in such a diverse dataset, the efect of materi- box axioms added. To that end, we sampled 100 axioms
alization is diferent, as it tends to add heterogeneous added with each of the two enrichment approaches, and
rather than homogeneous information to the walks. had them manually annotated as true or false by two
annotators. For the Wikidata set, the estimated precision is
65.5% (at a Cohen’s Kappa of 0.413), for the DL-Learner
dataset, the estimated precision is 61.5% (at a Cohen’s
Model / Dataset
500w_4d_200v
500w_4d_200v_Wikidata
500w_4d_200v_dllearner
500w_4d_500v
500w_4d_500v_Wikidata
500w_4d_500v_dllearner
500w_8d_200v
500w_8d_200v_Wikidata
500w_8d_200v_dllearner
500w_8d_500v
500w_8d_500v_Wikidata
500w_8d_500v_dllearner
Model / Dataset
500w_4d_200v
500w_4d_200v_Wikidata
500w_4d_200v_dllearner
500w_4d_500v
500w_4d_500v_Wikidata
500w_4d_500v_dllearner
500w_8d_200v
500w_8d_200v_Wikidata
500w_8d_200v_dllearner
500w_8d_500v
500w_8d_500v_Wikidata
500w_8d_500v_dllearner</p>
        <p>Results for the Document Similarity Task. w stands for number of walks, d stands for depth of walks, v stands for
dimension</p>
      </sec>
      <sec id="sec-1-11">
        <title>Kappa of 0.73). This shows that the majority of the axioms</title>
        <p>added to DBpedia are actually correct. Hence, we
conclude that a potential addition of erroneous axioms does
not explain the degradation in the downstream tasks.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Discussion: Missing</title>
    </sec>
    <sec id="sec-3">
      <title>Information – Signal or Defect?</title>
      <p>Since the results show that adding missing knowledge to
the knowledge graph actually results in worse RDF2vec
embeddings, we want to investigate the characteristics
of missing knowledge in DBpedia in general, as well as
its impact on RDF2vec and other algorithms.</p>
      <p>(b) depth=4, Wikidata
0.2
0.18
0.16
0.14
0.12
0.1
0.08
0.06
0.04
0.02
0 locatioloncationCountry country team rdf:type locationCity birthPlace groundregionServedrdfs:seeAlso
(e) depth=8, Wikidata
0.2
0.18
0.16
0.14
0.12
0.1
0.08
0.06
0.04
0.02
0 locatioloncationCountry country team rdf:type birthPlace locationCity groundregionServedrdfs:seeAlso
(f) depth=8, DL-Learner
4.1. Nature of Missing Information in</p>
      <p>Knowledge Graphs</p>
      <p>One first observation is that information in DBpedia and is contained in DBpedia, while its inverse
other knowledge graphs is not missing at random. For
a curated knowledge graph, a statement is contained in Nicolaus_Copernicus doctoralStudent
the knowledge graph because some person deemed it Georg_Joachim_Rheticus .
relevant.12 is not (since Nicolaus Copernicus is mainly known for</p>
      <p>Consider, e.g., the relation spouse. It is unarguably other achievements). Adding the inverse statement makes
symmetric, nevertheless, in DBpedia, only 9.8k spouse the random walks equally focus on the more important
relations are present in both directions, whereas 18.1k statements about Nicolaus Copernicus and the ones
cononly exist in one direction. Hence, the relation is noto- sidered less relevant.
riously incomplete, and a knowledge graph completion The transitive property adding most axioms to the
approach exploiting the symmetry of the spouse relation A-box is the isPartOf relation. For example, chains of
could directly add 18.1k missing axioms. geographic containment relations are usually
material</p>
      <p>One example of a spouse relation that only exists in ized, e.g., two cities in a country being part of a region, a
one direction is state, etc. ultimately also being part of that country. For
Ayda_Field spouse Robbie_Williams . once, this under-emphasizes diferences between those
cities by adding a statement making them more equal.</p>
      <p>Ayda Field is mainly known for being the wife of Robbie Moreover, there usually is a direct relation (e.g.,
counWilliams, while Robbie Williams is mostly known as a try) expressing this in a more concise way, so that the
musician. This is encoded by having the relation rep- information added is also redundant.
resented in one direction, but not the other. By adding
the reverse edge, we cancel out the information that the 4.2. Impact on RDF2vec and Other
original statement is more important than its inverse. Algorithms</p>
      <p>Adding inverse relations may have a similar efect. One
example in our dataset is the completion of doctoral advi- RDF2vec creates random walks on the graph, and uses
sors and students by exploiting the inverse relationship those to derive features. Assuming that all statements
between the two. For example, the fact in the knowledge graph are there because they were
considered relevant, each walk encodes a combination
of statements which were considered relevant.</p>
      <p>If missing information is added to the graph which
was not considered to be relevant, there is a number of
efects. First, the set of random walks encodes a mix of
12For the sake of this argument, we can also consider DBpedia
a curated knowledge graph, since the source it is created from, i.e.,
the infoboxes in Wikipedia, is curated. A statement is contained in
DBpedia if and only if somebody considers it relevant enough to be
added to an infobox in Wikipedia.
pieces of information which are relevant and pieces of
information which are not relevant. Moreover, since the
number of walks in RDF2vec is restricted by an upper
bound, adding irrelevant information also lowers the
likelihood of relevant information being reflected in a
random walk. The later representation learning will then
focus on representing relevant and irrelevant information
alike, and, ultimately, creates an embedding which works
worse.</p>
      <p>The efects are not limited to RDF2vec. Translational
embedding approaches are likely to expose a similar
behavior, since they will include both relevant and
irrelevant statements in their optimization target, which is
likely to cause a worse embedding.</p>
      <p>There are also other fields than embeddings where
missing information might actually be a valuable signal.</p>
      <p>Consider, for example, a movie recommender system
which recommends movies based on actors that played in
the movies. DBpedia and other similar knowledge graphs
typically contain the most relevant actors for a movie.13
If we were able to complete this relation and add all
actors even for minor roles, it would be likely that movie
recommendations were created on major and minor roles
alike – which are likely to be worse recommendations.</p>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion and Outlook</title>
      <sec id="sec-4-1">
        <title>In this paper, we have studied the efect of A-box ma</title>
        <p>terialization on knowledge graph embeddings created
with RDF2vec. The empirical results show that in many
cases, such a materialization has a negative efect on
downstream applications.</p>
        <p>Following up on those observations, we propose a
different view on knowledge graph incompleteness. While
mostly seen as a defect – i.e., a knowledge graph is
incomplete and hence needs to be fixed – we suggest that
such an incompleteness can also be a signal. Although
certain axioms could be completed by logical inference,
they might have been left out intentionally, since the
creators of the knowledge graph considered them less
relevant.</p>
        <p>A natural future step would be to conduct such
experiments on other embedding methods as well. While there
is a certain rationale that similar efects can be observed
on, e.g., translational embeddings as well, empirical
evidence is still outstanding.</p>
        <p>Overall, this paper has shown and discussed a
somewhat unexpected finding, i.e., that materialization an
A-box can actually do harm on downstream tasks, and
looked at various possible explanations for that
observation.</p>
        <p>13On average, a movie in DBpedia is connected to 3.7 actors.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ristoski</surname>
          </string-name>
          , H. Paulheim,
          <article-title>Rdf2vec: Rdf graph embeddings for data mining</article-title>
          ,
          <source>in: ISWC</source>
          , Springer,
          <year>2016</year>
          , pp.
          <fpage>498</fpage>
          -
          <lpage>514</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ristoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>A comparison of propositionalization strategies for creating features from linked open data</article-title>
          ,
          <source>in: LD4KD</source>
          , volume
          <volume>6</volume>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Eficient estimation of word representations in vector space (</article-title>
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          , node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          ,
          <source>in: KDD</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>855</fpage>
          -
          <lpage>864</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Perozzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Al-Rfou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Skiena</surname>
          </string-name>
          , Deepwalk:
          <article-title>Online learning of social representations</article-title>
          ,
          <source>in: KDD</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>701</fpage>
          -
          <lpage>710</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Knowledge graph refinement: A survey of approaches and evaluation methods</article-title>
          ,
          <source>Semantic Web</source>
          <volume>8</volume>
          (
          <year>2017</year>
          )
          <fpage>489</fpage>
          -
          <lpage>508</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Heist</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ringler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Knowledge graphs on the web - an overview</article-title>
          ,
          <source>in: Knowledge Graphs for eXplainable AI</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gangemi</surname>
          </string-name>
          ,
          <article-title>Serving dbpedia with dolce-more than just adding a cherry on top</article-title>
          ,
          <source>in: ISWC</source>
          , Springer,
          <year>2015</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cochez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ristoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Biased graph walks for rdf graph embeddings</article-title>
          ,
          <source>in: WIMS</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Saeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chelmis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. K.</given-names>
            <surname>Prasanna</surname>
          </string-name>
          ,
          <article-title>Extracting entity-specific substructures for rdf graph embeddings</article-title>
          ,
          <source>Semantic Web</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>1087</fpage>
          -
          <lpage>1108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Taweel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Towards exploiting implicit human feedback for improving rdf2vec embeddings</article-title>
          ,
          <source>in: DL4KGs</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Mai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Janowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Support and centrality: Learning weights for knowledge graph embedding models</article-title>
          ,
          <source>in: EKAW</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>212</fpage>
          -
          <lpage>227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ristoski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rosati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. D.</given-names>
            <surname>Leone</surname>
          </string-name>
          , H. Paulheim,
          <article-title>Rdf2vec: Rdf graph embeddings and their applications</article-title>
          ,
          <source>Semantic Web</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>721</fpage>
          -
          <lpage>752</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hladik</surname>
          </string-name>
          , H. Paulheim, Kgvec2go
          <article-title>- knowledge graph embeddings as a service</article-title>
          ,
          <source>in: LREC</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Vrandečić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krötzsch</surname>
          </string-name>
          ,
          <article-title>Wikidata: a free collaborative knowledgebase</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>57</volume>
          (
          <year>2014</year>
          )
          <fpage>78</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <article-title>Dl-learner: learning concepts in description logics</article-title>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>10</volume>
          (
          <year>2009</year>
          )
          <fpage>2639</fpage>
          -
          <lpage>2642</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Pellegrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cochez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Garofalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ristoski</surname>
          </string-name>
          ,
          <article-title>A configurable evaluation framework for node embedding techniques</article-title>
          ,
          <source>in: ESWC</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>156</fpage>
          -
          <lpage>160</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>