<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RDF2Vec-based Classi cation of Ontology Alignment Changes</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Media Unter den Eichen 5 65195 Wiesbaden</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>RheinMain University of Applied Sciences Department of Design</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>When ontologies cover overlapping topics, the overlap can be represented using ontology alignments. These alignments need to be continuously adapted to changing ontologies. Especially for large ontologies this is a costly task often consisting of manual work. Finding changes that do not lead to an adaption of the alignment can potentially make this process signi cantly easier. This work presents an approach to nding these changes based on RDF embeddings and common classi cation techniques. To examine the feasibility of this approach, an evaluation on a real-world dataset is presented. In this evaluation, the best classi ers reached a precision of 0.8.</p>
      </abstract>
      <kwd-group>
        <kwd>RDF Embedding</kwd>
        <kwd>Change Classi cation</kwd>
        <kwd>Ontology Alignment</kwd>
        <kwd>Ontology Mapping</kwd>
        <kwd>Mapping Adaption</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Finding alignments between ontologies, also known as ontology matching, is a
non-trivial task and has been an active area of research over the last ten years.
Several approaches in this area are based on the structure of the ontologies,
logical axioms or lexical similarity [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, once these alignments are found,
they will not necessarily stay untouched forever. Especially when alignments
connect large ontologies, adapting these alignments to changes is a work-intensive
task. In the area of biomedical ontologies, some alignments contain around 6500
correspondences that might be a ected by a change in one of the ontologies they
connect. Given a change in the ontology, detecting which parts of the alignment
are a ected by the change and need to be adapted is not a trivial task that usually
requires manual labour. The e ort required for this task can be signi cantly
reduced, if some changes can be excluded from it. However, it is usually not
clear how to identify changes that do not a ect the alignment.
      </p>
      <p>In this paper, we propose an approach to this problem based on RDF
embeddings and well-known classi cation techniques. The central aspect of this
approach is to represent changed concepts by their RDF embedding and classify
whether an alignment statement nearby should be changed. To gain evidence if
this approach works, we evaluate it using a dataset from the area of biomedical
ontologies. On this dataset, our approach is able to identify changes a ecting
alignment statements with a precision of 0.8.</p>
      <p>The remainder of this work is structured as follows: Section 2 discusses
foundations of our work and related approaches. The general approach is presented
in Section 3. Evaluation methodology, the dataset and results are shown in
Section 4. Section 5 discusses the results of our evaluation, and advantages and
disadvantages of our approach. A conclusion is given in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Foundations and Related Work</title>
      <p>
        An ontology alignment (sometimes also called ontology mapping) is a set of
correspondences between entities in di erent ontologies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To make it easier to
reason about these alignments, we use the following formal de nition in the style
of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for ontology mappings: An alignment between two ontologies O1 and O2
is de ned as
      </p>
      <p>AO1;O2 = f(c1; c2; semT ype)jc1 2 O1; c2 2 O2; semT ype 2 f ; ; gg
AO1;O2 is the set of all alignment statements. To denote a change of an
ontology over time, we use the prime symbol (e.g., a changed version of O is
denoted as O0). The alignment adaption problem for two ontologies O1 and O2
connected by AO1;O2 can then be stated as nding a new alignment A0O10;O20 ,
when O1 and O2 evolve to O10 and O20.</p>
      <p>
        In the area of ontology alignment adaption, several approaches are based on
rules or rule-based dependency analysis. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is focussed on nding which changes
are relevant to parts of the alignment using a dependency analysis. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] proposed
an incremental approach reacting to speci c changes in database schemas based
on rules. For each change pattern a speci c modi cation for the mapping is
de ned. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] proposed an approach that is based on a composition of alignments.
A new alignment A0O10;O20 is created by a composition of the alignment AO1;O2 and
A+O2;O20 , the alignment between O2 and O20. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] have shown that these techniques
can also be applied to ontologies. However, all of these approaches require a set
of rules that need to be constructed by a domain expert and are not necessarily
reusable for other domains. Also, these approaches are not able to identify which
changes in the ontologies are prone to causing an alignment change.
      </p>
      <p>
        The task of knowledge base completion shares some properties with the
problem we address in this work. In that area, classi ers are given a subject and a
predicate and try to predict an object [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Approaches like [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] also use
vector representations for prediction. However, this task does not take changes
in the knowledge bases into account and is not applied to ontology alignments.
      </p>
      <p>To our knowledge, no approach exists that predicts whether a given change
has an impact on the alignment without using a detailed set of rules. This issue
is at the core of our research.</p>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>
        Our general approach is based on the representation of changed resources using
RDF embeddings, a represenation of RDF nodes as vectors in a high-dimensional,
dense vector space. RDF embeddings are generated using RDF2Vec [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], an
approach based on random graph walks as input to Word2Vec [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The
RDF2VecModel is trained on an RDF graph consisting of the ontologies O1 and O2 as
well as the alignment AO1;O2 as de ned in section 2. With these embeddings, we
train a classi er on whether a changed resource a ected an alignment statement
and use this classi er to predict whether other changes will a ect the alignment.
We de ne a changed resource to lead to an alignment change, if a changed
alignment statement is within a distance of two in the RDF graph. This relatively
small measure is used to make it easier to exclude certain regions from the search
for a ected statements. For the same reason, only changes that are close to an
alignment axiom are regarded. The respective changes c are extracted using an
extension of the Protege plugin owl-di 1. By comparing the parts of AO1;O2 and
A0O10;O20 that are in the direct neighbourhood of c, it is possible to separate all
changes into two groups: (1) changes that caused an alignment change in their
neighbourhood and (2) changes that did not cause an alignment change in their
neighbourhood and therefore did not a ect the alignment.
      </p>
      <p>Each changed resource is represented by the corresponding RDF2Vec vector.
Hence, the input to the training of the classi er is a pair (v(c); k) consisting
of a vector v(c) and a class k. k determines whether c caused a change in its
direct neighbourhood. The task at hand is to correctly classify new vectors. To
solve this problem, we use several common classi cation techniques: Regression,
Naive Bayes, Tree-Based Algorithms as well as Support Vector Machines and
Multilayer Perceptrons. Each algorithm is trained on one set of changes and
evaluated on a di erent set.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>The research questions behind our evaluation are the following:
1. Can RDF embeddings be used for change classi cation with an acceptable
performance? This question tries to clarify, whether our approach is in
general applicable to the problem at hand.
2. Which classi ers can be used for this problem? This question is used to
identify the best classi ers for our problem.
4.1</p>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>The dataset used to answer these research questions in our experiments is a
real-word dataset from the domain of biomedical ontologies. It has been used</p>
        <sec id="sec-4-1-1">
          <title>1 https://github.com/mhfj/owl-diff</title>
          <p>
            in several works that deal with alignment adaption, e.g., [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ], [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]. The dataset
comprises three ontologies: SNOMED-CT, the NCI-Thesaurus and FMA. For
each ontology, yearly versions from 2009-2012 are available. Additionally, the
dataset contains alignments extracted from the UMLS metathesaurus between
the ontologies for each year. This dataset has been made publicly available2 by
the authors of [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ].
          </p>
          <p>For simplicity of our presentation, we will only present the alignment between
the ontologies NCI and FMA in the version change from 2009 to 2010. In the
formal notation introduced in Section 2, O1 and O2 refer to the ontologies NCI
and FMA as of 2009 and O10 and O20 as of 2010, respectively. The alignment from
2009 is denoted by AO1;O2 and the version from 2010 by A0O10;O20 .</p>
          <p>From 2009 to 2010, 924 changes are near alignment statements of which
47% require an adaption. These changes are used as a training set. The test set
consists of the changes from 2010 to 2011. This set contains 785 changes near
alignment statements, of which 36% lead to an alignment adaption.
4.2</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Methodology</title>
        <p>
          To generate RDF embeddings, the code from RDF2Vec [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] was used. The
embeddings were trained using the skip gram model, with 500 dimensions used
for the embeddings and random walks of length 8, as this was identi ed as the
best-performing variant in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. An overview regarding classi cation methods used
on these embeddings is given in Table 1. Standard scikit3 implementations are
used for the classi cation process. The classi ers are trained on changes from
2009-2010 and validated on changes from 2010-2011 of the dataset described in
Section 4.1. The performance of di erent classi cation techniques is evaluated
based on f1-measure, accuracy, precision and recall.
        </p>
        <sec id="sec-4-2-1">
          <title>2 https://dbs.uni-leipzig.de/de/research/projects/evolution_of_</title>
          <p>ontologies_and_mappings/ontology_mapping_adaption
3 http://scikit-learn.org/stable/</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Results</title>
        <p>The results of the described process are displayed in Table 2. Only changes
close to the alignment were included in this evaluation, since it would otherwise
be very easy to achieve accuracy values above 95%. Results for MLP did not
vary based on the structure of the hidden layers, so one row represents all MLP
results. All algorithms show a very similar performance regarding the evaluated
metrics. The highest achieved precision is 0.81, which can be reached using MLP
and linear SVM classi cation. These methods also reach the highest f1-measures
of 0.75. Accuracy of all classi ers is only marginally higher than what can be
achieved using random guessing, given the distribution of classes in the test set.
The results presented in section 4.3 give us some evidence on our rst research
question: Using RDF embeddings to represent changes seems to be a promising
approach to the mapping adaption problem, as we can see a precision around
0.8. In general, several classi cation approaches show a similar performance. This
precision can be achieved, although the approach uses no information regarding
the nature of changes, e.g., the algorithm can not distinguish the correction of
typos from major, structural changes.</p>
        <p>
          An important advantage of this approach is that no sophisticated change
model that is adapted to the domain is required. Approaches like [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] require a
rule-base that needs to be constructed from a detailed understanding of typical
changes in the domain the ontologies describe. Hence, the author of these rules
needs to be an expert in ontology engineering as well as the application domain.
Also, these rules need to be constantly adapted to evolving domains, whereas
an RDF-embedding based approach could learn new patterns autonomously.
However, to demonstrate these advantages, it is still required to show that this
approach is also applicable to other data sets and di erent application domains.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Outlook</title>
      <p>In this work, we presented an approach to ontology alignment adaption based
on RDF embeddings and common classi cation techniques. An evaluation on a
dataset from the biomedical domain provided some evidence, that the approach
is feasible. On the dataset, best-performing classi ers had a precision of 0.8.</p>
      <p>As future work, several extensions are possible: Further evaluations could
be performed on di erent datasets. Also, a combination of this approach with
existing mapping adaption approaches could be examined. Change types could
be used as another input to the classi cation process to improve classi cation
accuracy. Another aspect for future work is to determine, when embeddings need
to be updated, since embeddingis will become outdated when ontologies change.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Jero</surname>
          </string-name>
          <article-title>^me Euzenat and Pavel Shvaiko</article-title>
          . Ontology matching. Springer, Heidelberg,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Anika</given-names>
            <surname>Gro</surname>
          </string-name>
          , Julio Cesar dos Reis, Michael Hartung, Cedric Pruski, and
          <string-name>
            <given-names>Erhard</given-names>
            <surname>Rahm</surname>
          </string-name>
          .
          <article-title>Semi-automatic adaptation of mappings between life science ontologies</article-title>
          .
          <source>Lecture Notes in Computer Science (including subseries Lecture Notes in Arti cial Intelligence and Lecture Notes in Bioinformatics)</source>
          , 7970 LNBI:
          <volume>90</volume>
          {
          <fpage>104</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ernesto</surname>
            Jimenez-ruiz, Bernardo Cuenca Grau, Ian Horrocks, and
            <given-names>Rafael</given-names>
          </string-name>
          <string-name>
            <surname>Berlanga</surname>
          </string-name>
          .
          <article-title>Logic-based assessment of the compatibility of UMLS ontology sources</article-title>
          .
          <source>In JOURNAL OF BIOMEDICAL SEMANTICS</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          , Edouard Grave, Piotr Bojanowski, Maximilian Nickel, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Fast Linear Model for Knowledge Graph Embeddings</article-title>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Michel</given-names>
            <surname>Klein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Heiner</given-names>
            <surname>Stuckenschmidt</surname>
          </string-name>
          .
          <article-title>Evolution Management for Interconnected Ontologies</article-title>
          .
          <source>Workshop on Semantic Integration at ISWC</source>
          <year>2003</year>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <article-title>Je rey Dean. E cient estimation of word representations in vector space</article-title>
          .
          <source>CoRR, abs/1301.3781</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Maximilian</given-names>
            <surname>Nickel</surname>
          </string-name>
          , Kevin Murphy, Volker Tresp, and
          <string-name>
            <given-names>Evgeniy</given-names>
            <surname>Gabrilovich</surname>
          </string-name>
          .
          <article-title>A review of relational machine learning for knowledge graphs</article-title>
          .
          <source>Proceedings of the IEEE</source>
          ,
          <volume>104</volume>
          :
          <fpage>11</fpage>
          {
          <fpage>33</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Petar</given-names>
            <surname>Ristoski</surname>
          </string-name>
          and
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Paulheim</surname>
          </string-name>
          .
          <article-title>RDF2Vec: RDF graph embeddings for data mining</article-title>
          .
          <source>In The Semantic Web - ISWC 20162016</source>
          , pages
          <fpage>498</fpage>
          {
          <fpage>514</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Richard Socher, Danqi Chen, Christopher Manning, Danqi Chen, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Ng</surname>
          </string-name>
          .
          <article-title>Reasoning With Neural Tensor Networks for Knowledge Base Completion</article-title>
          .
          <source>Neural Information Processing Systems</source>
          (
          <year>2003</year>
          ), pages
          <fpage>926</fpage>
          {
          <fpage>934</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Yannis</surname>
            <given-names>Velegrakis</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Renee J</given-names>
            .
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Lucian</given-names>
            <surname>Popa</surname>
          </string-name>
          .
          <article-title>Mapping adaptation under evolving schemas</article-title>
          .
          <source>VLDB '03 Proceedings of the 29th international conference on Very large data bases -</source>
          Volume
          <volume>29</volume>
          , pages
          <fpage>584</fpage>
          {
          <fpage>595</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Bishan</surname>
            <given-names>Yang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wen-tau Yih</surname>
          </string-name>
          , Xiaodong He,
          <string-name>
            <surname>Jianfeng Gao</surname>
            ,
            <given-names>and Li</given-names>
          </string-name>
          <string-name>
            <surname>Deng</surname>
          </string-name>
          .
          <article-title>Embedding Entities and Relations for Learning and Inference in Knowledge Bases</article-title>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>Cong</given-names>
            <surname>Yu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lucian</given-names>
            <surname>Popa</surname>
          </string-name>
          .
          <article-title>Semantic Adaptation of Schema Mappings when Schemas Evolve</article-title>
          .
          <source>Very Large Data Bases</source>
          , pages
          <volume>1006</volume>
          {
          <fpage>1017</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>