<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DOME Results for OAEI 2019</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Confidence Adjustment</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science Group, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>DOME (Deep Ontology MatchEr) is a scalable matcher for instance and schema matching which relies on large texts describing the ontological concepts. The doc2vec approach is used to generate a vector representation of the concepts based on the textual information contained in literals. The cosine distance between two concepts in the embedding space is used as a con dence value. In comparison to the previous version of DOME it uses an instance based class matching approach. Due to its high scalability, it can also produce results in the largebio track of OAEI and can be applied to very large knowledge graphs. The results look promising if huge texts are available, but there is still a lot of room for improvement.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology Matching</kwd>
        <kwd>Knowledge Graph</kwd>
        <kwd>Doc2Vec</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>State, purpose, general statement</title>
      <p>The overall matching strategy of DOME is shown in gure 1. It starts with a
simple string matching followed by a con dence adjustment. This is applied for
0 Copyright c 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>Ontology 1
Ontology 2</p>
      <sec id="sec-2-1">
        <title>String matching</title>
        <p>final
alignment</p>
      </sec>
      <sec id="sec-2-2">
        <title>Instance based class matching</title>
      </sec>
      <sec id="sec-2-3">
        <title>Cardinality Filter</title>
      </sec>
      <sec id="sec-2-4">
        <title>Type Filter</title>
        <p>all classes, instances, and properties. The latter one includes owl:ObjectProperty,
owl:DatatypeProperty, and rdf:Property (as retrived by the jena1 method
OntModel.listAllOntProperties()). As a next step in the pipeline, an instance based
class matching is applied. It uses all matched individuals and based on those
types, tries to nd meaningful class mappings.</p>
        <p>The following type lter deletes all correspondences where the type of source
and target concept is di erent (like owl:DatatypeProperty - owl:ObjectProperty ).
This might happen because all properties (also rdf:Property ) can be matched
with each other. The nal cardinality lter ensures a one to one mapping by
sorting the correspondences by con dence and iterates over them in descending
order. If the source or target entity is not already matched, it counts a valid
correspondence - otherwise it will be dropped and will not appear in the nal
alignment.</p>
        <p>In the following, the rst three matching stages of DOME are discussed in
more detail.</p>
        <p>String matching As shown in gure 2, DOME uses multiple properties for
matching all types of resources. If a rdfs:label from ontology A matches the rdfs:label
from a resources in ontology B after the preprocessing, DOME creates a
mapping with a static con dence of 1.0. The same con dence is applied when a
skos:prefLabel matches. In case a URI fragment or skos:altLabel ts, a lower
con dence of 0.9 is used.</p>
        <p>The string preprocessing consists of tokenizing the text (also takes care of
CamelCase2 formatting), stopword removal and lowercasing. Afterwards the text
is concatenated together to form a new textual representation. In case the initial
text contains mostly numbers, the whole text is discarded.</p>
        <p>
          Con dence Adjustment The con dence adjustment stage of DOME iterates over
all correspondences and reassign a new con dence in case it is possible. The main
approach used here is doc2vec [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] which is based on word2vec [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. It allows to
compare texts of di erent lengths and represent them as a xed length vector.
A comparison of these vectors can be achieved with a cosine similarity.
        </p>
        <sec id="sec-2-4-1">
          <title>1 https://jena.apache.org 2 https://en.wikipedia.org/wiki/Camel_case</title>
          <p>Resource1</p>
          <p>URI Fragment</p>
          <p>RDFS:label
SKOS:prefLabel
SKOS:altLabel
String Literal</p>
          <p>String similarity (0.9)
String similarity (1.0)
String similarity (1.0)
String similarity (0.9)
Doc2Vec similarity</p>
          <p>RDFS:label
SKOS:prefLabel
SKOS:altLabel</p>
          <p>String Literal
Resource2</p>
          <p>
            In comparison to DOME submitted to OAEI 2018, the generation of the
text for a given resource has changed. In the current version, all statements
in the ontology are examined where a given resource has the subject position.
If the object is a literal and the datatype of it corresponds to xsd:string or
rdf:langString or contains a language tag, it will be selected. All those literals
are preprocessed in the same way as described in paragraph string matching and
concatenated together. This text forms a document which is used for training a
doc2vec model. DOME uses the DM sequence learning algorithm with a vector
size of 300 and window size of 5 as in the previous version of this matcher dla[
            <xref ref-type="bibr" rid="ref3">3</xref>
            ].
The minimal word frequency is set to one to allow all words contribute to the
concept vector. The adjusted con dence is later used in the cardinality lter to
create a 1:1 mapping.
          </p>
          <p>Instance based class matching After the class, instance, and property matching
an additional class alignment step is performed. The basic idea is to inspect the
types (classes) of already matched instances. If two individuals are the same,
there is a high probability that some of the corresponding types should be also
matched.</p>
          <p>
            We experimented with three di erent similarity metrics for two given classes
c1 and c2. The dice similarity metric [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] is de ned as follows:
          </p>
          <p>SimDICE (c1; c2) =
jIc1 j + jIc2 j
2 jIc1 \ Ic2 j 2 [0:::1]</p>
          <p>Ic1 and Ic2 denotes the set of instances which have c1 (c2) as one of its type.
Ic1 \ Ic2 corresponds to the matched instances which are typed with both c1 and
c2. SimDICE corresponds to the overlap of matched instances with both classes
and all instances of the two classes separately.</p>
          <p>
            [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] also includes a more relaxed version of the previous similarity called
SimMIN which is de ned as
          </p>
          <p>SimMIN (c1; c2) =</p>
          <p>jIc1 \ Ic2 j
min(jIc1 j; jIc2 j) 2 [0:::1]</p>
          <p>
            It interrelates the matched instances with both classes and the instances of
the smaller-sized class. As stated in [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] SimDICE is always smaller or equal to
SimMIN .
          </p>
          <p>
            A third possibility is SimBASE [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] which matches the classes c1 and c2 in
case at least one instance with those classes is matched:
          </p>
          <p>SimBASE(c1; c2) =
(1
0
if jIc1 \ Ic2 j &gt; 0
if jIc1 \ Ic2 j = 0 2 [0:::1]</p>
          <p>After experimenting with those measures, it turned out that SimBASE
introduces a lot of wrong correspondences because each error in the instance matching
is directly forwarded to the class matches. SimMIN needed a very low threshold
and ranks the classes suboptimal. Thus some similarity between SimBASE and
SimMIN is needed. One possible way is to incorporate the quality of the matcher
at hand - especially how many instance correspondences it nds. Thus another
similarity called SimMAT CH is used in DOME and de ned as follows:
SimMAT CH (c1; c2) = jIc1 \ Ic2 j 2 [0:::1]
jCI j
where CI represents all instance correspondences created by the matcher. The
threshold is set to 0:01 meaning that 1 % of the matches should have the
same class. If this is the case, the classes will be matched with a con dence
of SimMAT CH (c1; c2). This value is rather low. All correspondences generated
by this step are therefore scaled to minimum of 0:1 and maximum of 1:0.
1.2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Speci c techniques used</title>
      <p>
        The two main techniques used in DOME are the doc2vec approach [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for
comparing the textual representation of the resources and the instance based class
matching component.
1.3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Adaptations made for the evaluation</title>
      <p>
        As in the previous version of DOME for OAEI 2018 the DL4J3 (Deep Learning
for Java) library is used as an implementation of the doc2vec approach. Running
DOME with this dependency is not easy in SEALS. Therefore we use MELT[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to
package our matcher. The framework generates an intermediate matcher which
executes an external process (which is again in Java). This process runs now
in its own Java virtual machine (JVM) and allows to load system dependent
library les ( les with dll or so extension). [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] explains in more detail why this
is necessary.
1.4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Link to the system and parameters le</title>
      <p>DOME can be downloaded from
https://www.dropbox.com/s/1bpektuvcsbk5ph/DOME.zip?dl=0.</p>
      <sec id="sec-5-1">
        <title>3 https://deeplearning4j.org</title>
        <p>Results
This section discusses the results of DOME for each track of OAEI 2019 where
the matcher is able to produce results. The following tracks are included: anatomy,
conference, largebio, phenotype, and knowledge graph track.</p>
        <p>Similar to the previous version of DOME, the current matcher is not able to
match multiple languages and thus fail on multifarm track. Speci c interfaces
and matching strategies for the complex and interactive track are currently not
implemented.
2.1</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Anatomy</title>
      <p>For the anatomy track, DOME uses the string comparison method which
results in similar precision and recall as the baseline. Properties like oboInOwl:
hasRelatedSynonym or oboInOwl:hasDe nition are used to generate a textual
representation of the concepts but this does not introduce better con dence
values.</p>
      <p>DOME returns 948 correspondences. 932 matches with a con dence of 1.0
which are all correct. 12 correspondences scored with 0.9 are all false positives.
Therefore a con dence lter would make sense for this speci c track.</p>
      <p>The presented matcher has a very low runtime and scales to very huge
ontologies. The runtime of 23 seconds is the second best value in this track.</p>
      <p>Due to a slightly lower recall (0.007) and precision (0.001) DOME has a lower
F-Measure (0.006) than the baseline. The reason could be the di erent string
preprocessing techniques.
2.2</p>
    </sec>
    <sec id="sec-7">
      <title>Conference</title>
      <p>In the following analysis we refer to the rar24 reference alignment because it
contains more correspondences which are carefully resolved by an evaluator.</p>
      <p>When matching classes DOME is same as the edna baseline. Most
correspondences have a con dence of 0.9 because the conference track has mostly
all textual information in URL fragments. Only one mapping is scored with 1.0
which is &lt;edas:Country, iasted:Conference state, =, 1.0&gt;. It is generated by the
instance based class matching because both contain Mexico as an individual.
This mapping is a false positive. The instance based class matching could not
help here, because in most of the test cases no instances are available. Properties
are matched with an F1-measure of 0.22 which is better than the edna baseline
but lower than 5 other matchers. In comparison to the old version of this matcher,
the F1-measure is increased by 0.01. Figure 3 shows the result of DOME divided
into test cases. It shows that in four test cases (where the source ontology is
confOf ) the matcher is not able to return true positive correspondences.
4 http://oaei.ontologymatching.org/2019/results/conference/index.html
As the name already suggests, the largebio track needs matchers which scale
well. Test case four is a large test case which matches the whole FMA ontology
with a large fragment of SNOMED. The source ontology has 78,989 classes and
the target ontology 122,464 classes. This would result in more than 9 billion
comparisons when doing it naively. The runtime of DOME for this test case is
38 seconds which is the second best runtime. Moreover DOME is able to complete
all tasks within the given timeout.</p>
      <p>In task 3, 4, 5, and 6 DOME has the highest precision of all matchers but
misses a lot of correspondences in the gold standard and has therefore a lower
recall. In task one and two matcher Wiktionary have a higher precision.
Fmeasure wise DOME usually beats Wiktionary and AGM but AML and LogMap
variants are better.
2.4</p>
    </sec>
    <sec id="sec-8">
      <title>Phenotype</title>
      <p>In phenotype track, the matcher should nd alignments between disease and
phenotype ontologies. The matcher has the highest precision of 0.997 together
with FCAMapKG for test case HP-MP and second best for task DOID-ORDO.
With the low recall of 0.303 and 0.426 the F-measure is around 0.465 and 0.596.
In the second version of the knowledge graph track, the systems should be able
to match classes, properties and instances. DOME was able to run 4 out of 5 test
cases. The remaining test case could not be nished because of memory issues.</p>
      <p>In comparison to the previous version of the track, classes are more di cult
to match. DOME could achive an F-measure of 0.77 for classes (not counting the
un nished test case) and 0.96 for properties. Only FCAMap-KG and Wiktionary
are better in matching the latter one. Instances are matched with a F-measure
of 0.88 (again not counting the un nished test case). In average DOME returns
22 class, 75 property, and 4,895 instance mappings.
3
3.1</p>
      <p>General comments</p>
    </sec>
    <sec id="sec-9">
      <title>Comments on the results</title>
      <p>The discussion of the results shows that DOME is in a development phase. Some
improvements are already incorporated and some further ideas are discussed in
the next section.
3.2</p>
    </sec>
    <sec id="sec-10">
      <title>Discussions on the way to improve the proposed system</title>
      <p>
        One further improvment is still the ability to match di erent languages. As
stated in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] we could use cross lingual embeddings as shown in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Another
possibility would be to use a translation step in between.
      </p>
      <p>The con dence adjustment step can not only be done with doc2vec based
models but also with tf-idf or other document comparison methods. This should
be tried out in future version of this matcher.</p>
      <p>The memory issue in the knowledge graph track can be solved by writing all
text representations of all resources on disk and train the doc2vec model on this
le.
4</p>
      <p>Conclusions
In this paper, we have analyzed the results of DOME in OAEI 2019. It shows
that DOME is a highly scalable matcher which generates class, property and
instance alignments. With the new component DOME is able to match classes
based on instances and thus increase the recall of class alignments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>In: The semantic web</source>
          , pp.
          <volume>722</volume>
          {
          <fpage>735</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Idehen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berners-Lee</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linked data on the web (ldow2008)</article-title>
          .
          <source>In: Proceedings of the 17th international conference on World Wide Web</source>
          . pp.
          <volume>1265</volume>
          {
          <fpage>1266</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Dome results for oaei 2018</article-title>
          . In: OM@ ISWC. pp.
          <volume>144</volume>
          {
          <issue>151</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hertling</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Portisch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paulheim</surname>
          </string-name>
          , H.:
          <article-title>Melt - matching evaluation toolkit</article-title>
          .
          <source>In: SEMANTICS. Karlsruhe</source>
          . (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ho</surname>
            <given-names>art</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago2: A spatially and temporally enhanced knowledge base from wikipedia</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>194</volume>
          ,
          <fpage>28</fpage>
          {
          <fpage>61</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kirsten</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thor</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahm</surname>
          </string-name>
          , E.:
          <article-title>Instance-based matching of large life science ontologies</article-title>
          .
          <source>In: International Conference on Data Integration in the Life Sciences</source>
          . pp.
          <volume>172</volume>
          {
          <fpage>187</fpage>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of sentences and documents</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          . pp.
          <volume>1188</volume>
          {
          <issue>1196</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. van Rijsbergen,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Information retrieval (</article-title>
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Ruder</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vulic</surname>
          </string-name>
          , I., S gaard, A.:
          <article-title>A survey of cross-lingual word embedding models</article-title>
          .
          <source>arXiv preprint arXiv:1706.04902</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A survey of schema-based matching approaches</article-title>
          . In: Spaccapietra,
          <string-name>
            <surname>S</surname>
          </string-name>
          . (ed.)
          <source>Journal on Data Semantics IV, Lecture Notes in Computer Science</source>
          , vol.
          <volume>3730</volume>
          , pp.
          <volume>146</volume>
          {
          <fpage>171</fpage>
          . Springer Berlin Heidelberg (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>