<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sven Hertling</string-name>
          <email>sven.hertling@uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiko Paulheim</string-name>
          <email>heiko.paulheim@uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Ontology Matching, Knowledge Graph</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>1. The Tbox and Abox have</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Data and Web Science Group, University of Mannheim</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>In this paper we present the results of the ATBox matcher (ATMatcher for short) during the OAEI campain 2022. The system is able to match instances (Abox) as well as the schema (Tbox) of given knowledge graphs. It scales to large inputs by using efective and modular approaches implemented in the MELT framework. The system participates for the third time in the OAEI. First, two pipelines for matching the schema and instance are used to generate candidates. Afterwards the instance matches are improved and repaired by reusing the candidate class correspondences.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>TBox1
TBox 2
ABox1
ABox 2</p>
      <p>Stopword
Extraction
final
alignment
String Matching</p>
      <p>Synonym</p>
      <p>Extension
Cardinality Filter
Similar Neighbors</p>
      <p>Filter
Cosine Similarity</p>
      <p>Filter</p>
      <p>String Matching
Bounded Path</p>
      <p>Matching
Instance Filter</p>
      <p>Type Filter</p>
      <p>Common
Properties Filter
alignment. One of the main diferences in comparison to the system submitted last year is the
additional bounded path matching for classes.</p>
      <p>First have a look at the Tbox matching. It is applied for all classes and properties (o w l : O b j e c t
P r o p e r t y , o w l : D a t a t y p e P r o p e r t y , and r d f : P r o p e r t y ). They are retrieved by the jena1 methods
OntModel.listClasses() and OntModel.listAllOntProperties().</p>
      <p>The first step is to extract KG specific stopwords because in some cases the labels and/or
fragments contains tokens which appears very often like c l a s s , i n f o b o x etc. If these tokens
appears in more than 20 % of all classes/properties, then they are assumed to be stop words.</p>
      <p>
        The synonyms are extracted from the English Wiktionary via DBnary [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The extraction
process is detailed in the previous results paper[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] similarly to the string matching component.
After these components the new bound path matching is executed. This component will match
classes which are in between two already matched classes in a hierarchy. Thus it is a structural
approach which requires already matched resources. Figure 2 shows an example. The class
book is matched to class books and novel to novel. With this information, the class in between
is a candidate for another correspondence. Thus it will be added with the average confidence of
the other two correspondences.
      </p>
      <p>The instance matching (Abox - shown in the lower part of the figure 1) is kept the same in
comparison to the last submission. As a last step, all correspondences are combined and a final
cardinality filter ensures a one to one alignment by comparing the confidence scores.</p>
    </sec>
    <sec id="sec-2">
      <title>1.2. Specific techniques used</title>
      <p>
        We used the following matching components of MELT [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]:
one:Book
two:Books
rdfs:subClassOf
one:novel
crime
one:Fiction
      </p>
      <p>Book
rdfs:subClassOf
rdfs:subClassOf
one:novel
two:Enterta</p>
      <p>inment
two:Novel
rdfs:subClassOf
rdfs:subClassOf
• ScalableStringProcessingMatcher
• StopwordExtraction
• SimilarNeighborsFilter
• CommonPropertiesFilter
• CosineSimilarityConfidenceMatcher
• SimilarTypeFilter
• NaiveDescendingExtractor
• BoundedPathMatching</p>
    </sec>
    <sec id="sec-3">
      <title>1.3. Adaptations made for the evaluation</title>
      <p>ATBox matcher is available as a docker image. When starting a container given the image,
an HTTP endpoint is started on port 8080 which fulfills the requirements of the web based
interface described in the MELT user guide 2.</p>
    </sec>
    <sec id="sec-4">
      <title>1.4. Link to the system and parameters file</title>
      <sec id="sec-4-1">
        <title>ATBox matcher can be downloaded from https://www.dropbox.com/s/l344aawh0mw6rjm/atmatcher-1.0-web-latest.tar.gz?dl=0.</title>
        <sec id="sec-4-1-1">
          <title>2. Results</title>
          <p>This section discusses the results of ATBox for each track of OAEI 2022 where the matcher
is able to produce results. The following tracks are included: anatomy, conference, bio-ml,
commonKG and knowledge graph track.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>2https://dwslab.github.io/melt/matcher-packaging/web</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>2.1. Anatomy</title>
      <p>The F-Measure didn’t change in comparison to last years submission which was expected. It is
still 0.794. This beats the baseline but only by a small margin. The matcher is very precision
oriented and achieves the third highest value after the string baseline and ALIN. The recall can
be optimized by not only using synonyms from wordnet but also other external sources. We
hope to create a coherent alignment in the next submission by using the aforementioned repair
strategies.</p>
    </sec>
    <sec id="sec-6">
      <title>2.2. Conference</title>
      <p>
        In the conference track, ATBox matcher has a F-Measure of 0.59 using the rar2-M3 evaluation
setup [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] (which is a violation free version of the entailed reference alignment for classes and
properties). This is the fourth highest value after LogMap, GraphMatcher, and SEBMatcher.
Again the recall (with 0.51) is lower than precision (with 0.69).
      </p>
      <p>This year a new subtrack was also evaluated. The task is to match DBpedia to OntoFarm
ontologies. ATBox was one of six systems which was able to solve this task. The F-Measure is
0.55 (same with KGMatcher+ and LSMatch) and only LogMap was better in those test cases.</p>
    </sec>
    <sec id="sec-7">
      <title>2.3. Common Knowledge Graphs</title>
      <p>In this track the task is to align classes in two given KGs. This allows to use instance matches
to find useful class correspondences. In 2021 there was only one task where the input graphs
are Nell and DBpedia. This year a new task was added which aligns YAGO and Wikidata.</p>
      <p>In the first task, ATBox scored 0.89 (only Matcha, KGMatcher+ is better). In the second task
the situation is exactly the same - meaning that the proposed system is on the third place. This
also shows that the two matching tasks are very similar to each other.</p>
      <p>For this track it would be beneficial if classes matches are created with the help of instances
correspondences as already done by DOME matcher. The current version only uses the schema
matches to improve the instance alignment.</p>
    </sec>
    <sec id="sec-8">
      <title>2.4. Knowledge Graph</title>
      <p>In the KG track ATMatcher is the best matching system with an overall F-Measure of 0.84. In
previous years, some systems were able to beat this score by 0.03 like Wiktionary matcher.</p>
      <p>Due to the fact that scalability is one crucial factor for developing the system, it shows clearly
that it is the fastest one after the baselines. The proposed system only needs 19 minutes for all
test cases.</p>
      <sec id="sec-8-1">
        <title>3. General comments</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>3.1. Discussions on the way to improve the proposed system</title>
      <p>Due to the fact that the two matching pipelines are independent of each other, the runtime of
the system could be further decreased by parallelization.</p>
      <p>
        Another possible way to improve the system is to incorporate a transformer model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] as
already shown in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Furthermore, the created alignments can be logically checked at the
end of the pipeline. Possible approaches are the LogMap reapir strategy [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or the ALCOMO
component[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>Finally, one could not only improve instance matches by schema macthes but also the other
way around. This was not implemented in the presented approach but would help especially in
the Common KG track.</p>
      <sec id="sec-9-1">
        <title>4. Conclusions</title>
        <p>In this paper, we have analyzed the results of ATBox matcher in OAEI 2022. The system is
scalable and can generate class, property and instance alignments.</p>
        <p>
          Most of the components which are used in ATBox are included in the MELT framework[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
which allows other researchers to reuse and compose components in their own systems.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Atbox results for oaei 2021</article-title>
          ,
          <source>OM @ ISWC</source>
          <volume>3063</volume>
          (
          <year>2021</year>
          )
          <fpage>137</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Melt - matching evaluation toolkit</article-title>
          , in: SEMANTICS. Karlsruhe.,
          <year>2019</year>
          , pp.
          <fpage>231</fpage>
          -
          <lpage>245</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Solimando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Jimenez-Ruiz</surname>
          </string-name>
          , G. Guerrini,
          <article-title>Minimizing conservativity violations in ontology alignments: Algorithms and evaluation</article-title>
          ,
          <source>Knowledge and Information Systems</source>
          <volume>51</volume>
          (
          <year>2017</year>
          )
          <fpage>775</fpage>
          -
          <lpage>819</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jiménez-Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Meilicke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Horrocks</surname>
          </string-name>
          ,
          <article-title>Evaluating mapping repair systems with large biomedical ontologies</article-title>
          .,
          <source>Description Logics</source>
          <volume>13</volume>
          (
          <year>2013</year>
          )
          <fpage>246</fpage>
          -
          <lpage>257</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Sérasset</surname>
          </string-name>
          , Dbnary:
          <article-title>Wiktionary as a lemon-based multilingual lexical resource in rdf</article-title>
          ,
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <year>2015</year>
          )
          <fpage>355</fpage>
          -
          <lpage>361</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Atbox results for oaei 2020</article-title>
          ,
          <source>OM@ ISWC</source>
          <volume>2788</volume>
          (
          <year>2020</year>
          )
          <fpage>168</fpage>
          -
          <lpage>175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>O.</given-names>
            <surname>Zamazal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Svátek</surname>
          </string-name>
          ,
          <article-title>The ten-year ontofarm and its fertilization within the onto-sphere</article-title>
          ,
          <source>Journal of Web Semantics</source>
          <volume>43</volume>
          (
          <year>2017</year>
          )
          <fpage>46</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hertling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Portisch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          ,
          <article-title>Matching with transformers in melt</article-title>
          , in: OM@ ISWC,
          <year>2021</year>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Meilicke</surname>
          </string-name>
          ,
          <article-title>Alignment incoherence in ontology matching (</article-title>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>