<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>We Divide, You Conquer:</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ernesto Jime´nez-Ruiz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Asan Agibetov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Samwald</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerie Cross</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics, University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Miami University</institution>
          ,
          <addr-line>Oxford, OH 45056</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Section for Artificial Intelligence and Decision Support, Medical University of Vienna</institution>
          ,
          <addr-line>Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Alan Turing Institute</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p>An ontology matching task MT is composed of a pair of ontologies O1 and O2 and possibly an associated reference alignment MRA. The objective of a matching task is to discover an overlapping of O1 and O2 in the form of an alignment M. The size or search space of a matching task is typically bound to the size of the Cartesian product between the entities of the input ontologies. Large-scale ontology matching tasks still pose serious challenges to state-of-the-art ontology alignment systems [2]. In this paper we propose a novel method to effectively divide an input ontology matching task MT into several (independent) and more manageable (sub)tasks. This method relies on an efficient lexical index (as in LogMap [3]), a neural embedding model [4] and locality modules [5]. Unlike other state-of-the-art approaches, our method provides guarantees about the preservation of the coverage of the relevant ontology alignment.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1 Introduction
The approach presented in this paper relies on an ‘inverted’ lexical index (we will refer
to this index as LexI), commonly used in information retrieval applications, and also
used in ontology alignment systems like LogMap [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. LexI encodes the labels of all
entities of the input ontologies O1 and O2, including their lexical variations (e.g.,
preferred labels, synonyms), in the form of pairs key-value where the key is a set of words
and the value is a set of entity identifiers1 such that the set of words of the key appears in
one of the entity labels. Table 1 shows a few example entries of LexI for two ontologies.
2.1
      </p>
      <p>
        Creating matching subtasks from LexI
Deriving mappings from LexI. Each entry in LexI, after discarding entries pointing to
only one ontology, is a source of candidate mappings. For instance the example in
Table 1 suggests that there is a (potential) mapping between the entities O1:Serous acinus
and O2:Liver acinus since they are associated to the same entry in LexI facinusg. The
mappings derived from LexI are not necessarily correct but will link lexically related
? An extended version of this paper is available in arXiv.org [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>1 The indexation module associates unique numerical identifiers to entity URIs.</title>
      <p>entities. We refer to the set of all mappings suggested by LexI as MLexI. Note that</p>
      <sec id="sec-2-1">
        <title>MLexI represents a manageable subset of the Cartesian product between the entities of</title>
        <p>
          the input ontologies. Mappings outside MLexI will rarely be discovered by standard
matching systems as they typically rely on lexical similarity measures [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
Context as matching task. Logic-based module extraction techniques compute
ontology fragments that capture the meaning of an input signature (e.g., set of entities)
with respect to a given ontology. In this paper we rely on bottom-locality modules [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
which will be referred to as locality-modules or simply as modules. Locality
modules play an important role in ontology alignment tasks. For example they provide
the scope or context (i.e., sets of semantically related entities [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]) for the entities in
a given mapping or set of mappings. The context of the mappings MLexI derived from
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>LexI leads to two ontology modules O1LexI and O2LexI from O1 and O2, respectively. MT LexI = hO1LexI; O2LexIi is the (single) matching subtask derived from LexI.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The whole set of entries in LexI, however, may lead to a very large number of can</title>
      <p>
        didate mappings MLexI and, as a consequence, to large modules O1LexI and O2LexI. These
modules, although smaller than O1 and O2, can still be challenging for many ontology
matching systems. A solution is to divide the entries in LexI into more than one cluster.
Figure 1 shows an overview of the pipeline where LexI is split into n clusters and these
clusters lead to n matching subtasks DMnT = fMT1LexI; : : : ; MTnLexIg.
Clustering strategies. We have implemented two clustering strategies which we refer to
as: naive and neural embedding. The naive strategy implements a very simple algorithm
that randomly splits the entries in LexI into a given number of clusters of the same
size, while the neural embedding strategy aims at identifying more accurate clusters
and relies on the StarSpace toolkit and its neural embedding model [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Applied to
the lexical index LexI, the neural embedding model learns vector representations for
the individual words in the index keys, and for the individual entity identifiers in the
1.00
0.99
0.98
      </p>
      <sec id="sec-3-1">
        <title>In this section we provide empirical evidence of the suitability of the presented approach</title>
        <p>to divide an ontology matching task.2 We rely on the datasets of the Ontology
Alignment Evaluation Initiative (OAEI), more specifically, on the matching tasks provided in
the anatomy (AMA-NCIA), largebio (FMA-NCI, FMA-SNOMED, SNOMED-NCI)
and phenotype (HPO-MP, DOID-ORDO) tracks.</p>
        <p>Adequacy of clustering strategies. We have evaluated the clustering strategies in terms
of the coverage with respect to the OAEI reference alignments3 and the size of the
ontologies in the matching subtasks. We have compared the two strategies for different
number of matching subtasks n 2 f2; 5; 10; 20; 50; 100; 200g. Figure 2 shows the
coverage of the different divisions DMnT for the naive (left) and neural embedding (right)
strategies. The results are very good and, in the worst case, approx. 93% of the available
200 .
reference mappings in SNOMED-NCI are covered by the matching subtasks in DMT</p>
        <p>
          The scatter plots in Figure 3 visualize, for the FMA-NCI case, the size of the source
modules against the size of the target modules for the matching subtasks in each
division DMnT . For instance, the (orange) triangles represent points jSig(O1i)j; jSig(O2i)j
being O1i and O2i the ontologies (with i=1,. . . ,5) in the matching subtasks of DnM5T . The
naive strategy leads to rather balanced and similar tasks for each division DMT , while
the neural embedding strategy has more variability in the size of the tasks within a given
division DMnT . Nonetheless, on average, the size of the matching subtasks in the neural
embedding strategy are significantly smaller than the ones in the naive strategy.
Evaluation of OAEI systems. We have selected the following four systems from the
latest OAEI campaigns: Mamba, FCA-Map, KEPLER, and POMap. These systems were
unable to complete, given some computational constraints, some OAEI tasks. Table 2
n
shows the obtained results with different divisions DMT computed by the naive and
2 Extended evaluation material in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] and https://doi.org/10.5281/zenodo.1214149
3 A mapping is covered if it can (potentially) be discovered in one or more matching subtasks.
35000
30000
25000
e
l
u
od20000
m
t
e
rag15000
tze
iS10000
5000
0
35000
30000
25000
20000
15000
10000
5000
(a) Naive strategy (b) Neural embedding strategy
        </p>
        <p>Fig. 3: Source and target module sizes in the computed subtasks for FMA-NCI.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Mamba</title>
      </sec>
      <sec id="sec-3-3">
        <title>KEPLER FMA-NCI 2017 FCA-Map FMA-NCI 2016</title>
      </sec>
      <sec id="sec-3-4">
        <title>POMap FMA-NCI 2017</title>
        <p>neural embedding strategies, for the OAEI tasks that the selected systems failed to
compute results. For example, Mamba was able to complete the OAEI 2015 Anatomy track
with divisions DM20 T and DM50 T involving 20 and 50 matching subtasks, respectively.</p>
      </sec>
      <sec id="sec-3-5">
        <title>The subtasks generated by the neural embedding strategy lead to much lower times.</title>
      </sec>
      <sec id="sec-3-6">
        <title>The results are encouraging and suggest that the proposed method to divide an ontology matching task (i) leads to a very limited information loss (i.e., high coverage), and (ii) enables new systems to complete large-scale OAEI tasks.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Jimenez-Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agibetov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samwald</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cross</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Breaking-down the Ontology Alignment Task with a Lexical Index and Neural Embeddings</article-title>
          . arXiv (
          <year>2018</year>
          ) Available from: https://arxiv.org/abs/
          <year>1805</year>
          .12402.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Euzenat</surname>
          </string-name>
          , J.:
          <article-title>Ontology matching: State of the art and future challenges</article-title>
          .
          <source>IEEE Trans. Knowl. Data Eng</source>
          .
          <volume>25</volume>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          )
          <fpage>158</fpage>
          -
          <lpage>176</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jime</surname>
          </string-name>
          <article-title>´nez-</article-title>
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuenca Grau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>LogMap: Logic-Based and Scalable Ontology Matching</article-title>
          . In: International Semantic Web Conference. (
          <year>2011</year>
          )
          <fpage>273</fpage>
          -
          <lpage>288</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fisch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chopra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bordes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>StarSpace: Embed All The Things! arXiv (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Cuenca</given-names>
            <surname>Grau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Horrocks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Kazakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Sattler</surname>
          </string-name>
          ,
          <string-name>
            <surname>U.</surname>
          </string-name>
          :
          <article-title>Modular reuse of ontologies: Theory and practice</article-title>
          .
          <source>J. Artif. Intell. Res</source>
          .
          <volume>31</volume>
          (
          <year>2008</year>
          )
          <fpage>273</fpage>
          -
          <lpage>318</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>