<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The First Version of the OAEI Complex Alignment Benchmark</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elodie Thiéblin</string-name>
          <email>elodie.thieblin@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michelle Cheatham</string-name>
          <email>michelle.cheatham@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cassia Trojahn</string-name>
          <email>cassia.trojahn@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondřej Zamazal</string-name>
          <email>ondrej.zamazal@vse.cz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lu Zhou</string-name>
          <email>zhou.34@wright.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IRIT &amp; Université de Toulouse 2 Jean Jaurès</institution>
          ,
          <addr-line>Toulouse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Economics</institution>
          ,
          <addr-line>Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Wright State University</institution>
          ,
          <addr-line>Dayton</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present the first version of the complex benchmark of the Ontology Alignment Evaluation Initiative campaigns. This benchmark is composed of four datasets from different domains (conference, hydrology, geoscience and agronomy) and covers different evaluation strategies.</p>
      </abstract>
      <kwd-group>
        <kwd>complex ontology alignments</kwd>
        <kwd>evaluation dataset</kwd>
        <kwd>OAEI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Complex correspondences involve transformation functions of literal values or
logical constructors (e.g. 8x, ekaw:AcceptedPaper(x) 9y, cmt:acceptedBy(x,y)),
which make them more expressive than simple correspondences. Complex
alignments, composed of at least one complex correspondence, are therefore a
complement to simple alignments. Different approaches for complex matching have
emerged in the literature [
        <xref ref-type="bibr" rid="ref2 ref4 ref5 ref8">2,4,5,8</xref>
        ]. Most of them, however, have been evaluated
on tailored datasets (e.g., targeting a specific correspondence pattern). Most
efforts on systematic evaluation, in the context of the OAEI campaigns1, are still
dedicated to simple matchers.
      </p>
      <p>This paper presents the first version of the OAEI complex track, composed of
four datasets (Table 1) from different domains. This domain and correspondence
variety allows for better covering different kinds of heterogeneity between
ontologies. Different evaluation strategies aim at evaluating complex matchers under
different perspectives. The evaluation will be supported by the SEALS platform
and the output alignments must be in EDOAL. The detail of each dataset and
evaluation process can be found on the OAEI’s 2018 complex track webpage2,
and are introduced in the following.</p>
      <p>1http://oaei.ontologymatching.org/
2http://oaei.ontologymatching.org/2018/complex/index.html</p>
    </sec>
    <sec id="sec-2">
      <title>Conference consensual dataset</title>
      <p>
        This dataset is based on the OntoFarm dataset [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which is composed of 16
ontologies on the conference organisation domain and simple reference alignments
between 7 of them. Here, we consider 3 out of the 7 ontologies from the
reference alignments (cmt, conference and ekaw ), resulting in 3 alignment pairs. The
alignments involve both logical constructors (76 correspondences) and
transformations (3 correspondences). Examples are given in the following :
1. 8 x, ekaw:AcceptedPaper(x) 9 y, cmt:acceptedBy(x,y) is a correspondence
with the existential constructor.
2. 8 x,y, cmt:name(x,y) 9 y1, y2, conference:has_the_first_name(x,y1) ^
conference:has_the_last_name(x,y2) ^ concatenation(y,y1," ", y2), where
concatenation(a,b1, b2, ...) is a predicate ensuring that its first parameter
a is equal to the string concatenation of the others {b1, b2, ...}. It uses a
transformation function of the literal values.
      </p>
      <p>
        The alignments have been manually created by three experts in the domain,
following the methodology in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Four experts assessed the generated
correspondences to reach a consensus. The systems will be manually evaluated on
their output alignments to produce precision and recall scores. Only the
complex equivalence correspondences will be assessed. The systems can use a simple
reference alignment as input. Confidence scores of correspondences will not be
taken into account in the evaluation.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Hydrography dataset</title>
      <p>
        The hydrography dataset is composed of 4 source ontologies (Hydro3,
hydrOntology_native, hydrOntology_translated and Cree) that each should be aligned
to a single target Surface Water Ontology (swo). The source ontologies vary
in their similarity to the target ontology – Hydro3 is similar in both language
and structure, hydrOntology_native and hydrOntology_translated are similar
in structure but hydrOntology_translated is in Spanish rather than English, and
Cree is very different in terms of both language and structure. The alignments
were created by a geologist and an ontologist, in consultation with a native
Spanish speaker regarding the hydrOntology_translated, and consist of logical
relations such as the one shown below.
1. 8x, hydrOntology_translated:Aguas_Corrientes(x) swo:SurfaceFeature(x)
^ swo:Waterbody(x) ^ 9y, swo:hasFlow(x,y) ^ swo:Flow(y)
Performance on this dataset will be evaluated on three sub-tasks: 1)
identifying the atoms (classes and properties) from the target ontology involved in the
relations (e.g., swo:SurfaceFeature, swo:Waterbody, swo:hasFlow and swo:Flow
from the correspondence above), 2) when given the atoms, identifying the logical
relations that hold between them and 3) the full complex alignment task.
Evaluation of the first sub-task will use traditional F-measure, while the remaining
two subtasks will be evaluated on semantic F-measure [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>GeoLink dataset</title>
      <p>
        This dataset is from the GeoLink project3, which was funded under the U.S.
National Science Foundation’s EarthCube initiative. It is composed of 2 populated
ontologies: the GeoLink base ontology (gbo) and the GeoLink modular ontology
(gmo). The GeoLink project is a real-world use case of ontologies. The alignment
between the ontologies was developed in consultation with domain experts from
several Geoscience research institutions. The complex correspondences include
not only class and property subsumption and property chains (described in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]),
but also some that involve typecasting (c.f. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]), for example:
1. Property Chain: 8x,z, gbo:Award(x) ^ gbo:hasSponsor(x,z)
9y, gmo:FundingAward(x) ^ gmo:providesAgentRole(x,y) ^
gmo:SponsorRole(y) ^ gmo :performedBy(y,z)
2. Class Typecasting: 8x, gbo:PlaceType(x) rdfs:subClassOf(x, gmo:Place)
More information about this dataset can be found in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and the benchmark
and alignment can be downloaded here4. The performance of alignment systems
on this dataset will be evaluated in the same way as the hydrography dataset.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Taxon dataset</title>
      <p>
        This dataset is composed of 4 populated ontologies whose common scope is plant
taxonomy: AgronomicTaxon (agtx ), Agrovoc (agv and agronto), DBpedia (dbo)
and TaxRef-LD (txr ). This dataset extends the one proposed in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] by adding
the TaxRef-LD ontology. The alignments were manually created with the help
of one expert and involve only logical constructors, as for example:
1. 8x, agtx:GenusRank(x) agronto:hasTaxonomicRank(x,agv:c_11125)
2. 8x, agtx:GenusRank(x) 9y, dbo:Species(y) ^ dbo:genus(y,x) ^ dbo:Species(x)
The evaluation of this dataset is task-oriented. We will evaluate the generated
correspondences using a SPARQL query rewriting system and manually
measure their ability of answering a set of queries over each dataset. For example, a
competency question could be “Retrieve all the genus taxa”. For
AgronomicTaxon, as source ontology, the corresponding SPARQL query is SELECT ?x
WHERE {?x a agtx:GenusRank.} and the correspondences output by the
systems with Agrovoc as target ontology, should be able to translate the query
into: SELECT ?x WHERE {?x agronto:hasTaxonomicRank agv:c_11125.}
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>This paper has presented the first OAEI complex evaluation track, covering
different kinds of complex correspondences, domains and evaluation strategies.
For most datasets, the evaluation is still manually performed, opening directions
on how complex alignments can be automatically generated and evaluated.
Acknowledgements. We thank Catherine Roussey (IRSTEA) and Nathalie
Hernandez (IRIT) for their help on the Taxon dataset and Dalia Varanka (US
Geological survey) for her work on the hydrography dataset. Ondřej Zamazal
has been partially supported by the CSF grant no. 18-23964S. Creation of the
GeoLink dataset was funded by NSF 1440202.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Euzenat</surname>
          </string-name>
          , J.:
          <article-title>Semantic precision and recall for ontology alignment evaluation</article-title>
          .
          <source>In: IJCAI 2007, Proceedings of the 20th International Joint Conference on Artificial Intelligence</source>
          , Hyderabad, India, January 6-
          <issue>12</issue>
          ,
          <year>2007</year>
          . pp.
          <fpage>348</fpage>
          -
          <lpage>353</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lowd</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kafle</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dou</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Ontology matching with knowledge rules. In: Transactions on Large-Scale Data-</article-title>
          and
          <string-name>
            <surname>Knowledge-Centered</surname>
            <given-names>Systems</given-names>
          </string-name>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>95</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Krisnadhi</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janowicz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>On the capabilities and limitations of OWL regarding typecasting and ontology design pattern views</article-title>
          . In: Ontology Engineering - 12th
          <source>International Experiences and Directions Workshop on OWL, OWLED</source>
          <year>2015</year>
          ,
          <article-title>co-located with ISWC 2015, Bethlehem</article-title>
          , PA, USA, October 9-
          <issue>10</issue>
          ,
          <year>2015</year>
          , Revised Selected Papers. pp.
          <fpage>105</fpage>
          -
          <lpage>116</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Parundekar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knoblock</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ambite</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Discovering concept coverings in ontologies of linked data sources</article-title>
          .
          <source>In: ISWC</source>
          . pp.
          <fpage>427</fpage>
          -
          <lpage>443</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ritze</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meilicke</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Šváb Zamazal</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stuckenschmidt</surname>
          </string-name>
          , H.:
          <article-title>A pattern-based ontology matching approach for detecting complex correspondences</article-title>
          .
          <source>In: 4th OM workshop</source>
          . pp.
          <fpage>25</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Thiéblin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amarger</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roussey</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trojahn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Cross-querying lod datasets using complex alignments: an application to agronomic taxa</article-title>
          .
          <source>In: MTSR</source>
          . pp.
          <fpage>25</fpage>
          -
          <lpage>37</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Thiéblin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haemmerlé</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trojahn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Task-oriented complex ontology alignment - two alignment evaluation sets</article-title>
          .
          <source>In: ESWC</source>
          (
          <year>2018</year>
          ), (to appear)
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Walshe</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brennan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Bayes-recce: A bayesian model for detecting restriction class correspondences in linked open data knowledge bases</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems</source>
          <volume>12</volume>
          (
          <issue>2</issue>
          ),
          <fpage>25</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Zamazal</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Svátek</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The Ten-Year OntoFarm and its Fertilization within the Onto-Sphere</article-title>
          .
          <source>Journal of Web Semantics</source>
          <volume>43</volume>
          ,
          <fpage>46</fpage>
          -
          <lpage>53</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheatham</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krisnadhi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hitzler</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A complex alignment benchmark: Geolink dataset</article-title>
          .
          <source>In: ISWC</source>
          . Springer (
          <year>2018</year>
          ), (to appear)
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>