<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Benchmark for Testing Instance-Based Ontology Matching Methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Katrin Zaiss</string-name>
          <email>zaiss@cs.uni-</email>
          <email>zaiss@cs.uniduesseldorf.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sven Vater</string-name>
          <email>Sven.Vater@uni-</email>
          <email>Sven.Vater@uniduesseldorf.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Conrad</string-name>
          <email>conrad@cs.uni-</email>
          <email>conrad@cs.uniduesseldorf.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computer Science</institution>
          ,
          <addr-line>Universitaetsstr. 1, 40225 Duesseldorf</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The matching of ontologies is a problem solved by many di erent matching systems using various algorithms. To test di erent methods or a complete system or to compare the systems among each other a common data set is needed. There are already some benchmarks containing many test scenarios available, but they mainly focus on concept-based matching algorithms or on instance matching (the process of nding similar instances). Instance-based methods cannot be tested su ciently, because the ontologies do not contain instances at all or the number of instances is very small. In this poster we introduce a new benchmark, ONTOBI, which makes use of Wikipedia to create a benchmark test series with ontologies that contain many instances.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Ontologies represent knowledge in a structure way. In many
application areas there is a need to match ontologies, e.g.
in the eld of query answering on heterogeneous sources. In
the past many matching systems have been developed to
cope with this problem, an overview can be found in [ES07].
Generally, the used methods can be divided into concept-,
structure- and instance-based approaches and in most cases
a matching system uses a combination of those approaches.
To test the e ciency of single algorithms or complete
systems, or to compare systems among each other based on a
common data set, appropriate ontologies and test cases have
to be developed.</p>
      <p>Currently, there are already some benchmarks, e.g. the one
published by the OAEI [OAE09], the STBenchmark [ATV08]
or the IIMB [FLMV08]. The OAEI benchmark ontologies
only contain a very small number of instances for a few
concepts. In 2009 some instance-matching tasks have been
added, but they cannot easily be adapted for instance-based
ontology matching methods, because the concept
information remains the same for all tests and the reference
alignment only contains instance-to-instance correspondences. The
instances of the ontologies created with the STBenchmark
are created arti cially and do not contain instance
variations/modi cations and the IIMB benchmark is again
designed for instance matching tasks and only provides a small
ontology with no changes on the concept level. The lack of a
reasonable amount of instances in existing benchmarks
motivates the development of an additional benchmark, which
we present in this paper.
2. DEVELOPING THE BENCHMARK
Under consideration of the advantages and disadvantages of
existing benchmarks and regarding the personal experiences
with matching systems, we de ned a set of requirements that
our new benchmark should ful ll. First of all, the general
requirements for evaluation frameworks as described in [ES07]
should be considered, i.e. systematic procedure, continuity,
quality and equity, dissemination and intelligibility.
Additionally, we formulated some more criteria: bigger amount of
instances, varying structure, di erent data formats, spelling
mistakes and 1:n mappings. As described before, our focus
is set to a huge amount of instances to enable the
evaluation of instance-based matching methods, but we also want
to provide a complete benchmark with which all kinds of
methods and systems can be tested.</p>
      <p>Similar to the OAEI benchmark, our ONTOBI benchmark
consists of di erent test scenarios, whereas a reference
ontology provides the basis for each test case. The reference
ontologies gets modi ed by applying one or more of the
modi cations described in Table 1. The modi cations are
applied on di erent parts of the ontology (instance set,
concept names, etc.), and in most cases only a subset of the
according data set is changed. This modi ed ontology has
to be matched against the original reference ontology, an
overview of the process is given in Figure 1. The reference
alignment is given as well (the format for this alignment is
borrowed from the Ontology Alignment API [Euz06]), such
that the results can be evaluated by e.g. calculating
Precision and Recall.</p>
      <p>We decided to use Wikipedia as the basis for our reference
ontology, because Wikipedia provides a lot of information
within its structured infoboxes, and developed a tool, that
extracts concepts, attributes, relations and instances out
of these infoboxes. The reference ontology consists of 17
classes, 13 object properties and 128 data type properties.
It is constructed around di erent concepts describing the
geographical structure of our world, i.e. countries, states,</p>
      <p>modi cation
spelling mistakes
changed format
di erent naming conventions
suppressed comments</p>
      <p>no data types
overlapping data sets</p>
      <p>subset data sets
expanded structure
attened structure
another language
random names</p>
      <p>synonyms
disjunct data sets</p>
      <sec id="sec-1-1">
        <title>Test case reference ontology mods</title>
      </sec>
      <sec id="sec-1-2">
        <title>Alignment modified ontology</title>
        <p>cities and languages. Additionally there are concepts
describing di erent kinds of entertainment instruments, such
as books, movies or songs with their corresponding authors,
actors and singers. Another part of the ontology deals with
companies and their products, e.g. cars, mobile phones or
magazines. The most important issue for ONTOBI is the
instance set. Currently, the reference ontology contains more
than 3500 instances, but the number grows constantly.
The di erent combinations of modi cation that we applied
on the reference ontology and hence the di erent test cases
can be found in Table 2. All modi cations are executed
manually by using an ontology editor like Protege [Pro09].
The complete benchmark, including the ontologies and the
reference alignmenta, are available for download on demand.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. FUTURE WORK</title>
      <p>The work on this benchmark is still in progress. Currently,
we enhanced the quality of our benchmark by directly
deriving the ontologies from DBPedia [LBK+09] (see
www.dbs.cs.uni-duesseldorf.de/projekte/ONTOBI). We also
increased the number of concepts, attributes and instances,
the modi cations have been slightly changed and the test
cases have been reorganized. Additionally, an ontology
modi cator has been implemented which automatically applies
selected transformation on the reference ontology. In future
modi cation(s)</p>
      <p>M
S1
I2
L4
L2
L3
H1
H2
work we want to focus on developing more complex modi
cations on the instance level.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [ATV08]
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Alexe</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wang-Chiew Tan</surname>
          </string-name>
          , and Yannis Velegrakis.
          <article-title>STBenchmark: Towards a Benchmark for Mapping Systems</article-title>
          .
          <source>Proc. VLDB Endow</source>
          .,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <volume>230</volume>
          {
          <fpage>244</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [ES07]
          <article-title>Jer^ome Euzenat and Pavel Shvaiko</article-title>
          .
          <source>Ontology Matching</source>
          . Springer-Verlag, Heidelberg (DE),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Euz06]
          <article-title>Jero^me Euzenat. An API for ontology Alignment (version 2</article-title>
          .1). http://gforge.inria.fr/docman/ view.php/117/251/align.pdf,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[FLMV08] Al o Ferrara</source>
          , Davide Lorusso, Stefano Montanelli, and
          <string-name>
            <given-names>Gaia</given-names>
            <surname>Varese</surname>
          </string-name>
          .
          <article-title>Towards a Benchmark for Instance Matching</article-title>
          .
          <source>In Proceedings of the 3rd International Workshop on Ontology Matching (OM-2008)</source>
          <article-title>Collocated with the 7th International Semantic Web Conference (ISWC-</article-title>
          <year>2008</year>
          ), Karlsruhe, Germany, October
          <volume>26</volume>
          ,
          <year>2008</year>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [LBK+09]
          <string-name>
            <surname>Jens</surname>
            <given-names>Lehmann</given-names>
          </string-name>
          , Chris Bizer, Georgi Kobilarov, S ren Auer, Christian Becker, Richard Cyganiak, and
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Hellmann. DBpedia - A Crystallization</surname>
          </string-name>
          <article-title>Point for the Web of Data</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [OAE09]
          <string-name>
            <given-names>Ontology</given-names>
            <surname>Alignment Evaluation Initiative - OAEI-2009 Campaign</surname>
          </string-name>
          . http://oaei.ontologymatching.org/2009/,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Pro09]
          <article-title>The Protege Ontology Editor and Knowledge Acquisition System</article-title>
          . http://protege.stanford.edu/,
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>