<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Results of GeRoMeSuite for OAEI 2008</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christoph Quix</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Geisler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Kensche</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiang Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Informatik 5 (Information Systems) RWTH Aachen University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>221</fpage>
      <lpage>231</lpage>
      <abstract>
        <p>GeRoMeSuite is a generic model management system which provides several functions for managing complex data models, such as schema integration, definition and execution of schema mappings, model transformation, and matching. The system uses the generic metamodel GeRoMe for representing models, and because of this, it is able to deal with models in various modeling languages such as XML Schema, OWL, ER, and relational schemas. A component for schema matching and ontology alignment is also part of the system. We participated this year the first time in the OAEI contest in order to evaluate and compare the performance of our matcher component with other systems. Therefore, we focused our efforts on the 'benchmark' track.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>State, purpose, general statement</title>
      <p>As a generic model management tool, GeRoMeSuite provides several matchers which
can be used for matching models in general, i.e. our tool is not restricted to a
particular domain or modeling language. Therefore, the tool provides several well known
matching strategies, such as string matchers, Similarity Flooding, children and parent
matchers, matchers using WordNet, etc. In order to enable the flexible combination of
these basic matching technologies, matching strategies combining several matchers can
be configured in a graphical user interface.</p>
      <p>
        Because of its generic approach, GeRoMeSuite is well suited for matching tasks
across heterogeneous modeling languages, such as matching XML Schema with OWL.
We discussed in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] that the use of a generic metamodel, which represents the semantics
of the models to be matched in detail, is more advantageous for such heterogeneous
matching tasks than a simple graph representation.
      </p>
      <p>
        Furthermore, GeRoMeSuite is a holistic model management and not limited to schema
matching or ontology alignment. It supports also other model management tasks such as
schema integration [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], model transformation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], mapping execution and composition
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
1.2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Specific techniques used</title>
      <p>The basis of GeRoMeSuite is the representation of models (including ontologies) in the
generic metamodel GeRoMe. Any kind of model is transformed first into the generic
representation, then the model management operators can be applied to the generic
representation. The main advantage of this approach is that operators have to be
implemented only once for the generic representation. In contrast to other (matching)
approaches which use a graph representation without detailed semantics, our approach is
based on the semantically rich metamodel GeRoMe which is able to represent modeling
features in detail.</p>
      <p>For the OAEI campaign, we focused on improving our matchers for the special
case of ontology alignment, e.g. we added some features which are useful for
matching ontologies. For example, the generic representation of models allows the traversal of
models in several different ways. During the tests with the OAEI tasks, we realized that,
in contrast to other modeling languages, traversing the ontologies using another
structure than class hierarchy is not beneficial. Therefore, we configured all our matchers
that take the model structure into account just to work with the class hierarchy.
Furthermore, we implemented so called ‘children’ and ‘parent’ matchers, which propagate the
similarity of elements up and down in the class hierarchy.</p>
      <p>In addition, we also implemented a matcher using WordNet to discover synonyms
in the ontologies. However, as the benchmark track contains only one example which
uses synonyms, we did not include this matcher in the final configuration for the OAEI
campaign.
1.3</p>
    </sec>
    <sec id="sec-4">
      <title>Adaptations made for the evaluation</title>
      <p>As only one configuration can be used for all matching tasks, we worked on strategies
for measuring the quality of an alignment without having a reference alignment. We
compared several statistical measures (such as expected value, variance, etc.) of
alignments with different qualities in order to identify a ‘good’ alignment. Furthermore,
these values can be used to set thresholds automatically.</p>
      <p>During the tests, we made the experience that the expected value of all similarities,
the standard deviation, and the number of mappings per model element can be used to
evaluate the quality of an alignment.</p>
      <p>Furthermore, we experimented with histograms, i.e. a graphical representation of
the distribution of similarity values. Although we could not identify particular patterns
for histograms of ‘good’ or ‘bad’ alignments, we found the histogram quite useful for
working interactively with the matching component of GeRoMeSuite. Fig. 2 shows a
screenshot of GeRoMeSuite, including a window showing the histogram of a match
result. In addition, the upper part of the window shows some statistical values for the
current similarities. In another dialog, the filter can be adapted.</p>
      <p>On a technical level, we implemented a command line interface for the matching
component, as the matching component is normally used from within the GUI
framework of GeRoMeSuite. The command line interface can work in a batch modus in which
several matching tasks and configurations can be processed and compared.
The results for the OAEI campaign 2008 are available at http://www.dbis.rwth-aachen.
de/gerome/results.html
2</p>
      <sec id="sec-4-1">
        <title>Results</title>
        <p>As we participated the first time in the OAEI campaign, we just focused on the
benchmark track. The time used for each matching task was about 5 to 15 seconds.</p>
        <p>Fig. 2. GUI of the Matching Component of GeRoMeSuite with Histogram Dialog
2.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Benchmark</title>
      <p>Overall, our matching component achieved very similar values for precision and recall,
which seems to be rather unusual, if we compare our results with the results of other
systems for previous years, where the precision was usually higher than recall.
Tasks 101-104 These tasks were quite easy as we could achieve a high precision and
recall already with simple string matchers.</p>
      <p>Task Precision Recall
101 1,00 0,79
103 0,94 0,79
104 0,95 0,79</p>
      <p>For task 102 (irrelevant ontology), our matcher identified a few corresponding
elements, such as year and yearValue, or date and year. Depending on the application of
the mapping, such correspondences might be reasonable (e.g. for ontology merging).
Tasks 201-210 In these tasks, the linguistic information could not always be used as
labels or comments were missing. After including also comments into the matching
process, we could improve the match quality for these tasks significantly. For the
synonym task (205), we also tested a matcher which uses WordNet to detect synonyms.
However, as this matcher did not significantly improve the quality of the match result
and required about three times more time than all other matchers together, we dropped
the WordNet matcher from the final configuration.</p>
      <p>Overall, the results are satisfying, except for the case 202, where no linguistic
information at all was available.</p>
      <p>Task Precision Recall
201 0,87 0,79
201-2 0,95 0,79
201-4 0,93 0,79
201-6 0,97 0,79
201-8 0,91 0,79
202 0,17 0,06
202-2 0,74 0,77
202-4 0,93 0,67
202-6 0,94 0,60
202-8 0,77 0,48
203 1,00 0,79
204 1,00 0,79
205 1,00 0,79
206 0,93 0,78
207 0,93 0,78
208 0,99 0,77
209 0,58 0,45
210 0,61 0,73
2.2
The ontologies in these tasks lacked some structural information. As our matcher still
uses string similarity in a first step, the results in this section were still quite reasonable.</p>
      <p>Task Precision Recall
Tasks 232-266 These tasks are some combinations of the tasks before. For most of
the tasks, the performance of our matcher was satisfying, but for some tasks, especially
those without any linguistic information, it produced disappointing results. This gives
some hints for future improvements of our matcher component, e.g. taking into account
the overall structure of the ontology.</p>
      <p>Tasks 301-304 For tasks 301 and 304, our system produce quite reasonable results.
Further improvements could have been achieved, for example, by using the WordNet
matcher for detecting synonyms, but we did not include this matcher because of
performance reasons as explained above. Task 303 could not be processed by our system as
there was a problem with importing this ontology into our generic representation.
3</p>
      <sec id="sec-5-1">
        <title>Comments</title>
        <p>
          A structured evaluation and comparison of ontology alignment and schema matching
components is very useful for the development of such technologies. However,
mappings between models are constructed for various reasons which can result in very
different mapping results. For example, mappings for schema integration may differ
from mappings for data translation. Therefore, different semantics for ontology
alignments should be taken into account in the future, as it has been pointed out for schema
matching in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
4
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Conclusion</title>
        <p>As our tool is neither specialized on ontologies nor limited to the matching task, we
did not expect to deliver very good results. However, we were quite satisfied with the
overall results. In general, we need to work on an improvement of the recall value.
Furthermore, techniques used by other tools presented at the workshop would help us to
improve the quality of the matching result and the performance of our tool. For example,
identification of similar sub-structures in ontologies and semantic verification of the
identified correspondences seem to be promising techniques to improve the quality and
performance of the matching system.</p>
        <p>Acknowledgements: This work is supported by the DFG Research Cluster on Ultra
High-Speed Mobile Information and Communication (UMIC, http://www.umic.
rwth-aachen.de).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Evermann</surname>
          </string-name>
          .
          <article-title>Theories of Meaning in Schema Matching: A Review</article-title>
          .
          <source>Journal of Database Management</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>55</fpage>
          -
          <lpage>82</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          .
          <article-title>Transformation of Models in(to) a Generic Metamodel</article-title>
          .
          <source>Proc. BTW Workshop on Model and Metadata Management</source>
          , pp.
          <fpage>4</fpage>
          -
          <lpage>15</lpage>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Chatti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jarke. GeRoMe: A Generic Role</surname>
          </string-name>
          <article-title>Based Metamodel for Model Management</article-title>
          .
          <source>Journal on Data Semantics</source>
          , VIII:
          <fpage>82</fpage>
          -
          <lpage>117</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li. GeRoMeSuite: A System for Holistic Generic Model Management. C. Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gehrke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Garofalakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aberer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Deshpande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Florescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Chan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ganti</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-C. Kanne</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Klas</surname>
          </string-name>
          , E. J. Neuhold (eds.),
          <source>Proceedings 33rd Intl. Conf. on Very Large Data Bases (VLDB)</source>
          , pp.
          <fpage>1322</fpage>
          -
          <lpage>1325</lpage>
          . Vienna, Austria,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jarke</surname>
          </string-name>
          .
          <article-title>Generic Schema Mappings</article-title>
          .
          <source>Proc. 26th Intl. Conf. on Conceptual Modeling (ER'07)</source>
          , pp.
          <fpage>132</fpage>
          -
          <lpage>148</lpage>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. Li. Generic Schema</given-names>
            <surname>Merging. J. Krogstie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Opdahl</surname>
          </string-name>
          , G. Sindre (eds.),
          <source>Proc. 19th Intl. Conf. on Advanced Information Systems Engineering (CAiSE'07)</source>
          , LNCS, pp.
          <fpage>127</fpage>
          -
          <lpage>141</lpage>
          . Springer-Verlag,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>C.</given-names>
            <surname>Quix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kensche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Matching of Ontologies with XML Schemas using a Generic Metamodel</article-title>
          .
          <source>Proc. Intl. Conf. Ontologies, DataBases, and Applications of Semantics (ODBASE)</source>
          .
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>