<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Aggregation of Similarity Measures in Ontology Matching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lihua Zhao</string-name>
          <email>lihua@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryutaro Ichise</string-name>
          <email>ichise@nii.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Informatics</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an aggregation approach of similarity measures for ontology matching called n-Harmony. The n-Harmony measure identifies top-n highest values in each similarity matrix to assign a weight to the corresponding similarity measure for aggregation. We can also exclude noisy similarity measures that have a low weight and the n-Harmony outperforms previous methods in our experimental tests.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Ontology matching is a promising research field that discovers similarities
between two ontologies and is widely used in applications such as semantic web,
biomedical informatics and software engineering. Most of current ontology
matching systems combine different similarity measures. For instance, the authors in
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] applied the Ordered Weighted Average(OWA) to combine similarity measures
and Ichise[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] proposed a machine learning approach to aggregate 40 similarity
measures.
      </p>
      <p>
        Harmony[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] measure is a state-of-the-art adaptive aggregation method that
assigns a higher weight to reliable and important similarity measure and a lower
weight to those fail to map similar ontologies. The harmony weight for a
similarity measure is calculated according to the number of the highest values in the
corresponding similarity matrix. However, the harmony measure has drawbacks
when there exist other similarity measures that are as important as the ones
with the highest similarity value. Hence, we extended the harmony measure by
considering top-n values in each row and column of similarity matrices and we
call this method as n-Harmony measure. The top-n is calculated according to
the number of concepts in two ontologies. Our extended n-Harmony considers
more values in similarity matrices and only aggregates similarity measures that
have a high harmony weight.
We applied 13 different similarity measures for aggregation which include 4
string-based, 1 structure-based and 8 WordNet-based similarity measures[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The
final aggregated similarity matrix is Pk(nHk×|SSMMaattrriixx|k(Os,Ot)) , where nHk is
nHarmony weight and SMatrix is the similarity matrix of each similarity measure
between ontology Os and Ot. Before combining the similarity matrices, we
remove min(L-1, nH × L) lowest values in each row and column of similarity
matrix, where L is the minimum number of concepts in two ontologies and nH
represents harmony weight of corresponding similarity matrix. Furthermore, only
those similarity matrices with a high harmony weight are aggregated for the final
similarity matrix. The final decision of whether a ontology pair is matching or
not depends on the final similarity matrix and manually tuned threshold.
      </p>
      <p>
        Directory data sets3 and Benchmark data sets4 from OAEI5 are tested with
our system. The n-Harmony measure returns best result on Directory data sets
when the threshold is 0.45 with 0.86 recall and 0.70 F-measure while the
original harmony measure returns 0.81 recall and 0.68 F-measure. This result is also
better than the results of best systems in OM2009[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], such as ASMOV which
reaches 0.65 recall and 0.63 F-measure on the Directory data sets. On the
Benchmark data sets, n-Harmony performs the same as the original harmony measure
or returns slightly better recalls and F-measures than harmony. Comparing with
the ASMOV, n-Harmony performs almost the same on data sets #101-104 and
#221-247 and returns higher precisions on #302-304, but slightly lower recalls.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Conclusions and Future Work</title>
      <p>Experimental results show that our n-Harmony outperforms original harmony
measure on most of the Directory and Benchmark data sets and also comparable
with the best systems attended in OM2009. However there are still rooms to
improve our n-Harmony measure by exploring advanced structure-based similarity
measures and by investigating automatic threshold selection method rather than
manually tuning the threshold to find out the best performance.
3 http://oaei.ontologymatching.org/2009/directory/
4 http://oaei.ontologymatching.org/2009/benchmarks/
5 http://oaei.ontologymatching.org/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Results of the ontology alignment evaluation initiative 2009</article-title>
          .
          <source>In: Proceedings of the 4th International Workshop on Ontology Matching</source>
          . pp.
          <fpage>82</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Euzenat</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shvaiko</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Ontology matching. Springer-Verlag, Heidelberg (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ichise</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>: Machine learning approach for ontology matching using multiple concept similarity measures</article-title>
          .
          <source>In: Proceedings of the 7th IEEE/ACIS International Conference on Computer and Information Science</source>
          . pp.
          <fpage>340</fpage>
          -
          <lpage>346</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haase</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qi</surname>
          </string-name>
          , G.:
          <article-title>Combination of similarity measures in ontology matching using the owa operator</article-title>
          .
          <source>In: Proceedings of the 12th International Conference on Information Processing and Management of Uncertainty in Knowledge-Based Sytems</source>
          . pp.
          <fpage>243</fpage>
          -
          <lpage>250</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spring</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An adaptive ontology mapping approach with neural network based constraint satisfaction</article-title>
          .
          <source>In: Web Semantics: Science, Services and Agents on the World Wide Web</source>
          . pp.
          <fpage>14</fpage>
          -
          <lpage>25</lpage>
          . No.
          <volume>8</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>