<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontology Matching Evaluation: A Statistical Perspective</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Majid Mohammadi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wout Hofman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yao-hua Tan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Technical Science, The Netherlands Institute of Applied Technology (TNO)</institution>
          ,
          <addr-line>Soesterberg</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Technology</institution>
          ,
          <addr-line>Policy and Management</addr-line>
          ,
          <institution>Delft University of Technology</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper proposes statistical approaches to test if the difference between two ontology matchers is real. Speci cally, the performances of the matchers over multiple data sets are obtained and based on their performances, the conclusion can be drawn whether one method is better than one another or not. To do so, the paired t-test and Wilcoxon signed rank test are proposed and the comparisons over six recently proposed methods are reported.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology alignment</kwd>
        <kwd>evaluation</kwd>
        <kwd>statistical inference</kwd>
        <kwd>paired t-test</kwd>
        <kwd>Wilcoxon signed rank</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>where P^i is the average performances of the matcher i.</p>
      <p>Before running any statistical test, the signi cant level must be determined. the is the
probability of rejecting null hypothesis when the null hypothesis is true. To the best of
our knowledge, no statistical techniques have been employed to test the above-mentioned
hypothesis. Firstly, the widely-used paired t-test is presented with more detail. Having
hard preconditions to be satis ed, it must be warned that t-test might be inappropriate
and statistically unsafe. Thus, the Wilcoxon signed-rank test is presented which is able
to detect more difference even though the number of samples are not large enough.
2.1</p>
      <p>Paired t-test
A common way to check if the difference between two matchers on different data sets is
not random is to compute the paired t-test. Let di = Pi1 Pi2 be the difference between the
performances of two matchers over i th data set. The t statistics is computed as t = x^dx^
where x^ and ^d are sample average and standard deviation of samples, respectively. This
statistics is distributed according to the Student distribution with N 1 degree of freedom.
After obtaining the probability of observing the data given that H0 being true (p-value)
according to the Student distribution, the H0 can be rejected if p value &lt; and then
Ha is accepted.
2.2</p>
      <p>Wilcoxon Signed Rank test
The non-parametric alternative to the paired t-test is Wilcoxon singed rank test. This
method ranks the absolute values of performance differences of two matchers. Then, it
compares the rank of positive and negative differences. After computing the difference
between two matchers over the the i th data set, di , the differences are ranked based
on the values of di , disregarding its sign. if di = 0 it is ignored and the average ranks are
assigned if the performances over one data set ties. Assume W + = ∑di&gt;0 rank(di) and
W</p>
      <p>= ∑di&lt;0 rank(di) and T = min(W +; W ). Then z = √ 214TN(41NN+(1N)(+21N)+1) is distributed
according to the normal distribution.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Experimental Results</title>
      <sec id="sec-2-1">
        <title>XMAP</title>
      </sec>
      <sec id="sec-2-2">
        <title>XMAP</title>
        <p>AML 0.640625
AML2014 0.00647436 0.01596065
CroMatcher 0.00097656 6.10E-05</p>
        <p>edna 0.000822 0.011231
refalign 0.000977 6.10E-05</p>
        <p>AML AML2014 CroMatcher edna refalign
0.526403 0.23326767 0.00094182 0.000972 0.000939
0.05359674 0.00079181 0.113909 0.000697
0.00026227 0.243871 0.000243
2.83E-06 0.01664</p>
        <p>4.75E-06</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>