<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RiMOM-IM Results for OAEI 2014</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chao Shao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Linmei Hu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juanzi Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Tsinghua University</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the results of RiMOM-IM in the Ontology Alignment Evaluation Initiative (OAEI) 2014. We only participated in IM@OAEI2014. We first describe the overall framework of our matching System (RiMOM-IM); then we detail the techniques used in the framework for instance matching. Last, we give a thorough analysis on our results and discuss some future work on RiMOM-IM.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>mismatched instances since instances in two different knowledge bases are usually
described by different numbers of RDF triples.</p>
      <p>
        In order to solve the above challenges in large-scale instance matching, we propose
an iterative instance matching framework RiMOM-IM (RiMOM-Instance Matching),
which is developed based on our ontology matching system RiMOM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The main
idea behind the framework is to maximize the utilization of distinctive and available
matching information. RiMOM-IM presents a novel blocking method to improve the
efficiency and employs a weighted exponential function based similarity aggregation
method to guarantee high accuracy of instance matching.
1.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>State, purpose, general statement</title>
      <p>This section describes the overall framework of RiMOM-IM. The overview of the
instance matching system is shown in Fig. 1. The system includes five modules, i.e.,
Initial Interactive Configuration, Candidate Pair Generation, Matching Score
Calculation, Instance Alignment and Validation. The annotated numbers in the figure show the
sequences of the process. We illustrate the process as follows.</p>
      <p>1. Configuration</p>
      <p>Candidate Pair Generation</p>
      <p>2. Data
Preprocessing
3. Blocking</p>
      <p>Matching Score Calculation
4. SPimreidlaicriattyesover 4A.gSgirmegilaatriiotny
5. For unique instance sets, we iteratively use “Unique Subject Matching” and
“Oneleft Object Matching” to generate aligned set until no new aligned instances are
generated. These aligned instances will then be used to find new candidate pairs
and new unique instances, thus updating candidate set and unique instance sets.
Correspondingly, matching scores for related instance pairs and the priority queue
will be updated.
6. For the priority queue, we use “Score Matching” to generate only one aligned pair
with the highest score above the threshold. If there is a newly aligned instance pair,
we will generate new unique instances, which will be taken as input to step 5. If
there is no new aligned pair, we continue step 7.
7. If Validation module is chosen in step 1, we will conduct validation on all aligned
pairs. Otherwise, terminate.
This year we only participate in the IM@2014 track. We will describe specific
techniques used in this track.</p>
      <p>Data Preprocessing: First, we translate all the languages used in the whole datasets
to English by using google translator. Then we remove special symbols like “♯, *, !”,
etc. and stop words like “a, of, the”, etc. Afterwards, we calculate the TF-IDF values of
words in each knowledge base.</p>
      <p>Blocking: Blocking aims to pick a relatively small set of candidate pairs from all
pairs. Due to the large scale of knowledge bases, it is impossible to calculate matching
scores of all instance pairs. In our blocking method, we take the predicate as well as
top 10 words of the object (ordering by tf-idf values in the knowledge base) as index
keys of instances. It should be noticed that if the object is an instance, the entire URI is
considered as a word. Owing to the novel blocking method which restricts the candidate
pairs with identical distinctive information (predicate and distinctive object features),
we greatly reduce the number of similarity comparisons and improve the efficiency.</p>
      <p>Similarity over Predicates: The similarity function varies with different predicates.
For example, we can use indicator function for the predicate of birthdate, when the
value are the same, the indicator is 1, otherwise, 0. For the predicate of comments, we
compute cosine similarity based on the tf-idf vectors. In system configuration, we can
specify a similarity function for each aligned predicate.</p>
      <p>
        Similarity Aggregation: For each instance pair, after acquiring similarity values
in terms of multiple aligned predicates, we need to aggregate the similarities to get
final matching score. AVG aggregates the similarities by computing the average value
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. SIGMOID(SIG) aggregates the similarities by computing the average similarities
transformed by a sigmoid function [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These methods do not adapt to the case when
different instance pairs have different numbers of aligned predicates. In this work, we
propose a weighted exponential aggregation function, ExpAgg to aggregate the
similarities S, which is a set of similarities of all aligned predicates. The function is as follows:
ExpAgg(S) =
Σsi∈S wi′ exp(wi′′ si)
Σsi∈S wi′ exp(wi′′ 1)
(1)
Among them, si is the similarity score in terms of the ith aligned predicate. We set the
weights of the predict “label” and the other predicts as as 16 and 1, respectively.
      </p>
      <p>Score Matching: In this task, we don’t use the modules of “Unique Subject
Matching” and “One-left Object Matching”. Each time we choose the pair with the highest
score as the aligned pair, we will then update the matching score of each instance pair
in the prior queue. With the greedy algorithm of extracting only the most matching pair
every time, we control error propagation to some extent. As we can not guarantee a
global optimization with the greedy algorithm, we add the later process of validation.</p>
      <p>Validation:Since many objects of instances are URIs referring to other instances,
there still exists some nondeterminacy in aligning two instances due to the uncertainty
in the alignment situation of their compatible neighbors, and we also find some rules
are very useful. We add validation module to correct some mistakes by some useful
rules. In this track, we find that if two instances both contain the predict of “lable”, their
“label” predicts shall share at least a same token.
1.3</p>
    </sec>
    <sec id="sec-3">
      <title>Link to the system and parameters file</title>
      <p>The RiMOM-IM system can be found at http://keg.cs.tsinghua.edu.cn/
project/RiMOM/.
2</p>
      <sec id="sec-3-1">
        <title>Results</title>
        <p>The IM@2014 track contains two subtasks. we present the results and related analysis
for the two subtasks in the following subsections.
2.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Identity Recognition sub-task</title>
      <p>The goal of the Identity Recognition sub-task is to determine whether two OWL
instances refer to the same real-object. Due to a lack of training data, it is very difficult
for us to tuning our parameters. First, we use the default setup to get a preliminary
result, and then we check the information of some aligned pairs. We find out that the
predict of “label” is very important, so we increase the weight of the “label” predict.
Finally, we get 1103 instance pairs as matching ones.</p>
      <p>As show in figure 2, the results for the identity task are: Precision 0.65, Recall 0.49,
Fmeasure 0.56, which is much lower than we expected. But we are pretty sure that if
we have some training set, we can tuning a much better result.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Similarity Recognition sub-task</title>
      <p>The goal of the Similarity Recognition sub-task is to determine the degree of
similarity between two OWL instances, even when the two instances describe different
realobjects. In our system, we use the traditional cosine similarity measurement, however,
if one predicate have many similarity values, we use the maximum value. So in
summary, we use maxpooling+cosine similarity. We can find that if two instances describe the
same real-objects, their similarity value will usually be larger than that of two instances
which describe different real-objects. Therefore, it’s reliable to use the similarity value
as a measure to judge whether two instances describe the same real-objects. And we find
that if two instances describe different real-objects while have a high similarity value,
then their labels are usually different. We can use these two observations for instance
matching. As show in figure 3, this similarity strategy is much close to the
crowdsourcing activities’. We have chosen other complexity similarity measurements, but it turns
out that this simple measurements works better.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Discussions on the way to improve the proposed system</title>
      <p>Our system need the aligned predicates to select the candidate instance pairs. Our
system will use the aligned instance pairs to calculate the similarity values of other instance
pairs, which will also need the information of aligned predicates. We need to invent an
algorithm to automatically align the predicates. Although there are some algorithms that
can align the instances by measuring the similarity values of predicates, none of them
use the aligned instance pairs to help to update the similarity values for predicates. We
will develop an algorithm that do not need any aligned predicates, but can iteratively
use the aligned instance pairs to align the predicates, which will in turn advance the
instance alignment.
This task is a cross-lingual instance matching task. We find out that we can significantly
improve the result by using translation method. And we find that the blocking method
also improves the precision of the result. Because the “datatype” of the “object” is
always “String”, we do not have any relations between any two instances. So we can’t
use the relation information to improve the recall. This year we use ten keywords for
every predict to get more candidate pairs to ensure a high recall.
3</p>
      <sec id="sec-6-1">
        <title>Conclusion and future work</title>
        <p>In this paper, we present the system of RiMOM-IM in OAEI 2014 Campaign. We
participate in one track this year. We described specific techniques we used during this
campaign. In our project, we design a new framework to do the instance matching task.
Our method effective and efficient.</p>
        <p>For now, we need to tune the parameter manually, we will improve it by making the
tuning process automatic in the future work.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dbpedia - A crystallization point for the web of data</article-title>
          .
          <source>J. Web Sem</source>
          .
          <volume>7</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hoffart</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>YAGO2: A spatially and temporally enhanced knowledge base from wikipedia</article-title>
          .
          <source>Artif. Intell</source>
          .
          <volume>194</volume>
          (
          <year>2013</year>
          )
          <fpage>28</fpage>
          -
          <lpage>61</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Tang</surname>
          </string-name>
          , J.:
          <article-title>Xlore: A large-scale english-chinese bilingual knowledge graph</article-title>
          .
          <source>In: Proceedings of the ISWC 2013 Posters &amp; Demonstrations Track</source>
          , Sydney, Australia, October
          <volume>23</volume>
          ,
          <year>2013</year>
          . (
          <year>2013</year>
          )
          <fpage>121</fpage>
          -
          <lpage>124</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          :
          <article-title>Rimom: A dynamic multistrategy ontology alignment framework</article-title>
          .
          <source>IEEE Trans. Knowl. Data Eng</source>
          .
          <volume>21</volume>
          (
          <issue>8</issue>
          ) (
          <year>2009</year>
          )
          <fpage>1218</fpage>
          -
          <lpage>1232</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jean-Mary</surname>
            ,
            <given-names>Y.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shironoshita</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kabuka</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          :
          <article-title>Ontology matching with semantic verification</article-title>
          .
          <source>J. Web Sem</source>
          .
          <volume>7</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <fpage>235</fpage>
          -
          <lpage>251</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>