<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The importance of cross-lingual information for matching Wikipedia with the Cyc ontology. ?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksander Smywi«ski-Pohl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krzysztof Wrbel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Chair in Computational Linguistics, Jagiellonian University</institution>
          ,
          <addr-line>ul. ojasiewicza 4, 30-348 Krakw</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Computer Science, Electronics and Telecommunications, AGH University of Science and Technology</institution>
          ,
          <addr-line>al. Mickiewicza 30, Krakw</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Polish National Center for Research and Development under LIDER/37/69/L- 3/11/NCBR/2012 grant and partly by the Faculty of Management and Social Communication, Jagiellonian University in Krakow</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we try to answer the question how cross-lingual evidence may improve matching between dierent classication schemas. We concentrate specically on the task of mapping between Wikipedia categories and Cyc terms as well as the classication of Wikipedia articles to the Cyc taxonomy and show how this process may be improved by consuming the evidence that is available in dierent editions of Wikipedia. The results show that the performance of the mapping procedure may be improved from 0.6 to 4.9 percentage points, depending on the number of external Wikipedia editions and the given task.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>ontology mapping</kwd>
        <kwd>classication</kwd>
        <kwd>multilingual data</kwd>
        <kwd>Wikipedia</kwd>
        <kwd>Cyc</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        To answer the question how the additional Wikipedia editions inuence the
performance of the mapping between Wikipedia and Cyc (cf. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) we have dened
the following tasks: 1) mapping of the Wikipedia categories to Cyc terms; 2)
classication of the Wikipedia articles to the Cyc ontology based on the rst
sentences. In each case the decision of selecting the corresponding Cyc term
requires disambiguation of some English expressions against the Cyc ontology. This
decision is based on the contextual data that are available for each Wikipedia
article and category. Consulting of the supplementary Wikipedia editions extends
the context available when making the decision and in general should improve
the performance of the corresponding algorithms.
      </p>
      <p>
        In case of the category mapping (based on the identication of plural head
nouns in category names cf. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]), when an English category is mapped, the
corresponding Dutch, German, etc. categories are inspected. Then the parent and
      </p>
      <p>Aleksander Smywi«ski-Pohl and Krzysztof Wrbel
child categories as well as articles of the corresponding categories in the other
editions are looked up in a reverse interlingual mapping index and if there is an
English Wikipedia page, that was not present in the original context, it is
included in the new, extended context. Then a support value used to disambiguate
the category is computed against the extended context.</p>
      <p>
        In case of article classication (based on the rst sentence parsing, cf. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ])
the supplementary Wikipedia editions provide additional categories for the
classied article, that are used to verify the disambiguation decision. The manner of
operation is similar to that from the previous task the corresponding articles
in other Wikipedia editions are consulted, their categories are translated back
to English and these new categories are included in the extended context.
2
      </p>
      <p>Results
There was a small improvement (F1 increased from 86.8% to 78.4%) in the
performance of the category mapping when the English Wikipedia is supported
by three other Wikipedias ( de,nl,sv). However providing the algorithm with
more data from other Wikipedia editions, increased the computation time, but
did not further improve the results.</p>
      <p>On the other hand the inuence of the additional Wikipedia editions in the
task of the classication of the articles into the Cyc ontology was much stronger.
Not only the additional Wikipedia editions improved the recall, but also the
precision. The maximum precision was achieved for 5 and 6 additional Wikipedias
(96.6% compared to 95.8% for the sole English Wikipedia). The F1 was the
largest for the 8 additional Wikipedias resulting in an increase from 66.2% to
71.1%.</p>
      <p>The overall conclusion from the results is that the inuence of the
supplementary Wikipedias is task dependant and in general the extra time necessary
to pre-process the data and the increase of the computation time may not be
justied. However the task of articles classication shows also that such
supplementary data may be very valuable and may increase both the precision and the
recall of the results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Gangemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nuzzolese</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Presutti</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Draicchio</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musetti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciancarini</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Automatic typing of DBpedia entities</article-title>
          .
          <source>In: The Semantic WebISWC</source>
          <year>2012</year>
          , pp.
          <fpage>6581</fpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Pohl</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Classifying the Wikipedia Articles into the OpenCyc Taxonomy</article-title>
          . In: Rizzo,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Charton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Hellmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Kalyanpur</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the Web of Linked Entities Workshop in conjuction with the 11th International Semantic Web Conference</source>
          . pp.
          <volume>516</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>YAGO: a core of semantic knowledge</article-title>
          .
          <source>In: Proceedings of the 16th international conference on World Wide Web</source>
          . pp.
          <fpage>697706</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>