<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Boosting MultiFarm Track with Turkish Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Abderrahmane Khiat</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Beyza Yaman</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanna Guerrini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ernesto Jime´nez-Ruiz</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naouel Karam</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DIBRIS, University of Genoa</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Human-Centered Computing Lab, Freie Universita ̈t Berlin</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The evolution of semantic structured data, such as those behind the deep web or social networks, requires mapping between sources to enable a high level integration. Several ontology matching systems have been developed to establish mappings between multilingual ontologies, however, employing these systems in real world requires an assessment of the ontologies capability and performance which is conducted by the MultiFarm Track. Yet, this track still lacks of ontologies from different language families. In this paper, we contribute to the OAEI initiative with a Turkish dataset to extend the coverage of languages for the matching systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The Turkish language comes from a different branch of the language family than the
existing datasets in the Multifarm Track, we believe it will add a different perspective
to the assessments of matching systems. The dataset is valuable for the OM domain for
several reasons: i) Since the Mutltifarm track is composed of a set of ontologies of the
conference domain, to the best of our knowledge no such dataset exists for the
conference domain in Turkish, ii) we will integrate Turkish datasets to the OAEI campaign
to assess the performance of cross-lingual ontology alignment systems along with other
languages and iii) we will close the gap of lacking datasets for Turkish in the OAEI.</p>
      <p>
        We followed the steps detailed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to create our dataset, to validate it and then to
generate the reference alignments for other ontologies. The Multifarm track has been
translated from English to Turkish semi-automatically, and then reviewed and corrected
by a professional English-Turkish speaker. During dataset generation, we have taken
advantage of our experiences of generating Arabic datasets [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We first generated the
Turkish ontologies via regular expressions with the regex API and, then, translated
entities are replaced by the original ones. Finally, the alignments between Turkish and other
languages are constructed by replacing English entity IDs by Turkish entity IDs.
      </p>
      <p>However, we have a different level of difficulty for generating alignments between
English-Turkish than we have experienced when we were constructing the alignments
for Arabic datasets. The difficulty for English-Arabic datasets was to find an automatic
solution for generating alignments between files where each one contains a high number
of semantic correspondences. On the other hand, thanks to our automation of the
framework, it was easier to create English-Turkish alignments by implementing a generator
for the solution. The solution consists of (1) considering the dataset that contains a high
number of semantic correspondences (e.g. English-French), (2) replacing English
entities by the new language (e.g. Turkish) (3) replacing entities of the other language (e.g.
French) with corresponding English entities, using the alignments of the same
ontologies between English and the other language (e.g. French in this case).
3</p>
    </sec>
    <sec id="sec-2">
      <title>Experiments</title>
      <p>
        The experimental study conducted on the Turkish datasets is performed using the CroLOM
system due to its good results (ranked third) obtained in the OAEI2016 edition. CroLOM[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
uses the Yandex translator, NLP techniques and a similarity computation based on the
categories of words and synonyms. The experimental results are presented in Table 1
for each language pair. The results are good for the pairs English and Spanish. However,
they are less satisfactory for the pairs Chinese, Arabic and German. This is explained
by the fact that CroLOM uses English as pivot to align multilingual ontologies. We can
also observe that, on average Turkish ontologies bring an additional complexity to the
Multifarm track, w.r.t. the results without Turkish dataset obtained via CroLOM.
      </p>
      <p>The CroLOM system completed all the tests involving the Turkish language and
the experimental study shows that the dataset is suitable to evaluate state-of-the-art
ontology matching systems.</p>
      <p>Acknowledgements We would like to thank Lecturer Nuriye In for her contributions
to the corrections of the datasets.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>A.</given-names>
            <surname>Khiat</surname>
          </string-name>
          .
          <article-title>Crolom: cross-lingual ontology matching system results for OAEI 2016</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Ontology Matching co-located with the 15th International Semantic Web Conference (ISWC</source>
          <year>2016</year>
          ), Japan, pages
          <fpage>146</fpage>
          -
          <lpage>152</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A.</given-names>
            <surname>Khiat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Diallo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yaman</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Jime´nez-</article-title>
          <string-name>
            <surname>Ruiz</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Benaissa</surname>
          </string-name>
          .
          <article-title>Abom and adom: Arabic datasets for the ontology alignment evaluation campaign</article-title>
          .
          <source>In ODBASE 2015</source>
          , pages
          <fpage>545</fpage>
          -
          <lpage>553</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>C.</given-names>
            <surname>Meilicke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Garcia-Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Freitas</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. R. van Hage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Montiel-Ponsoda</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . R. de Azevedo,
          <string-name>
            <given-names>H.</given-names>
            <surname>Stuckenschmidt</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>Sva´b-</article-title>
          <string-name>
            <surname>Zamazal</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Sva´tek, A</article-title>
          . Tamilin, C. T. dos
          <string-name>
            <surname>Santos</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Multifarm: A benchmark for multilingual ontology matching</article-title>
          .
          <source>J. Web Sem</source>
          .,
          <volume>15</volume>
          :
          <fpage>62</fpage>
          -
          <lpage>68</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>