<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>OTTO { Ontology Translation System</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Mihael Arcan Kartik Assoja Housam Ziad Paul Buitelaar Insight Centre for Data Analytics @ NUI Galway</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>To enable knowledge access across languages, ontologies that are often represented only in English, need to be translated into di erent languages. For this reason, we present OTTO, an OnTology TranslatiOn System, which enhances ontologies with multilingual information. Rather a di erent task than the classic document translation, ontology label translation faces highly speci c vocabulary and lack contextual information. Therefore, OTTO takes advantage of the semantic information of the ontology to improve the translation of labels.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The task of ontology translation involves generating an appropriate translation
for the lexical layer, i.e. labels stored in the ontology. Most of the previous related</p>
      <sec id="sec-1-1">
        <title>1 https://translate.google.com/</title>
      </sec>
      <sec id="sec-1-2">
        <title>2 Translation performed on 25.06.2015</title>
        <p>
          work focused on accessing existing multilingual lexical resources, like
EuroWordNet or IATE [
          <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
          ]. Their work focused on the identi cation of the lexical overlap
between the ontology and the multilingual resources, which guarantees a high
precision but a low recall. Consequently, external translation services like
BabelFish, SDL FreeTranslation tool or Google Translate were used to overcome
this issue [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ]. Additionally, [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] performed ontology label
disambiguation, where the ontology structure was used to annotate the labels with their
semantic senses. Di erently to the aforementioned approaches, which rely on
external knowledge or services, we focus on how to gain adequate translations with
a domain-aware SMT system, which is supported by the ontology hierarchy.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>System Implementation</title>
      <p>
        Based on the lexical and semantic overlap with the ontology labels, the OnTology
TranslatiOn System { OTTO3 identi es, from a large set of parallel corpora, the
most relevant source sentences containing the labels to be translated. The goal
is to translate the ontology labels within the textual context of the targeted
domain, rather than in isolation. For instance, with this selection approach, we
aim to retain relevant sentences, where the English word vessel or injection
belongs to the medical domain, but not to the technical domain.
Statistical Machine Translation For the translation approach, OTTO
engages the Moses toolkit [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. To have a broader domain coverage of the generic
parallel dataset necessary for training the SMT system, we merged the JRC-Acquis
3.0 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Europarl v7 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and OpenSubtitles2013 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], thus obtaining a training
corpus of 8.5M parallel sentences for English-German, 18.9M for English-Italian
and 33.6M for the English-Spanish translation directions. To train OTTO for the
(under-resourced) English-Irish translation direction, we collected around 723K
parallel sentences from various parallel corpora, like DGT (DG Translation at
the European Commission), EUbookshop or KDE4, from the OPUS webpage.4
Relevant Sentence Selection In order to improve the translation of ontology
labels, we select from the concatenated corpus only those source sentences, which
are most relevant to the labels to be translated. The rst criterion for relevance
is the n-gram overlap between a label and a source sentence coming from the
generic corpus. Due to the speci city of the ontology labels, just an n-gram
overlap approach is not su cient to select all the useful sentences. For this
reason, we follow the idea of extending the semantic information of the labels
using Word2Vec5 for computing distributed representations of words [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
technique is based on a neural network that analyses the textual data provided
as input, in our experiment ontology labels and source sentences, and outputs
a list of semantically related words [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Each input string is vectorized and
compared to other vectorized sets of words in a multi-dimensional vector space,
which was trained with Word2Vec on the Wikipedia articles.6
      </p>
      <p>To further improve the disambiguation of short labels, the related words of
the label are concatenated with the related words of its direct parent in the
ontology hierarchy. Given a label and a source sentence from the generic corpus,
3 http://server1.nlp.insight-centre.org/otto/
5 https://code.google.com/p/word2vec/</p>
      <sec id="sec-2-1">
        <title>4 http://opus.ling l.uu.se/</title>
      </sec>
      <sec id="sec-2-2">
        <title>6 Wikipedia dump id enwiki-20141106</title>
        <p>related words and their weights are extracted from both of them, and used as
entries of the vectors to calculate the cosine similarity. Finally, the most similar
source sentence and the label should share the largest number of related words.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 OTTO Demo</title>
      <p>OTTO takes as input an ontology represented in OWL or RDF and extracts
the labels stored in it. To improve the translations of the labels stored in the
ontology, the most relevant sentences, which contain the labels, are obtained from
the concatenated generic corpus. Once the labels are identi ed in the context of
the relevant sentences, OTTO engages Moses and translates the labels within the
context into German, Italian, Spanish and Irish. After the translation process is
done, the translated labels are identi ed in the relevant target sentences.</p>
      <p>Since the translation of the extracted labels may take several minutes (or
even hours), the OTTO user can provide an optional e-mail,7 which is used to
inform the user about the completion of the translation process, as well as the
web address where the provided data will be stored. Without this information,
the address of the stored data is given trough the OTTO interface (Figure 1).</p>
      <p>
        In the last step, the translated labels are represented in di erent formats,
for example in a HTML table and CSV le to allow a better visualisation.
Furthermore, the multilingual information is injected into the original monolingual
ontology and represented as a multilingual ontology as well as in lemon8 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], a
model for linking linguistic information with ontologies.
      </p>
    </sec>
    <sec id="sec-4">
      <title>5 Conclusion</title>
      <p>This paper is aimed at showing OTTO, an OnTology TranslatiOn System for
multilingual enrichment of semantically structured data, i.e. ontologies or
taxonomies. The system is based on an approach to identify the most relevant
source sentences from a large generic parallel corpus, giving the possibility to
automatically translate highly speci c ontology labels in context without
particular in-domain parallel data. The demonstrated approach reduces the ambiguity
of expressions in the selected sentences, which consequently generates better</p>
      <sec id="sec-4-1">
        <title>7 The provided e-mail is stored as a variable and is deleted after the process nishes.</title>
      </sec>
      <sec id="sec-4-2">
        <title>8 http://lemon-model.net/</title>
        <p>translations of ontology labels. As an ongoing work, we further focus on
improving the extraction of the lexical knowledge stored in ontologies. Additionally, we
plan to enable knowledge enrichment for existing multilingual ontologies.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>This publication has emanated from research conducted with the nancial
support of Science Foundation Ireland (SFI) under Grant Number SFI/12/RC/2289.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arcan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turchi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Knowledge portability with semantic expansion of ontology labels</article-title>
          .
          <source>In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing. ACL</source>
          , Beijing, China (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Espinoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A note on ontology localization</article-title>
          .
          <source>Appl. Ontol</source>
          .
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <volume>127</volume>
          {137 (Apr
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vela</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gantner</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manzano</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , D-Saarbru
          <article-title>cken: Multilingual lexical semantic resources for ontology translation</article-title>
          .
          <source>In: In Proceedings of the 5th International Conference on Language Resources and Evaluation</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Espinoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ontology localization</article-title>
          .
          <source>In: Proceedings of the Fifth International Conference on Knowledge Capture. K-CAP '09</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brennan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>Cross-lingual ontology mapping - an investigation of the impact of machine translation</article-title>
          . In:
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>Y</given-names>
          </string-name>
          . (eds.)
          <source>ASWC. Lecture Notes in Computer Science</source>
          , vol.
          <volume>5926</volume>
          . Springer (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vila-Suero</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguado-de Cea</surname>
          </string-name>
          , G.:
          <article-title>Guidelines for multilingual linked data</article-title>
          .
          <source>In: Proceedings of the 3rd International Conference on Web Intelligence</source>
          , Mining and
          <string-name>
            <surname>Semantics. ACM</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buitelaar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Challenges for the multilingual web of data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>11</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Koehn</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Europarl: A Parallel Corpus for Statistical Machine Translation</article-title>
          .
          <source>In: Conference Proceedings: the tenth Machine Translation Summit. AAMT</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Koehn</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callison-Burch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Federico</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertoldi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cowan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moran</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zens</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojar</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Constantin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herbst</surname>
          </string-name>
          , E.: Moses:
          <article-title>Open source toolkit for statistical machine translation</article-title>
          .
          <source>In: Proceedings of the 45th Annual Meeting of the ACL on Interactive Poster and Demonstration Sessions. Stroudsburg</source>
          , PA, USA (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Espinoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montiel-Ponsoda</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguado-de Cea</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Combining statistical and semantic approaches to the translation of ontologies and taxonomies</article-title>
          . In: Fifth workshop on Syntax,
          <source>Structure and Semantics in Statistical Translation (SSST-5)</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spohr</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Linking lexical resources and ontologies on the semantic web with lemon</article-title>
          .
          <source>The Semantic Web: Research and Applications</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>CoRR abs/1301</source>
          .3781 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Steinberger</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouliquen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Widiger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ignat</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erjavec</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tu s</surname>
          </string-name>
          , D.,
          <string-name>
            <surname>Varga</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The JRC-Acquis: A multilingual aligned parallel corpus with 20+ languages</article-title>
          .
          <source>In: Proceedings of the 5th International Conference on Language Resources and Evaluation</source>
          (LREC'
          <year>2006</year>
          ) (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tiedemann</surname>
          </string-name>
          , J.:
          <article-title>Parallel data, tools and interfaces in opus</article-title>
          . In: Chair),
          <string-name>
            <given-names>N.C.C.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Dogan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.U.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the Eight International Conference on Language Resources and Evaluation</source>
          . Istanbul, Turkey (may
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>