<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards the Integration of Multilingual Terminologies: an Example of a Linked Data Prototype</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Montiel-Ponsoda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julia Bosque-Gil</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Gracia</string-name>
          <email>jgracia@fi.upm.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guadalupe Aguado-de-Cea</string-name>
          <email>lupe@fi.upm.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Vila-Suero</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ontology Engineering Group, Universidad Politécnica de Madrid</institution>
          ,
          <addr-line>Spain Campus de Montegancedo sn, Boadilla del Monte 28660 Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>205</fpage>
      <lpage>206</lpage>
      <abstract>
        <p>Many language resources are nowadays available in machine readable formats, but still contained in isolated silos. Current Semantic Web-based techniques enable the transformation and linking of those resources to become a navigable graph of linked language resources, which can be directly consumed by third-party applications. The prototype we have developed builds on a web user interface and SPARQL endpoint initially developed to query a single terminological database (Terminesp), now extended to navigate a set of multilingual terminologies. The vocabulary used to represent these terminologies into the linked data format is lemon-ontolex, a de facto standard for representing lexical information relative to ontologies and for linking lexicons and machine-readable dictionaries to the Semantic Web.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The Linguistic Linked Open Data (LLOD) cloud1
is a sub-cloud of linguistic resources provided in an
interoperable way (using the Resource Description
Framework or RDF data model), freely accessible
and linked with each other. In its current state, the
LLOD Cloud contains monolingual and multilingual
dictionaries, lexicons, thesauri and even corpora.
English is the best represented language, and some
languages are underrepresented or not present at all.</p>
      <p>
        With Terminesp (a multilingual terminological
database created by the Spanish Association for
Terminology, AETER), we aimed at validating the
lemon-ontolex model as a representation scheme for
1 http://linguistic-lod.org/
lexical resources, specifically, the so-called
vartrans module, a dedicated module that
accounts for terminological variation and translation
relations among entries
        <xref ref-type="bibr" rid="ref1">(Bosque-Gil et al., 2015)</xref>
        .
Building on that experience, we have now
transformed additional multilingual terminological
resources, namely, a set of freely available
terminology databases2 from the Catalan
Terminological Centre, TERMCAT, into linked data
(LD) using lemon-ontolex as underlying data format,
and aim to showcase the benefits of integrating
terminological resources.
      </p>
      <p>In this paper, we focus on the design decisions
taken in the transformation and linking steps, and on
the impact they have in the search and navigation of
the resulting linked terminological data.</p>
      <p>In Section 2, we introduce lemon-ontolex and the
vartrans module. In section 3, we describe the
design decisions taken in the transformation process.
In section 4, we refer to the benefits of browsing and
navigating linked multilingual terminologies.
2.</p>
      <p>lemon-ontolex</p>
      <p>
        The lemon-ontolex model is the resulting work of
the efforts made by the W3C Ontology Lexica
Community Group since 2011 to build a rich model
to represent the lexicon-ontology interface. It is
largely based on the lemon model
        <xref ref-type="bibr" rid="ref2">(McCrae et al.,
2012)</xref>
        and consists of a core set of classes and
several modules3. The vartrans module has been
developed to record lexico-semantic relations across
entries in the same or different languages (Fig. 1.):
those among senses and those among lexical entries
and/or forms. Lexico-semantic relations among
senses are of semantic nature and include
2 http://www.termcat.cat/es/terminologiaoberta/
3 See lemon-ontolex final model specifications at
http://www.w3.org/community/ontolex/wiki/Final_Model_Sp
ecification
terminological relations (dialectal, register,
chronological, discursive, and dimensional variation)
and translation relations. In contrast, relations among
lexical entries and/or forms concern the surface form
of a term and encode morphological and
orthographical variation, among other aspects.
      </p>
      <p>
        Migration and linking of the resources
For the transformation of TERMCAT terminology
repertoires to the LD format and linking to
Terminesp we followed these steps: data exploration,
URI naming strategy, data modeling, RDF
generation and linking
        <xref ref-type="bibr" rid="ref3">(Vila-Suero et al., 2014)</xref>
        .
      </p>
      <p>Data exploration. TERMCAT terminology
repertoires are divided by domain. Each database
consists of a list of entries in Catalan and their
translations into Spanish, English, French, etc., along
with the term type (full form or abbreviation),
references to associated terms, synonyms, and,
sometimes, definitions. Data for part-of-speech,
gender and number in nouns, and subcategorization
of verbs, is also available.</p>
      <p>URI naming strategy. Inspired by the work in
the Apertium dictionaries4, the term itself, its part of
speech and the language of the term are part of the
URI of the lexical entry. For lexical senses, the
domain is included in the URI.</p>
      <p>Modeling. For the modeling process, we regard
each term in a set of translations as a specific sense
of a lexical entry, a sense that is mapped to a concept
in a particular domain. This allows us to have a
unique lexical entry red (network), for instance,
which occurs both in the lexicon Internet i societat
de la informació as in the lexicon Indústria
electrònica i dels materials elèctrics, with different
senses that we extract from each domain lexicon.
This results in a number of RDF lexica that matches
the number of languages available in TERMCAT
data, and each lexical entry will have a different
number of senses depending on its use across
domains. In this way, the lexical entry :red-n-es will
be mapped to a sense :red-n-es-Internet-sense, as
well as to a :red-n-es-Industria-sense, etc. Each of
these senses refers to a skos:Concept with a
particular definition and domain. Regarding
translations, the vartrans module represents them
as relations across lexical senses of the entries of
each lexicon. Parts of speech, subcategorization,
gender and number are accounted for as well.</p>
      <p>Generation and linking. For the transformation
we used the data cleaning and transformation tool
OpenRefine5 with its extension for LD. We linked to
lexinfo6 to cover morphosyntactic information, and
to Terminesp at the lexical entry level. Linking to
DBpedia is also planned as a next step.
4.</p>
      <p>Browsing multilingual terminologies</p>
      <p>We reuse the Terminesp web user interface (see
Fig. 2.) and SPARQL endpoint to browse and query
this set of integrated terminologies7. Benefits are
related to easy access and reuse of linguistic data by
end users (translators, terminologists) and
semanticaware software agents.</p>
      <p>Acknowledgements. This work is supported by
the FP7 EU project LIDER (610782), and the
Spanish 4V project (TIN2013-46238-C4-2-R).
5 http://openrefine.org/index.html
6 http://lexinfo.net/
7 http://linguistic.linkeddata.es/terminesp/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Bosque-Gil</surname>
          </string-name>
          et al. (
          <year>2015</year>
          ).
          <article-title>Applying the OntoLex Model to a Multilingual Terminological Resource</article-title>
          .
          <source>In Proc. of ESWC 2015</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>J. McCrae</surname>
          </string-name>
          et al. (
          <year>2012</year>
          ).
          <article-title>Interchanging lexical resources on the semantic web</article-title>
          .
          <source>Language Resources and Evaluation</source>
          , vol.
          <volume>46</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Vila-Suero</surname>
          </string-name>
          et al. (
          <year>2014</year>
          ).
          <article-title>Publishing Linked Data on the Web: the Multilingual Dimension</article-title>
          . In P. Cimiano &amp; P. Buitelaar (Eds.) Towards the Multilingual Semantic Web. Springer.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>