<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>One ontology to bind them all: The META-SHARE OWL ontology for the interoperability of linguistic datasets on the Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>John P. McCrae</string-name>
          <email>jmccraeg@cit-ec.uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Penny Labropoulou</string-name>
          <email>penny@ilsp.athena-innovation.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jorge Gracia</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marta Villegas</string-name>
          <email>marta.villegas@upf.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V ctor Rodr guez-Doncel</string-name>
          <email>vrodriguezg@fi.upm.es</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Cimiano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cognitive Interaction Technology</institution>
          ,
          <addr-line>Excellence Cluster</addr-line>
          ,
          <institution>Bielefeld University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ILSP/\Athena" RC</institution>
          ,
          <addr-line>Athens</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ontology Engineering Group, Universidad Politecnica de Madrid</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University Pompeu Fabra</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>META-SHARE is an infrastructure for sharing Language Resources (LRs) where signi cant e ort has been made into providing carefully curated metadata about LRs. However, in the face of the ood of data that is used in computational linguistics, a manual approach cannot su ce. We present the development of the META-SHARE ontology, which transforms the metadata schema used by META-SHARE into an open world ontology that can better handle the diversity of metadata found in legacy and crowd-sourced resources. We show how this model can interface with other more general purpose vocabularies for online datasets and licensing, and apply this model to the CLARIN VLO, a large source of legacy metadata about LRs. Furthermore, we demonstrate the usefulness of this approach in two public metadata portals for information about language resources.</p>
      </abstract>
      <kwd-group>
        <kwd>language resources and evaluation</kwd>
        <kwd>metadata</kwd>
        <kwd>ontologies</kwd>
        <kwd>harmonization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The study of language and the development of natural language processing
(NLP) applications requires access to language resources (LRs). Recently,
several digital repositories that index metadata for LRs have emerged, supporting
the discovery and reuse of LRs. One of the most notable of such initiatives is
META-SHARE5 [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], an open, integrated, secure and interoperable exchange
infrastructure where LRs are documented, uploaded, stored, catalogued,
announced, downloaded, exchanged and discussed, aiming to support reuse of LRs.
5 http://www.meta-share.eu
Towards this end, META-SHARE has developed a rich metadata schema that
allows aspects of LRs accounting for their whole lifecycle from their production
to their usage to be described. The schema has been implemented as an XML
Schema De nition (XSD) 6 and descriptions of speci c LRs are available as XML
documents.
      </p>
      <p>
        Yet, META-SHARE is not the only source for discovering LRs and their
descriptions; other sources include the catalogs of agencies dedicated to LRs
promotion and distribution, such as ELRA7 and LDC8, other infrastructures such
as the CLARIN Virtual Language Observatory (VLO)9 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the Language Grid10
and Alveo11, the Open Language Archives Community (OLAC)12, catalogs with
crowd-sourced metadata, such as the LREMap13 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and, more recently,
repositories coming from various communities (e.g. OpenAire14, EUDAT15 etc.). The
metadata schemes of all these sources vary with respect to their coverage and
the set of speci c metadata captured. Currently, it is not possible to query all
these sources in an integrated and uniform fashion. The Web of Data is a natural
scenario for exposing LRs metadata in order to allow their automated discovery,
share and reuse by humans or software agents and the bene ts of this model
including interoperability, federation, expressivity and dynamicity were laid out
by Chiarcos et al.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        In this paper we contribute to the interoperability of these repositories by
developing an ontology in the Web Ontology Language (OWL) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] that allows
us to represent the metadata schemes of these repositories under an extensible,
open-world model.16 The proposed ontology is based on the ontology
developed by Villegas et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] for the University Pompeu Fabra's (UPF)
METASHARE node (covering part of the original schema), which is extended to the
complete schema (in order to cover all relevant LRs) and incorporates the
consensus reached in the context of the W3C Linked Data for Language
Technologies (LD4LT) Community Group17. We show how this model interacts with the
DCAT [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] vocabulary as well as the most prominent models in the CLARIN
VLO data. Further, we describe the application of the model in two portals,
rstly the IULA LOD catalogue and secondly Linghub18.
      </p>
      <p>The remainder of this paper is structured as follows: in Section 2 we will
describe the related work in the elds of LR metadata harmonization. The
de6 https://github.com/metashare/META-SHARE/tree/master/misc/schema/v3.0
7 http://www.elra.info/en
8 https://www.ldc.upenn.edu/
9 http://catalog.clarin.eu/vlo/?1
10 http://langrid.org/en/index.html
11 http://alveo.edu.au
12 http://www.language-archives.org
13 http://www.resourcebook.eu/searchll.php
14 https://www.openaire.eu/
15 http://eudat.eu
16 http://purl.org/net/def/metashare
17 https://www.w3.org/community/ld4lt
18 http://linghub.org/
velopment of the META-SHARE ontology is described in section 3 and its
application in Section 4. Finally, in section 5 we consider the broader impact of
this ontology as a tool for computational linguists and as a method to realize an
architecture of (linked) data-aware services.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The task of nding common vocabularies for linguistics is of wide interest and
several general ontologies for linguistics have been proposed. The General
Ontology for Linguistic Description [9, GOLD] was proposed as a common model
for linguistic data, but its relatively limited scope and low coherence has not led
to wide-spread adoption. An alternative approach that has been proposed is to
use ontologies to create coherence among the resources, in particular by using
ontologies to align di erent linguistic schemas [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        This lack of consensus resides also in the description of LRs, even for
nonlinguistic concepts. In fact, there are as many metadata schemas for their
descriptions as catalogs and repositories for their presentation (e.g. those used
by ELRA and the LDC) and communities describing them (e.g. TEI [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] or
CES [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). The most widely used schema for the exchange of LRs is the one
suggested by the Open Language Archives Community [1, OLAC], which builds
on the Dublin Core metadata and has been criticized as too reductionistic.
Differences between the schemas lie in the range of features used and their labels
and datatypes.
      </p>
      <p>
        An important e ort to harmonize metadata has been the ISO Data Category
Registry (ISOcat DCR) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], intended as a registry where metadata providers
could register their concepts (Data Categories) and link them to those of other
providers. A subset thereof were selected by metadata experts as the core
elements for the description of LRs (\Athens Core"). The Component Metadata
Infrastructure [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed by CLARIN extends this principle of a common
registry to include \components" and \pro les": \components" consist of
semantically close elements to be shared among di erent communities when producing
\pro les" for speci c LR types. However, as we observe in section 3.5, this has
in practice merely resulted in each contributing institute using its own scheme,
with very little commonality between di erent institutes. To improve this
situation it was recently proposed that the conversion of these CMDI schemas to
RDF would enable better interoperability [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        A di erent approach was taken for the design of the META-SHARE schema
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which was based on a comparative study of the most widespread metadata
schemas and catalog descriptions, analysis of user needs and discussions with
metadata providers and experts in order to arrive at a common schema, taking
into account previous initiatives and recommendations (cf. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]).
      </p>
    </sec>
    <sec id="sec-3">
      <title>The META-SHARE OWL Ontology</title>
      <sec id="sec-3-1">
        <title>Original MS XSD schema</title>
        <p>
          The META-SHARE schema [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] has been designed not only as an aid for LR
search and retrieval, but also as a means to foster their production, use and
reuse by bringing together knowledge about LRs and related objects and processes,
thus encoding information about the whole lifecycle of the LR from production
to usage. The central entity of the META-SHARE schema is the LR per se,
which encompasses both data sets (e.g., textual, audio and
multimodal/multimedia corpora, lexical data, ontologies, terminologies, computational grammars,
language models) and technologies (e.g., tools, services) used for their
processing. In addition to the central entity, other entities are also documented in
the schema; these are reference documents related to the LR (papers, reports,
manuals etc.), persons/organizations involved in its creation and use (creators,
distributors etc.), related projects and activities (funding projects, activities of
usage etc.), accompanying licenses, etc., all described with metadata taken as far
as possible from relevant schemas and guidelines (e.g. BibTex for bibliographical
references). The META-SHARE schema proposes a set of elements to encode
speci c descriptive features of each of these entities and relations holding
between them, taking as a starting point the LR. Following the CMDI approach,
these elements are grouped together into \components". The core of the schema
is the resourceInfo component (Figure 1), which subsumes
{ administrative components relevant to all LRs, e.g. identificationInfo
(name, description and identi ers), distributionInfo (licensing and
intellectual property rights information), usageInfo (information about the
intended and actual use of the LR).
{ components speci c to the resource type (corpus, lexical/conceptual
resource, language model, tool/service) and media type (text, audio, video,
image), which support the encoding of information relevant to
resource/media combinations, e.g. text or audio parts of corpora, lexical/conceptual
resources etc., such as language, formats, classi cation.
        </p>
        <p>The META-SHARE schema recognises obligatory elements (minimal version)
and recommended and optional elements (maximal version). An integrated
environment supports the description of LRs, either from scratch or through
uploading of XML les adhering to the META-SHARE metadata schema, as well
as browsing, searching and viewing of the LRs.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Formal modelling and mapping issues</title>
        <p>
          In the META-SHARE XSD schema, elements are formalized as simple
elements whereas components are formalized as complex-type elements. When
mapping the XSD schema to RDF, elements can be naturally understood as
properties (e.g. name, gender, etc.). Components (i.e. complex-type elements),
however, deserve a careful analysis. General mapping rules from XSD to RDF
establish that a local element with complex type translates into an object
property and a class. We observed that the straightforward application of such a
principle may derive unnecessarily verbose graphs. Thus, following Villegas et
al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ], we identi ed potentially removable nodes before undertaking the actual
RDFication process. Embedded complex elements with cardinality of exactly
one are identi ed as potentially removable, provided they contain neither text
nor attributes. This allows for a simpli cation of the model, for example in
the chain resourceInfo identificationInfo resourceName, the
identificationInfo property is not needed. Interestingly enough, the
removal of the super uous wrapping elements has also led to a change of philosophy
in the schema and a need for restructuring in order to ensure that properties
are attached to the most appropriate node, as exempli ed and discussed in
Section 3.4. Beyond this, we made extensions to our mapping strategy in order to
improve the ontology, such as the following:
{ Removal of the Info su x from the names of wrapping elements of
components.
{ Improvement of names that created confusion, as already noted by the
META-SHARE group and/or the LD4LT group; thus, resourceInfo was
renamed LanguageResource, restrictionsOfUse became conditionsOfUse.
{ Generalization of concepts, e.g. notAvailableThroughMetashare with
availableThroughOtherDistributor;
{ Development of novel classes based on existing values, for example:
        </p>
        <p>Corpus 9resourceType:corpus
{ Grouping similar elements under novel superclasses, e.g. annotationType
and genre values are structured in classes and subclasses better re ecting
the relation between them. Indicatively, the superclass SemanticAnnotation
can be used to bring together semantic annotation types, such as semantic
roles, named entities, polarity, and semantic relations.
{ Extension of existing classes with new values and new properties (e.g. licenseCategory
for licences).</p>
        <p>The actual mapping was achieved by means of a custom domain-speci c
language inspired by XML Stylesheet Transforms.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Interface with DCAT and other vocabularies</title>
        <p>The META-SHARE model can be considered broadly similar to DCAT in that
there are classes that are nearly an exact match to the ones in DCAT for
three out of four classes. DCAT's dataset corresponds nearly exactly to the
resourceInfo tag and similarly, distributions are similar to distributionInfo
classes and catalogRecord is similar to metadataInfo. Thus, we
introduced equivalent class relations between these elements. The fourth main class,
catalog covers a level not modelled by META-SHARE. DCAT uses Dublin
Core properties for many parts of the metadata, and often these properties are
found deeply nested in the META-SHARE description. For example, language
is found in several places deeply nested under six tags19. In META-SHARE this
allows di erent media types in the resource to have di erent languages, e.g., the
dialogues and the scripts of a video may be in English, whereas the subtitles can
be in French and German. We still include this ne-grained metadata but also
add the property at the resource level to indicate if any part of the resource is in
19 e.g., resourceInfo resourceComponentType corpus corpusMediaType
corpusVideoInfo languageInfo ! dc:language
&lt;&gt; a dcat:Dataset ;
dc:description "Cette base de donnes..."@fr
dc:language "tur" ;
dc:source "META-SHARE" ;
ms:corpusInfo &lt;#corpusInfo&gt; ;
ms:distributionInfo &lt;#distributionInfo&gt; ;
rdfs:seeAlso &lt;http://metashare.elda.org/reposit...ac770/&gt; .
&lt;#corpusInfo&gt; a ms:CorpusInfo , ms:CorpusAudioInfo ;
dc:language "tur" ;
ms:languageName "Turkish" ;
ms:mediaType ms:audio .
the stated language. Similarly, it is also the case that some Dublin Core
properties are not directly speci ed in the META-SHARE model, but can be inferred
from related properties, e.g., Dublin Core's `contributor' follows (by means of a
property chain) from people indicated as `annotators', `evaluators', `recorders' or
`validators'. Furthermore, several DCAT speci c-properties, such as `download
URL', are nearly exactly equivalent to those in Metashare but occur in places
that do not t the domain and range of the properties. In this particular case,
it was a simple x to move the property to the enclosing DistributionInfo
class. Inevitably, several properties from DCAT did not have equivalences in
META-SHARE, notably `keyword'.</p>
        <p>In addition to DCAT, we used also other vocabularies to establish
equivalences to parts of the model. In particular, we mapped to the Friend of a Friend
(FOAF) ontology to describe people and organizations and the Semantic Web
for Research Communities (SWRC) ontology to describe scienti c publications.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Licensing module</title>
        <p>
          A speci c area where we made a signi cant e ort to improve the modelling was in
the licensing information in order to allow the formulation of a clear and concise
rights information of the LRs. Some languages already exist for this purpose,
and among them, ODRL 2.1 (Open Digital Rights Language) was chosen and
extended, which is a policy and rights expression language speci ed by the W3C
ODRL Community Group20 which de nes a model for representing permissions,
prohibitions and duties. The most common licenses (for software, data or general
works) have been already expressed in ODRL in the RDF License dataset [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]
and can be pointed to when an LR is licensed with any of these. Extensions to the
20 https://www.w3.org/community/odrl/
ODRL vocabulary have been made to represent some of the speci cities of the LR
domain. The speci cation also suggested changes, some of them structural, to the
previous META-SHARE modelling, and to this extent we combined the existing
META-SHARE licensing vocabulary with ODRL. In addition, we extended the
model by adding some new properties and individuals based on requirements
from the LD4LT community group21. In particular, the generic
conditions-ofuse values of the META-SHARE schema have been exploited for creating RDF
codes for non-standard licenses and are mapped to ODRL actions (e.g. the duty
to attribute), and included in an RDF document, as shown in gure 4. This
module has been published both as independent module22 and as part of the
META-SHARE ontology.
&lt;#distributionInfo&gt; a ms:DistributionInfo , dcat:Distribution ;
dct:license &lt;http://purl.org/NET/rdflicense/ms-c-nored-ff&gt; ;
dcat:accessURL &lt;http://catalog.elra...&gt; .
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Harmonizing other resources with META-SHARE</title>
        <p>
          While a basic level of interoperability can be established by using standard
vocabularies such as DCAT and Dublin Core, this can only be done by sacri cing
completeness and ignoring all metadata particular to language resources. For this
reason, we use the META-SHARE model to represent and harmonize the
metadata relating speci cally to the domain of linguistics and language resources. As
a proof-of-concept, we show how the META-SHARE ontology supports the
harmonization of data from the CLARIN VLO. The CLARIN repository describes
its resources using a small common set of metadata and a larger description
de ned by the Component Metadata Infrastructure [4, CMDI]. These metadata
schemes are extremely diverse as shown in Table 1. We will focus on the top ve
of these types for which we have created corresponding mappings. Two of these
schemes are only Dublin Core properties and so do not have speci c language
resource metadata. The most frequent tag 'Song' tag is used to describe records
of a database consisting of musical recordings. While many of the properties used
by this tag (e.g., `number of stanzas') have no correspondence in Dublin Core,
they can be described with respect to existing elements of the META-SHARE
Ontology. The Session tag is used to provide IMDI metadata [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and has a very
loose correspondence to META-SHARE: For instance, there are no
corresponding properties to describe the participants of a media recording. This highlights
the advantage of taking an open world, ontological approach as opposed to a
21 https://www.w3.org/community/ld4lt/wiki/Metashare vocabulary for licenses
22 http://purl.org/NET/ms-rights
        </p>
        <p>
          xed schema, in that we can easily introduce new properties while still reusing
the META-SHARE properties where they are appropriate. The MODS metadata
scheme [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] is in fact a general domain metadata framework. We found that 28
entities from META-SHARE corresponded to elements used in the MI `Sing'
metadata, and 37 to the IMDI metadata, although there was only minor overlap
with the MODS scheme (in particular 4 entities used to describe language) as
this scheme is not speci c to language resources.
        </p>
        <p>Component Root Tag Institutes Frequency
Song 1 (MI) 155,403
Session 1 (MPI) 128,673
OLAC-DcmiTerms 39 95,370
MODS 1 (Utrecht) 64,632
DcmiTerms 2 (BeG,HI) 46,160
SongScan 1 (MI) 28,448
media-session-pro le 1 (Munich) 22,405
SourceScan 1 (MI) 21,256
Source 1 (MI) 16,519
teiHeader 2 (BBAW, Copenhagen) 15,998</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Applications</title>
      <sec id="sec-4-1">
        <title>IULA LOD Catalogue</title>
        <p>
          The IULA-UPF CLARIN Competence Centre23 aims to promote and support
the use of technology and text analysis tools in the Humanities and Social
Sciences research. The centre includes a Catalogue24 with information on language
resources and technology. The Catalogue is based on the initial linked open data
(LOD) version of the META-SHARE model as described in [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and the original
data originate from the UPF META-SHARE node25. The source XML records
were converted into RDF and augmented with service descriptions (not included
in the UPF META-SHARE node) and relevant documentation (appropriate
articles, documentation, sample data and results, illustrative experiments,
examples from outstanding projects, illustrative use cases, etc) to encourage potential
users to embrace digital tools. Finally, the data was enriched with internal and
23 http://www.clarin-es-lab.org/index-en.html
24 http://lod.iula.upf.edu/
25 http://metashare.upf.edu
external links. The resulting linked data maximised the information contained
in the original repository and enabled data mashup techniques that get relevant
data from the DBpedia and the DBLP26. The catalogue demonstrates the
bene ts of the LOD framework and how this can be easily used as the basis for a
web browser application that maximizes information and helps users to navigate
throughout the datasets in a comprehensive way.
Linghub is a portal designed to allow common querying of metadata from
multiple highly heterogeneous repositories. Currently, it draws not only from
METASHARE, but also from the LRE-Map [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], the CLARIN VLO [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and DataHub
and is regularly updated with new/changed information. The repository
currently bases itself mostly on the DCAT and Dublin Core vocabularies, however
these models do not capture any speci c linguistic information. For this reason,
the ontology described in this paper is currently being integrated into the
system to allow users to use META-SHARE as the basic vocabulary for querying
linguistic information about language resources, and the mappings previously
described have already been applied to data from LRE-Map and the CLARIN VLO.
Linghub supports browsing and querying by several means, including faceted
browsing, full-text search, SPARQL querying and related item search. As such,
we believe that the portal, while not a direct collector of metadata, will enable
users to nd more language resources and do so more easily. The Linghub portal
is thus a proof-of-concept for the level of harmonization that the use of a
common ontology provides, as metadata originating from di erent repositories can
be uniformly queried in Linghub in an integrated fashion. We adhere to an open
architecture in which not only Linghub but other discovery services aggregate
and index data could potentially be developed.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>
        This work represents only a rst starting point for the harmonization of language
resources by providing a standard ontology that can be used in the description
of metadata of linguistic resources and there are still a number of challenges
ahead of us to be addressed. Firstly, the next step would be to make sure that
not only metadata, but the actual data is available on the Web in open web
standards such as RDF so that data can be automatically crawled and analyzed.
Secondly, it should be required that linguistic data published on the Web should
ideally follow the same format (e.g. RDF) so that it can be easily integrated
and data can be queried across datasets. This presupposes the agreement on
best practices for data publication and formats and the Natural Language
Processing Interchange Format (NIF)[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is an obvious candidate for that. Thirdly,
harmonization should be extended to the description of NLP services so that
26 http://dblp.uni-trier.de/db/index.html
NLP services can be distributed across providers and repositories. The
mechanisms for description of the functionality of NLP services should be extremely
light-weight. Finally, input and output formats for services should be
standardized and homogenized so that services can be easily composed to realize more
complex work ows, without relying on too much parametrization. Work ows
of services should be easily executable `on the cloud'. In order to scale, services
should support parallelization, streaming and non-centralized processing. We
believe that the development of common vocabularies such as the one presented in
paper should enable the emergence of a new paradigm supporting the discovery
and exploitation of linguistic data and services across repositories.
Acknowledgments. We are very grateful to the members of the W3C Linked
Data for Language Technologies (LD4LT) for all the useful feedback received
and for allowing this initiative to be developed as an activity of the group. This
work is supported by the FP7 European project LIDER (610782), by the Spanish
Ministry of Economy and Competitiveness (project TIN2013-46238-C4-2-R and
a Juan de la Cierva grant), the Greek CLARIN Attiki project (MIS 441451) and
the H2020 project CRACKER (645357).
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simons</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The OLAC metadata set and controlled vocabularies</article-title>
          .
          <source>In: Proceedings of the ACL 2001 Workshop on Sharing Tools and Resources-</source>
          Volume
          <volume>15</volume>
          . pp.
          <volume>7</volume>
          {
          <fpage>18</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Broeder</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kemps-Snijders</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Uytvanck</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Windhouwer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Withers</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zinn</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A data category registry-and component-based metadata framework</article-title>
          .
          <source>In: Proceedings of the Seventh Conference on International Language Resources and Evaluation</source>
          . pp.
          <volume>43</volume>
          {
          <issue>47</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Broeder</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , O enga, F.,
          <string-name>
            <surname>Willems</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The IMDI metadata set, its tools and accessible linguistic databases</article-title>
          .
          <source>In: Proceedings of the IRCS Workshop on Linguistic Databases</source>
          . pp.
          <volume>11</volume>
          {
          <issue>13</issue>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Broeder</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Windhouwer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Uytvanck</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goosen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trippel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>CMDI: a component metadata infrastructure</article-title>
          . In:
          <article-title>Describing LRs with metadata: towards exibility and interoperability in the documentation of LR</article-title>
          . pp.
          <volume>1</volume>
          {
          <issue>4</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Gratta</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francopoulo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubino</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russo</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soria</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The LRE Map</article-title>
          .
          <article-title>Harmonising community descriptions of resources</article-title>
          .
          <source>In: Proceedings of the Eighth Conference on International Language Resources and Evaluation</source>
          . pp.
          <volume>1084</volume>
          {
          <issue>1089</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chiarcos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Ontologies of linguistic annotation: Survey and perspectives</article-title>
          .
          <source>In: Proceedings of the Eighth International Conference on Language Resources and Evaluation</source>
          . pp.
          <volume>303</volume>
          {
          <issue>310</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chiarcos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCrae</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Towards open data for linguistics: Lexical Linked Data</article-title>
          , pp.
          <volume>7</volume>
          {
          <fpage>25</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cieri</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langendoen</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pustejovsky</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A road map for interoperable language resource metadata</article-title>
          . In: Chair),
          <string-name>
            <given-names>N.C.C.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Rosner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Tapias</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Valletta, Malta (may
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Farrar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langendoen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>A common ontology for linguistic concepts</article-title>
          .
          <source>In: Proceedings of the Knowledge Technologies Conference</source>
          . pp.
          <volume>10</volume>
          {
          <issue>13</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gartner</surname>
          </string-name>
          , R.: MODS:
          <article-title>Metadata object description schema</article-title>
          .
          <source>JISC Techwatch report TSW</source>
          pp.
          <volume>3</volume>
          {
          <issue>6</issue>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Gavrilidou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labropoulou</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Desipri</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papageorgiou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monachini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frontini</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Francopoulo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arranz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mapelli</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The META-SHARE metadata schema for the description of language resources</article-title>
          .
          <source>In: Proceedings of the Eighth International Conference on Language Resources and Evaluation</source>
          . pp.
          <volume>1090</volume>
          {
          <issue>1097</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Brummer, M.:
          <article-title>Integrating NLP using linked data</article-title>
          .
          <source>In: Proceedings of the 12th International Semantic Web Conference</source>
          , pp.
          <volume>98</volume>
          {
          <fpage>113</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ide</surname>
          </string-name>
          , N.:
          <article-title>Corpus encoding standard: SGML guidelines for encoding linguistic corpora</article-title>
          .
          <source>In: Proceedings of the First International Language Resources and Evaluation Conference</source>
          . pp.
          <volume>463</volume>
          {
          <issue>470</issue>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veronis</surname>
          </string-name>
          , J.:
          <article-title>Text encoding initiative: Background and contexts</article-title>
          , vol.
          <volume>29</volume>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Kemps-Snijders</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Windhouwer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wittenburg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>ISOcat: Corralling data categories in the wild</article-title>
          .
          <source>In: Proceedings of the Seventh Conference on International Language Resources and Evaluation</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Maali</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erickson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Archer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Data catalog vocabulary (DCAT)</article-title>
          .
          <source>W3C recommendation, The World Wide Web Consortium</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Motik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel-Schneider</surname>
            ,
            <given-names>P.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bock</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fokoue</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haase</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoekstra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruttenberg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sattler</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>OWL 2 web ontology language structural speci cation and functional-style syntax</article-title>
          .
          <source>W3C recommendation, The World Wide Web Consortium</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The META-SHARE language resources sharing infrastructure: Principles, challenges, solutions</article-title>
          .
          <source>In: Proceedings of the Eighth Conference on International Language Resources and Evaluation</source>
          . pp.
          <volume>36</volume>
          {
          <issue>42</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Rodriguez-Doncel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Villata</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A dataset of RDF licenses</article-title>
          .
          <source>In: Proceedings of the 27th Int. Conf. on Legal Knowledge and Information System (JURIX)</source>
          . pp.
          <volume>187</volume>
          {
          <issue>189</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Soria</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monachini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quochi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bel</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Odijk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The language resource strategic agenda: the arenet synthesis of community recommendations</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>48</volume>
          (
          <issue>4</issue>
          ),
          <volume>753</volume>
          {
          <fpage>775</fpage>
          (
          <year>2014</year>
          ), http://dx.doi.org/10.1007/s10579-014-9279-y
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Durco</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Windhouwer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>From CLARIN component metadata to linked open data</article-title>
          .
          <source>In: Proceedings of the 3rd Workshop on Linked Data in Linguistics</source>
          . pp.
          <volume>13</volume>
          {
          <issue>17</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Villegas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melero</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bel</surname>
          </string-name>
          , N.:
          <article-title>Metadata as linked open data: mapping disparate XML metadata registries into one RDF/OWL registry</article-title>
          .
          <source>In: Proceedings of the Ninth International Conference on Language Resources and Evaluation</source>
          . pp.
          <volume>393</volume>
          {
          <issue>400</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>