<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linghub: a Linked Data Based Portal Supporting the Discovery of Language Resources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>John P. McCrae1,2</string-name>
          <email>john@mccr.ae</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Philipp Cimiano2</string-name>
          <email>cimiano@cit-ec.uni-bielefeld.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>1Insight Centre for Data Analytics, National, University of Ireland</institution>
          ,
          <addr-line>Galway, Galway</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>2Cognitive Interaction Technology, Cluster of</institution>
          ,
          <addr-line>Excellence</addr-line>
          ,
          <institution>Bielefeld University</institution>
          ,
          <addr-line>Bielefeld</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>88</fpage>
      <lpage>91</lpage>
      <abstract>
        <p>Language resources are an essential component of any natural language processing system and such systems can only be applied to new languages and domains if appropriate resources can be found. Currently the task of finding new language resources for a particular task or application is complicated by the fact that records about such resources are stored in different repositories with different models, different quality and search mechanisms. To remedy this situation, we present Linghub, a new portal that aggregates and indexes data from a range of sources and repositories and applied the Linked Data Principles to expose all the metadata under a common interface. Furthermore, we use faceted browsing and SPARQL queries to show how this can help to answer real user problems extracted from a mailing list for linguists.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Linked Data</kwd>
        <kwd>Language Resources</kwd>
        <kwd>SPARQL</kwd>
        <kwd>Faceted Browsing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Language resources are essential for nearly all tasks in natural
language processing (NLP) and in particular for the adaptation of
resources and methods to new domains and languages. In order to
use language resources for new purposes they must first be
discovered and this can only be done if there is a comprehensive list of all
resources that may be available. To this there have been a number
of projects that have attempted to collect such a catalogue using
various methods and with differing degrees of data quality. We
present a new portal, Linghub, that aims to integrate all these data
from different sources by means of linked data and thus to create
a website, whereby all information about language resources can
be included and queried using a common methodology. The goal
of Linghub is thus to enable wider discovery of language resources
for researchers in NLP, computational linguistics and linguistics.
Currently, two approaches to metadata collection for language
resources can be distinguished. Firstly, we distinguish a curatorial
approach to metadata collection in which a repository of language
resource metadata is maintained by a cross-institution organization
such as META-SHARE [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] or CLARIN project’s Virtual Language
Observatory [17, VLO]. This approach is characterized though
highquality metadata that are entered by experts, at the expense of
coverage. A collaborative approach, on the other hand, allows
anyone to publish language resource metadata. Examples of this are
the LREMap [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or Datahub1. A process for controlling the
quality of metadata entered is typically lacking for such collaborative
repositories, leading to less qualitative metadata and
inhomogeneous metadata resulting from free-text fields, user-provided tags
and the lack of controlled vocabularies.
      </p>
      <p>
        Given the nature of this difference we wish to make data available
from multiple sources in a homogeneous manner and we saw the
development of a new linked data portal as the primary method to
achieve this. To this end we adopted a model based on the DCAT
data model [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] along with properties from Dublin Core [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In
addition, we used the RDF version [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] of the META-SHARE
model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to provide for metadata properties that are specific to
language data and linguistic research. As such, in this paper we
describe the creation of the largest collection of information about
language resources and briefly describe its publication on the Web
by means of linked data principles.
      </p>
      <p>The rest of the paper is structured as follows: firstly, we will
describe related work in Section 2, then we will describe the
collection and processing of data in Section 3. Next, we will describe
the portal and how we envision users can access the data in
Section 4 and examine how real user queries could be answered with
Linghub in Section 5. Finally we conclude in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>
        There have been several attempts to collect metadata about
language resources mostly associated with large infrastructure projects.
CLARIN has been collecting resources under a project called the
Virtual Language Observatory [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], using the Component
Metadata Infrastructure [3, CMDI] to collect common metadata values
from multiple sources. A similar project is META-SHARE [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
from the META-NET project where language resources are
collected and high-quality, manual entries are created for each record.
Similarly, the Open Languages Archives Community [2, OLAC]
collects data from a number of sources although the metadata
collected is not itself open. Another related project called SHACHI
has also collected some metadata [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. There has also been an
attempt to track language resources by means of assigning them an
International Standard Language Resource Number (ISLRN)
similar to an ISBN used to track books [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>Source</title>
        <p>
          Datahub
LRE-Map
META-SHARE
CLARIN VLO
All
On the contrary some resources have instead collected data directly
from creators of the resources, for example the LRE-Map [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]
collects data from authors of papers submitted to conference, such as
LREC. Similarly, Datahub collects resources directly from those
submitted to the website, but focusses primarily on linked data
resources.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. DATASET</title>
      <p>
        In order to ensure that all the data from many sources can be queried
in a homogenous manner we made sure that the metadata from all
the repositories mentioned in Table 1 was available as RDF. In
doing this, we aligned the proprietary schemas used in these
repositories to well-known semantic Web vocabularies and fixed
existing modeling errors, such as using percent-encoded URIs for titles
of resources or introducing URL links that would never resolve.
Two of our resources, LRE-Map and Datahub, were already
available in RDF, so that the conversion mainly involved developing an
appropriate URL schema so that datasets were uniquely identified
and thus to avoid collisions when uploading data into the Linghub
portal. A number of quality issues were also fixed in doing this
transformation, such as deciding whether property values should be
literals or URIs, reducing the number of blank nodes and reusing
existing metadata vocabularies such as VoID [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The other resources (CLARIN VLO and META-SHARE) were
available in XML. We developed a custom converter for each of these
resources building on a transformation language similar to XSLT,
which we developed. For META-SHARE, this was a challenging
task as there were nearly a thousand unique tags defined and each
one was examined to see if it was similar to an existing
Semantic Web vocabulary, and in fact we ended up mapping to FOAF 2,
SWRC 3 and the Media Ontology 4. In the case of CLARIN, there
was actually a significant difference between the XML schemas
used by each contributing instance, with only a small common
section giving the resource title and download link. We thus developed
distinct mappings for the largest 5 institutes.</p>
      <p>Two key issues emerge when collecting data from a heterogenous
set of sources such as we are doing. Firstly, the data is likely to
be noisy and inconsistent in the properties it uses and more
importantly in the values that these properties have. For example,
languages may be represented by their English names or alternatively
by means of the codes such as the ISO 639 codes5.</p>
      <p>
        Secondly, it happens relatively frequently that dataset descriptions
are duplicate as they are contained in multiple source repositories
(currently this affects 5.0% of resources). Furthermore, also
intrarepository duplicates exist, resulting from the fact that in some
2http://xmlns.com/foaf/spec/
3http://ontoware.org/swrc/
4http://www.w3.org/TR/mediaont-10/
5http://www.iso.org/iso/home/standards/
language_codes.htm
repositories one metadata record is created for each language a
resource is available in (this is the case for CLARIN for instance
and represents 35.0% of all resources). In order to remove these
duplications we used state-of-the-art word sense disambiguation
techniques, including Babelfy [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to identify common controlled
vocabularies and duplicate entries. For the case of properties we
mapped to several existing resources, including LexVo [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for
languages, and BabelNet for resource types. Duplicate entries were
not removed from the dataset but instead were marked with the
addition of the Dublin Core property is replaced by. In the case that
these entries were subsets of resources the target of this link would
be a new combined record for the entire resource and in the case
of duplicate records collected from distinct sources we referred to
the most complete triple, that is the record with the most triples.
The harmonization and description is described in more detail in
McCrae et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>Currently, there is no direct method for users to provide metadata to
the repository, however it is foreseen that users could submit valid
DCAT files to Linghub. We do note that Datahub allows any user
to submit a dataset and such datasets will quickly be picked up by
Linghub and added to the repository in this manner.</p>
    </sec>
    <sec id="sec-4">
      <title>4. THE LINGHUB PORTAL</title>
      <p>In order to enable users to quickly and easily discover datasets,
we set up a portal for browsing the dataset. Naturally we set this
up as a site that publishes the individual records as either RDF or
HTML, with the actual content delivered to the client decided by
means of content negotiation. We developed templates that render
the RDF in a readable manner, while still appearing close to the data
in such a way that users would get a consistent view of a dataset
record even if it came from a different original source and hence
had very different properties. In addition, we provide a number of
mechanisms by which users and automated agents can discover a
dataset. For users, we allowed resources to be discovered by means
of faceted browsing by enabling users to select properties and their
values. We fixed the list of properties in advance to those that have
been harmonized so as to not overload the user with choices for
properties that only occur for a few datasets and also to enable the
compilation of indexes to speed up page load times. In addition the
front page of Linghub contains a free-text search engine allowing
the users to query fields by a property. This free-text search engine
is powered by a separate index which includes not only the text
of data properties but also the labels of URIs which appear as the
value of object properties. Machine-based agents may access the
endpoint by means of SPARQL querying, although the endpoint
limits the agents to a subset of the SPARQL query language. The
goal of this is to enable constant query-time without overloading
our server. The nature of SPARQL makes it very easy for users
to write queries that are of a complexity that would not be easy to
answer and other sites have attempted to handle this by enforcing
timeouts on SPARQL queries. In general we find this solution to
be sub-optimal as it means that queries may fail unpredictably if
the server has many concurrent connections. Instead, we limit the
complexity of the queries themselves by requiring that the triples
have certain properties that can be easily answered. These include:</p>
      <sec id="sec-4-1">
        <title>1. A required limit on the number of results; 2. The property may not be a variable, thus limiting the number of results;</title>
        <p>
          3. The query must be a ‘tree’ in that every triples should be
connected from a single root node.
can search for both language and subject with the following
query7:
Furthermore, the SPARQL endpoint also by default returns
SPARQLJSON results[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], so that the results may be easily applied. This is
based on the fact that many clients, notably client-side Javascript
in browsers, will not accept XML due to security concerns. Other
clients may still obtain SPARQL-XML by supplying the
appropriate header or parameter in the query.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. USE CASES</title>
      <p>As a proof-of-concept for Linghub, we discuss a number of realistic
use cases that demonstrate the type of queries that can be answered
using Linghub. In order to get realistic use cases, we collected
queries for language resources from the Corpora List, a mailing list
used by researchers in corpus linguistics to discuss corpora. From
questions posed in February 2015, 3 queries are considered below
as they are clear and well-stated questions that would have feasible
answers. We chose these queries to provide an illustrative example
of queries that can be directly answered using the Linghub portal,
while discarding many other questions that were vague, unclearly
stated or misused linguistic terminology. We discuss these queries
and show how they can be formalized as SPARQL queries against
Linghub and discuss whether reasonable answers were retrieved.
1. “[...] desparately needs an Igbo corpus.” (Thapelo J.
Otlogetswe, Feb. 5th 20156)
Igbo is a language of Nigeria and Equatorial Guinea and is
identified with the language code ibo. Simply typing “Igbo”
into the search interface of Linghub finds a number of
resources that could be used. For many of these resources Igbo
is the value of the Dublin Core property subject. Although,
there is a language property some sources decided not to use
this Dublin Core category. In addition these resources are
marked with a type that is mapped to the META-SHARE
corpus individual even though the resources do not
originate from META-SHARE due to our harmonization. We
6http://mailman.uib.no/public/corpora/
2015-February/021993.html</p>
      <p>SELECT ?resource WHERE {
?resource
dct:language iso639:ibo |
dc:subject "Igbo" ;
dct:type metashare:corpus .</p>
      <p>}
2. “I am looking for a Lithuanian gigaword corpus for a
research project.” (Márton Makrai, Feb. 24th 20158)
Finding a corpus for a European language such as Lithuanian
is generally not a challenge, however this user also has the
requirement that the resource has over one billion words. We
can easily use the META-SHARE properties to return the
user a list of corpora with their associated sizes, as follows:
SELECT ?resource ?size WHERE {
?resource
ms:corpusInfo [
ms:languageInfo [
dct:language iso639:lit ;
ms:sizePerLanguage [
ms:size ?size ;
ms:sizeUnit ms:words
}</p>
      <p>]
Unfortunately, the results of this query show that no resource
in Linghub is over one billion words in size for Lithuanian.
3. “I am looking for freely available geotagged tweets
collection for research purpose.” (Md. Hasanuzzaman, Feb. 16th
20159)
7Note: We have implemented some syntactic extensions to
SPARQL. The | operator is a UNION with the same subject.
8http://mailman.uib.no/public/corpora/
2015-February/022103.html
9http://mailman.uib.no/public/corpora/
2015-February/022044.html
Several of the search terms here are unfortunately not found
anywhere in our data, namely ‘geotagged’ and ‘tweets’. It
would still be possible for this query to be answered by
looking at related keywords such as ‘Twitter’, and other aspects
of this query can be handled (e.g., ‘for research purpose’),
can be handled by means of the META-SHARE vocabulary.
In summary, we saw that in two of the three cases the users’ need
could be clearly expressed as a SPARQL query and that in one of
those cases, the query would return an answer as required, in the
second case no suitable dataset is recorded. In the final case, the
user’s query does not match the structured data found in Linghub,
but related resources can be found by using free text search. As
such, we see that Linghub enables users to better find their
resources than with previous approaches, although it is still not
satisfactory for all user queries. In particular, the crucial defect in the
final query is that there is no specific metadata that would indicate
if a resource is from a social media site or not, and this would
require deeper understanding of the textual components of resource
descriptions to better handle.</p>
    </sec>
    <sec id="sec-6">
      <title>6. CONCLUSION</title>
      <p>Linghub is a new site that collects data from a large number of
sources and makes it queriable through a common mechanisms.
Furthermore, the data has not only been converted to RDF it has
also been homogenized and linked to other bubbles in the
Linguistic Linked Open Data Cloud. As such, this resource is likely to pay
a pivotal role in enabling not only humans but also software agents
to find new resources and use them for applications in natural
language processing and artificial intelligence.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been funded by the LIDER project funded under the
European Commission Seventh Framework Program (FP7-610782),
the Cluster of Excellence Cognitive Interaction Technology ‘CITEC’
(EXC 277) at Bielefeld University, which is funded by the German
Research Foundation (DFG), and the Insight Centre for Data
Analytics which is funded by the Science Foundation Ireland under
Grant Number SFI/12/RC/2289.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Alexander</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hausenblas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>Describing linked datasets with the VoID vocabulary</article-title>
          .
          <source>Technical report, The World Wide Web Consortium</source>
          ,
          <year>2011</year>
          . Interest Group Note.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Simons</surname>
          </string-name>
          .
          <article-title>Extending dublin core metadata to support the description and discovery of language resources</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>37</volume>
          (
          <issue>4</issue>
          ):
          <fpage>375</fpage>
          -
          <lpage>388</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Broeder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Windhouwer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Van</given-names>
            <surname>Uytvanck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Goosen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Trippel</surname>
          </string-name>
          .
          <article-title>CMDI: a component metadata infrastructure. In Describing LRs with metadata: towards flexibility and interoperability in the documentation of LR workshop programme</article-title>
          , page
          <volume>1</volume>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Calzolari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Del Gratta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Francopoulo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rubino</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Russo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Soria</surname>
          </string-name>
          .
          <article-title>The LRE Map</article-title>
          .
          <article-title>Harmonising community descriptions of resources</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1084</fpage>
          -
          <lpage>1089</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Arranz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hamon</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          .
          <article-title>Using the international standard language resource number: Practical and technical aspects</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>50</fpage>
          -
          <lpage>54</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6] G. de Melo.
          <article-title>Lexvo.org: Language-related information for the linguistic linked data cloud</article-title>
          .
          <source>Semantic Web, page 7</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Federmann</surname>
          </string-name>
          , I. Giannopoulou,
          <string-name>
            <given-names>C.</given-names>
            <surname>Girardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Mavroeidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Minutoli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Schröder</surname>
          </string-name>
          .
          <article-title>META-SHARE v2: An open network of repositories for language resources including data and tools</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>3300</fpage>
          -
          <lpage>3303</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gavrilidou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Labropoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Desipri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Piperidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Papageorgiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Monachini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Frontini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Francopoulo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Arranz</surname>
          </string-name>
          , et al.
          <article-title>The META-SHARE metadata schema for the description of language resources</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1090</fpage>
          -
          <lpage>1097</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kunze</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <article-title>The Dublin Core metadata element set</article-title>
          .
          <source>RFC 5013</source>
          , Internet Engineering Task Force,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Maali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Erickson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Archer</surname>
          </string-name>
          .
          <article-title>Data catalog vocabulary (DCAT)</article-title>
          .
          <source>W3C recommendation, The World Wide Web Consortium</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          , V. Rodrig´ uez Doncel,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vila-Suero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gracia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Matteis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abele</surname>
          </string-name>
          , G. Vulcu, and
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          .
          <article-title>Reconciling heterogeneous descriptions of language resources</article-title>
          .
          <source>In Proceedings of the 4th Workshop on Linked Data in Linguisitcs</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Labrapoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gracia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Doncel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>One ontology to bind them all: The META-SHARE OWL ontology for the interoperability of linguistic datasets on the Web</article-title>
          .
          <source>In Proceedings of the 4th Workshop on the Multilingual Semantic Web</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Raganato</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          .
          <article-title>Entity linking meets word sense disambiguation: a unified approach. Transactions of the Association for Computational Linguistics (TACL</article-title>
          ),
          <volume>2</volume>
          :
          <fpage>231</fpage>
          -
          <lpage>244</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Piperidis</surname>
          </string-name>
          .
          <article-title>The META-SHARE language resources sharing infrastructure: Principles, challenges, solutions</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>36</fpage>
          -
          <lpage>42</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Seaborne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Feigenbaum</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Torres. SPARQL</surname>
          </string-name>
          <year>1</year>
          .
          <article-title>1 query results JSON format</article-title>
          .
          <source>W3C recommendation, The World Wide Web Consortium</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tohyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kozawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Uchimoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Matsubara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Isahara</surname>
          </string-name>
          .
          <article-title>Shachi: A large scale metadata database of language resources</article-title>
          .
          <source>In Proceedings of the 1st International Conference on Global Interoperability for Language resources</source>
          , pages
          <fpage>205</fpage>
          -
          <lpage>212</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>D. Van Uytvanck</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Stehouwer</surname>
            , and
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Lampen</surname>
          </string-name>
          .
          <article-title>Semantic metadata mapping in practice: the virtual language observatory</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1029</fpage>
          -
          <lpage>1034</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>