<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Introducing FREME: Deploying Linguistic Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Felix Sasaki</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tatiana Gornostay</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milan Dojchinovski</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Osella</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erik Mannens</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giannis Stoitsis</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Phil Ritchie</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kevin Koidl</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>felix.sasaki@dfki.de</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tilde</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>tatiana.gornostay@tilde.lv</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>InfAI</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>milan.dojchinovski@fit.cvut.cz</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>osella@ismb.it</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>iMinds</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>erik.mannens@ugent.be</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Agro-Know</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>stoitsis@agroknow.gr</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>VistaTEC</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>philr@vistatec.ie</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wripl</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>kevin@wripl.com</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper introduces the FREME project, a new Horizon 2020 innovation action. It aims at building an open framework of e-Services for multilingual and semantic enrichment of digital content, based on a reusable set of open Application Programme Interfaces and Graphical User Interfaces to FREME enrichment services. In addition, the paper discusses how the project deploys Linguistic Linked Data (LLD), especially existing LLD resources, LLD best practices and the LLD reference architecture.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The growing amount of digital content across languages, sectors and domains leads to
both challenges and business opportunities for many industries. Linked data (LD) and
language technology (LT) solutions exist, providing e.g. machine translation, entity
recognition, or multilingual linked data sets. These solutions face several issues, e.g.:
a plethora of content formats to process; adaptability and “silo solution” dependency;
and usability in an industry application scenario: the lack of adequate tooling for a
given or new group of user types (authors, translators, data wranglers or scientists
etc.) in selected business scenarios.</p>
      <p>The FREME framework addresses these issues by providing a reusable set of open
Application Programme Interfaces and Graphical User Interfaces to FREME
enrichment services. In this way, the project will improve the existing processes of
digital content management. The improvement will go through the whole content
value chain: content creation (or authoring), content translation/localization,
publishing and access to content including cross-language sharing and personalized content
recommendations. Thus, the goal is to open new opportunities for all sectors that are
involved in digital content management.</p>
    </sec>
    <sec id="sec-2">
      <title>3.1 Overview</title>
      <p>The main goal is to provide a set of interfaces for enrichment of digital content. We
understand digital content as any type of content that exists in a digital form (text,
video, audio, images, and others). Digital content is stored in various formats, for
example, textual content can be stored in structured formats (e.g. using the linked data
technology stack) and unstructured formats (using e.g. PDF, PPTX etc.). The share of
unstructured representation of digital content still prevails to a great extent. We will
work with the textual type of digital content in its structured and unstructured formats.
We aim at transforming unstructured content into its structured representation.
Enrichment services. By enrichment we understand annotation of content with
additional information. Content can be enriched on any step of its value chain that makes
it intelligent, i.e., discoverable, interoperable, and aggregatable further on.
Multilingual enrichment. Under multilingual enrichment we understand annotation
of content with additional linguistic information in a language or languages other than
the language of content itself. The following languages of the project partners and / or
their customers are in focus: English, German, Dutch, French, Italian, Spanish, Greek,
Latvian, Lithuanian, and Estonian.</p>
      <p>Semantic enrichment. Under semantic enrichment we understand annotation of
content with additional information that transforms unstructured content into its
structured representation. Semantic enrichment deploys the semantic richness provided by
linked open vocabularies, as well as data sources that are not (yet) available as linked
data.</p>
      <p>
        In FREME, enrichment information is stored in two ways: first, as information using
the Natural Language Processing Interchange Format (NIF), see
        <xref ref-type="bibr" rid="ref1">Hellmann et al.
(2013)</xref>
        . This storage is independent of the underlying format: the enrichment
information is stored in a stand-off manner, with pointers to the original location of content
items.
      </p>
      <p>Second, FREME allows storing information inside the enriched content itself. The
approach for storing the information then depends on the format in question. E.g. in
the case of HTML5 and XML, the project relies on the Internationalization Tag Set
(ITS) 2.0, see http://www.w3.org/TR/its20/ .</p>
    </sec>
    <sec id="sec-3">
      <title>3.2 e-Services</title>
      <p>The FREME framework will develop and integrate e-Services based on existing and
mature technologies. By e-Services we understand in most cases RESTful
webservices and graphical user interfaces. Details for each of the six e-Services will be
provided below. The framework is designed in an extensible manner, so that both
project partners as well as external partners are able to add more services.
e-Translation is based on cloud machine translation services for building custom
machine translation systems. The service takes content to be translated and the source
+ target language parameters as input.
e-Terminology: is based on cloud terminology services for terminology management
and terminology annotation. The service takes content as input and enriches the
content with information from terminology data bases.
e-Entity is based on entity recognition and existing linked entity datasets. The service
takes content as input and enriches it with information related to entities (e.g. names
of persons, places, etc.).
e-Internationalization is not a web service but the ability of other e-services to
handle Internationalization Tag Set (ITS) 2.0 so called “data categories”: these are
metadata items for handling the multilingual content production cycle. For example,
the “Translate” data category specifies whether a given piece of content should be
translated or not. This is of relevance e.g. for the e-Translation service. The
“Terminology” data category provides a standardized way to store the output of
eTerminology as part of a NIF representation or in content formats. The “Text
Analysis” data category allows storing the output of e-Entity.
e-Link is based on NIF and various (linked open) data sets. e-Link receives NIF
documents (with annotated entities) and performs enrichment relying on the data sets. An
example usage would be: have a NIF document with the entity “Berlin” annotated,
and enrich it with information from DBpedia about the current population of Berlin.
The difference to querying linked data directly via SPARQL is that the e-Link query
approach uses a query template mechanism. It hides the details of the actual query.
The user of e-Link just has to provide parameters to the template. In the above
example we assume a template called “query the population of a given place”. The
identifier for the place (e.g. http://dbpedia.org/resource/Berlin) is a parameter. The template
based approach replies to the needs of business case partners. It enables them to
enrich digital content with information from data sources without having to become
linked data experts. During configuration of e-Link, these experts set up the templates
for a given application.
e-Publishing has two aspects. First, it is a cloud based content authoring
environment, allowing content authors to deal with e-Services in a WYSIWYG manner.
Second, e-Publishing is a web service. Input is digital content in various forms (e.g.
plain text, HTML). Output is content made available in the EPUB 3 format, see
http://idpf.org/epub/30 . We use EPUB 3 since it is the standardised format for
representing digital book content.
e-Services can be deployed independently of each other. In addition, some e-Services
can benefit from processing the enrichment information created via other e-Services.
This will be made possible by using NIF as the format for storing and pipelining
enrichment information. Example chains of e-Services that are of interest for business
case partners (as of writing) are e.g.:
• e-Entity, e-Link
• e-Entity, e-Link, e-Translate
• e-Entity, e-Link, e-Translation, e-Publishing
• e-Entity, e-Link, e-Terminology
The framework will not hard wire these chains. Processing and then storing the
enrichment information using NIF will enable these and other chains. NIF will then be
both input and output of e-Services.</p>
    </sec>
    <sec id="sec-4">
      <title>3.3 Business Cases</title>
      <p>The innovation, robustness and usability of the FREME framework of e-Services will
be shaped by the four FREME real world business cases.</p>
    </sec>
    <sec id="sec-5">
      <title>BC 1: Authoring and publishing multilingually and semantically enriched</title>
      <p>eBooks. For publishing companies, digital content itself is exploding and is loosing
value. Via the project partner iMinds, we will build the e-Services so that they
provide additional value, going beyond digital publications. Initial discussions hint that
enhanced search engine optimisation via FREME e-Services could be an attractive
FREME application scenario for many publishing companies.</p>
      <p>BC 2: Translation and Localisation. In the translation and localisation industry,
demand for translation and the need for speed and quality are increasing. At the same
time prices being paid are going down. Via the project partner VistaTEC, FREME
will allow for integrating enrichment functionalities in localisation workflows. The
outcome will enable localisation companies to provide value beyond translation, e.g.
by adding information from linked (open) data sources via e-Link to translated
content.</p>
      <p>BC 3 Agriculture and food metadata. In the area of public sector information, the
discovery of data often is difficult due to missing multilingual metadata. E.g. many
metadata items in the agriculture area are not in the language of the person who wants
to use the metadata for search. Via the project partner Agro-know, a key player in the
agriculture and food data domain, FREME will tackle this challenge. The
eTranslation service will allow for (metadata) automated translation and in this way
will foster cross-language, public information access.</p>
      <p>BC 4 Web site personalisation. Web site personalisation is a field with many
emerging solutions and companies. Currently many solutions focus on English speaking
markets. Via the project partner Wripl, FREME will demonstrate how to deploy
underlying technologies e.g. via e-Entity in a larger number of languages, enabling
SMEs and start-ups to reach out to global markets.</p>
    </sec>
    <sec id="sec-6">
      <title>4 Linguistic Linked Data and FREME</title>
      <p>The data value chains that will be built with FREME rely heavily on linguistic linked
data sets (LLD). The LIDER project is crucial in providing the basis for a linguistic
linked data cloud1. LIDER fosters LLD as a basis for content analytics tasks of
unstructured multilingual cross-media content. The e-Services, especially e-Entity,
eLink and E-Terminology, can be seen as prototypical examples of content analytics
tasks. By providing these services together with e-Translation and several metadata
items relevant for translation workflow information (via e-Internationalisation),
FREME provides a technology stack that spans across content analytics and machine
translation technologies.</p>
      <p>The relevance of LIDER work on LLD can be seen in three areas: creation of LLD
data sets, best practices on multilingual linguistic linked data, and deployment of the
LIDER reference architecture.</p>
    </sec>
    <sec id="sec-7">
      <title>4.1 Linguistic Linked Data Sets</title>
      <p>Data sets are relevant for FREME in two ways. First, as content to be enriched via
FREME. These data sets mostly come from business case partners and are specific to
their needs and customers. Second, data sets to be used in enrichment e-Services.
Here, figure 2 provides an overview of relevant data sets.
1 See the LIDER homepage at http://lider-project.eu/ and an overview of the LLD at
http://linguistic-lod.org/llod-cloud .</p>
      <p>Data set
DBpedia</p>
      <p>Data type
Linked Data
(RDF)</p>
      <p>Data volume
500GB RDF data</p>
      <p>Sector
Multi-domain</p>
      <p>Language
119
NIF (RDF)
100k sentences total</p>
      <p>NLP</p>
      <p>mostly EN e-Entity
About 3.2 M terms
About 7 K concepts
32 K concepts
2 billion triples</p>
      <p>Multi-domain
Multi-domain
Agriculture and
food safety</p>
      <p>Geography
180GB web crawl</p>
      <p>Multi-domain
mostly EN e-Entity
Not all data sets are linguistic linked data sets. E.g. the LetsMT! parallel corpora are
not represented in RDF. Some of the data sets are moving towards the linguistic
linked data cloud and will be made available via FREME as linked data in the
technical sense, e.g. the terminological resource TaaS database.</p>
      <p>Initial discussions in the project have shown that in some cases a non-linked data
representation of linguistic resources with a clear path towards linked data (e.g. by
providing URIs for all data items) is the preferable approach for technical reasons.
For example, tooling for machine translation training or for training of statistical
named entity recognition currently is far more efficient relying on non-linked data
representations. On the other hand, for exchanging data sets and for enriching them
with additional information, linked data representations are the more adequate
approach.</p>
      <p>Similar lessons have been learned in the FALCON project, see
http://falcon-project.eu/ , with a focus on using LLD in translation and localisation
workflows. In FREME we will take a similar approach, driven by tooling available in
the four business cases. An additional goal then is to make this tooling linguistic
linked data aware, e.g. providing linked data enabled machine translation systems.</p>
    </sec>
    <sec id="sec-8">
      <title>4.2 Best Practices for the Creation of LLD</title>
      <p>As discussed in the previous section, many data sets are not yet available as linguistic
linked data. The LIDER project is working on best practises for creating LLD. This
endeavour is undertaken under the helm of the W3C “Best Practises for Multilingual
Linked Open Data” (BPMLOD) community group. As of writing, three best practises
have been drafted; see http://bpmlod.github.io/report/ for details.
• General Patterns: a set of common practices and patterns that can be applied to
publish linked data in a multilingual context.
• Guidelines for creating bilingual dictionaries.
• Guidelines for creating multilingual dictionaries.</p>
      <p>
        All of these best practices are relevant for FREME partners. The business case
partners have their own data sets that they want to deploy in e-Services. In e-Translation,
data sets can be used to provide translations for given lexical items. Currently there is
a plethora of formats for such data sets. As part of deploying the best practices, we
will rely on LEMON to representing bilingual and multilingual dictionaries.
BabelNet, see
        <xref ref-type="bibr" rid="ref2">Navigli and Ponzetto (2012)</xref>
        , is a resource that demonstrates the
approach towards multilingual dictionaries. BabelNet is crucial for building general,
domain independent multilingual and semantic enrichment applications. The tool
Babelfy shows how to deploy BabelNet for such applications. Babelfy also
demonstrates the approach (see section 4.1) of relying on LLD resources (here BabelNet) not
in a native RDF representation but using them in as part of other tooling, i.e. for
statistical training of named entity recognition.
      </p>
      <p>For the conversion of LLD resources, off-the-shelf tooling is crucial. In the realm of
the BPMLOD group, a TBX2RDF converter has been created. This implementation
will help FREME to tackle conversion tasks, e.g. for the forehand mentioned TaaS
database. In addition, it demonstrates the best practice of using the LEMON model for
representing terminological resources as linguistic linked data.</p>
    </sec>
    <sec id="sec-9">
      <title>4.3 FREME and the LIDER Reference Architecture</title>
      <p>Within the LIDER project, a reference architecture for working with linguistic linked
data has been created, cf. Koidl et al. (2014), esp. section 4.2. FREME instantiates
several parts of the architecture.
e-Services as LLD aware services. By using NIF as the interchange format between
e-Services, FREME provides e-Services as LLD aware services. In the terminology of
the reference architecture the e-Services allow to constitute LLD based workflows.
LLD publishing via e-Link. The forehand described conversion of TBX to RDF is a
publication of non LLD resources as LLD. It realises the best practices (see section
4.2) and relies on migrators like the forehand mentioned TBX2RDF convertor.</p>
    </sec>
    <sec id="sec-10">
      <title>Service composition via using NIF as interchange format. The combination of e</title>
      <p>Services via FREME is an example of linked data service composition. The
eServices are LLD aware: both service workflow input/output and the actual interfaces
comply to linked data standards and best practices. The current approach in FREME
does not foresee a declarative description for composing services. This is left to the
software client using the e-Services.</p>
      <p>We connect to the LIDER reference architecture for two reasons. First, it eases the
task of knowledge and technology transfer. Via the architecture, several FREME
partners learn more easily how to build linguistic linked data enabled applications.
Second, the reference architecture can also be seen as providing input to
standardisation activities within W3C or other organisations. The W3C LD4LT community
group serves as a forum also for LIDER and now also FREME to discuss this and
other potential standardisation tasks. In a long term the e-Services may become the
basis for standardised processing of both data and language technologies on the Web.
But this is not a main work item of FREME.</p>
    </sec>
    <sec id="sec-11">
      <title>5 Conclusions and Next Steps</title>
      <p>This paper introduced the FREME project: its motivation and goals, the outline of
eServices, and the four business cases. We then discussed the role of linguistic linked
data for FREME, including existing data sets, best practices and tooling for new data
sets, and the LIDER reference architecture for working with linguistic linked data.
As of writing, early prototypes of e-Services are available. The e-Services are being
developed in an agile manner, taking feedback from the four business cases into
account. Next steps will be including this feedback. A special focus that relates to the
topic of this paper is linguistic linked data sources. FREME is looking for working
with data set providers who could make their data set available via the e-Services,
data set users who want to use LLD for multilingual and semantic enrichment, and
providers of multilingual and semantic technologies. The last group could benefit
from FREME by making components available for a larger audience, crossing the
realms of data and language technologies as well as several industry sectors.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer and M. Brümmer</surname>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Integrating NLP using Linked Data</article-title>
          .
          <source>In: Proceedings of the 12th International Semantic Web Conference</source>
          , Sydney, Australia, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Navigli</surname>
            and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Buitelaar</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <source>LIDER Deliverable D3.1</source>
          .1:
          <string-name>
            <given-names>Linguistic</given-names>
            <surname>Linked Data Reference Architecture - Phase</surname>
          </string-name>
          <string-name>
            <surname>I</surname>
          </string-name>
          . Available at http://lider-project.eu/sites/default/files/D3.1.
          <fpage>1</fpage>
          -
          <lpage>v1</lpage>
          .0.
          <string-name>
            <surname>pdf</surname>
            <given-names>Navigli</given-names>
          </string-name>
          , R. and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ponzetto</surname>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>BabelNet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>193</volume>
          ,
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          ,
          <year>2012</year>
          , pp.
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>