<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collection and Curation of Language Data within the Eu- ropean Language Resource Coordination (ELRC)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Lösch</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valérie Mapelli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khalid Choukri</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Giagkou</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stelios Piperi- dis</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prokopis Prokopidis</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vassilis Papavassiliou</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miltos Deligiannis</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aivars Ber- zins</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrejs Vasiljevs</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eileen Schnur</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thierry Declerck</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josef van Genabith</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DFKI GmbH</institution>
          ,
          <addr-line>Stuhlsatzenhausweg 3, 66123 Saarbrücken</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ELDA</institution>
          ,
          <addr-line>9 rue des Cordelières, 75013 Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ILSP/Athena RC, Epidavrou &amp; Artemidos</institution>
          ,
          <addr-line>Maroussi, Athens</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Value of Language Data</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Tilde</institution>
          ,
          <addr-line>Vienibas gatve 75a, LV1004, Riga</addr-line>
          ,
          <country country="LV">Latvia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In order to help improve the quality, coverage and performance of automated translation solutions for current and future Connecting Europe Facility (CEF) digital services, the European Language Resource Coordination (ELRC) was set up in 2015 through a service contract operating under the European Commission's CEF SMART 2014/1074 programme. Since then, ELRC initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. All resources shared by the contributors were gathered and curated in the ELRC-SHARE Repository, after having passed the validation process developed by ELRC. This paper provides insights into the overall data collection and curation process (including both technical and legal validation of resources) employed within ELRC. The ELRC Helpdesk provides both technical and legal guidance (e.g. Intellectual Property Rights (IPR) clearance support) to potential data contributors, thus enabling the sustainable sharing of language data.</p>
      </abstract>
      <kwd-group>
        <kwd>ELRC</kwd>
        <kwd>language data</kwd>
        <kwd>LR evaluation</kwd>
        <kwd>LR validation</kwd>
        <kwd>LR curation</kwd>
        <kwd>PSI Directive</kwd>
        <kwd>Open Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>including also public service applications. Especially Natural Language Understanding
(NLU) (in particular chatbot applications) as well as domain- and application-specific
machine translation (MT) are expected to be of increasing importance in the European
LT-Market.</p>
      <p>In a report issued in September 2018 on a resolution passed by the European
Parliament (“Language equality in the digital age”3), an explanatory statement stresses that
especially smaller or minority languages are the ones to gain most from language
technologies, tools and resources. A clause in the adopted motion calls on the EC to “make
as a priority of language technology those Member States which are small in size and
have their own language” (European Parliament (2018), p.13) which coincides with the
recommendation of the CEF Market Study, according to which EU developers have a
clear competitive advantage with regard to building strong experience for such
languages thanks to the multilingual market. As such, investing in AI-driven LT research
for smaller European languages does not only yield significant impact and potential for
the European LT industry, but also for European public services and citizens who are
enabled to participate in and contribute to the European Digital Single Market.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The importance of a European Language Resource</title>
    </sec>
    <sec id="sec-3">
      <title>Coordination (ELRC)</title>
      <p>The development of language technologies in general – and machine translation
systems in particular – require substantial amounts of language data which for most
domains and smaller languages simply do not exist in sufficient volumes. On the other
hand, every day, European, national and regional public administrations in all EU
Member States deal with a huge amount of multilingual textual information in original
and translated form. With the European Language Resource Coordination (ELRC)4, the
European Commission (EC) has taken a decisive step towards minimizing language
barriers across Europe and enabling the development of European language
technologies by supporting the collection of this language data for all EU official languages,
Norwegian, Icelandic, and other languages of interest to the EU Member States (e.g.,
Chinese, Russian, Turkish etc.).</p>
      <p>
        ELRC was set up through the Connecting Europe Facility’s SMART 2014/1074
programme in April 2015 and since then is coordinated by DFKI5 (Deutsches
Forschungszentrum für Künstliche Intelligenz, Germany), in partnership with ELDA6
(Evaluations and Language Resources Distribution Agency, France), ILSP/Athena RC7
(Institute for Language and Speech Processing/Athena Research Centre, Greece),
TILDE8 (Latvia) and recently CrossLang9 (Belgium). It is governed by the Language
3 http://www.europarl.europa.eu/doceo/document/A-
        <xref ref-type="bibr" rid="ref8">8-2018</xref>
        -0228_EN.html
4 http://www.lr-coordination.eu
5 http://dfki.de/en
6 http://elda.org/en
7 http:// www.ilsp.gr/en
8 http://tilde.com
9 http://www.crosslang.com
Resource Board (LRB) – an oversight body consisting of National Anchor Points
(NAPs), i.e., leading technological and public service representatives for each CEF
affiliated country10. The main activities of ELRC include the collection of language
resources, the provision of corresponding language data sharing facilities, in principle
through the ELRC-SHARE repository for language data11, the support of data sharing
through awareness-raising events (country-specific workshops, European
conferences12) and the ELRC Technical and Legal Helpdesk13. By supporting the sharing of
this language data and by turning it into standardized, machine-readable formats and
actionable language resources (LRs), ELRC directly contributes to improving the
quality, coverage and performance of CEF eTranslation and other MT systems that need
multilingual LRs as training data.
3
3.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Data Collection and Curation Process within ELRC</title>
      <sec id="sec-4-1">
        <title>Data Collection Process</title>
        <p>Language data refers to any textual, audio or audio-visual data produced using human
language or data about human language (such as lexica, raw or annotated corpora,
language models etc.). The collection of one or more language data sets grouped together
according to certain criteria constitutes a language resource. ELRC collects in particular
the following types of language resources:
 Corpora, a set of documents or a text in one or more languages such as:
official documents in the official administration (decisions, legal acts etc.);
articles, reports, magazines, newsletters, etc.; sets of documents and their
translations (parallel/comparable); Translation Memories (i.e. aligned text
segments in the source and target language)
 Language or translation models
 Lexical and Conceptual resources, such as terminologies, glossaries,
thesauri, wordnets, ontologies.</p>
        <p>
          As mentioned above, ELRC mainly seeks to unlock language data that reside in public
organisations across Europe, but also to collect sharable data from other potential data
owners, such as research institutions, NGOs, etc. For instance, the Press and
Information Office of Cyprus uploaded, among others, two datasets14: a) “Bilingual
publications of the Press and Information Office of Cyprus” consisting of 9 pairs of EN-EL
PDF documents, and b) “Press Releases (0
          <xref ref-type="bibr" rid="ref1">1.2018</xref>
          -01.2019) of the PIO” consisting of
13 XML files, which contain (in its internal specific structure) the Press Releases in
Greek and the translation of some of them in English.
        </p>
        <p>Considering them as raw datasets, ELRC exploited its internal pipeline to process
the first dataset (i.e. text extraction from PDF files, identification of sentence pairs and
10 http://lr-coordination.eu/anchor-points
11 https://elrc-share.eu
12 http://www.lr-coordination.eu/events
13 http://www.lr-coordination.eu/helpdesk
14 PIO publications and press releases
parallel corpus filtering) with the purpose of constructing a precision-high parallel
corpus15 in TMX format16. Similarly, ELRC developed custom scripts to parse the XML
files of the second raw dataset, and then generated a parallel corpus of 5162 translation
units (TUs).</p>
        <p>ELRC additionally acts as the focal one-stop point where the language resources
created by relevant CEF-funded projects are gathered, documented and made available
to the EC or to the wider public, depending on their conditions of use.</p>
        <p>Given the diversity of the potential contributors and stakeholders, ELRC put in place
a multifaceted, yet straightforward and simple process for collecting and sharing LRs
(Lösch et al., 2018). Depending on the size of the data set, the technological readiness,
and the needs of potential contributors, participating organisations and individuals may:
 Upload the data directly to the ELRC-SHARE Repository17, as a zip file,
through the corresponding online contribution form (see
https://www.elrcshare.eu/repository/contribute).
 Send data and metadata files through the ELRC-SHARE-client API.
 Send data and metadata files through the ELRC-SHARE CEF eDelivery
access point.</p>
        <p>
          ELRC-SHARE
          <xref ref-type="bibr" rid="ref10">(Piperidis et al., 2018)</xref>
          is a web-based platform designed to cover the
whole life cycle of LR sharing: uploading, documentation, uploading of accompanying
documents, monitoring, and reporting, updating, browsing, delivery and downloading.
        </p>
        <p>
          The process is built on and inspired by META-SHARE
          <xref ref-type="bibr" rid="ref13">(Piperidis, 2012)</xref>
          and is
essentially an extension and adaptation of its latest version, mainly in terms of the
employed metadata schema, the user management module and the largely simplified
operational workflow.
        </p>
        <p>
          In addition to collecting contributions from external stakeholders, ELRC collects
and processes language data from scratch with web crawling techniques (Papavassiliou
et al., 201
          <xref ref-type="bibr" rid="ref3">3; Papavassiliou et al., 2018</xref>
          ). Web crawling is conducted using ILSP Focused
Crawler (ILSP-FC), a comprehensive end-to-end solution for the acquisition of
domain-specific monolingual and bilingual corpora from the web.
        </p>
        <p>The ELRC partners initially identified and documented public administration
websites (e.g., websites of ministries, local authorities, museums, etc.) as candidate sources
for the extraction of content relevant to the CEF Digital Service Infrastructures (DSIs)
and subsequently deployed ILSP-FC to acquire language resources for specific (EN-X)
language pairs, where X stands for any official EU languages in CEF-affiliated
countries.</p>
        <p>Starting from a list of seed URLs (i.e., the homepages of the identified websites), the
crawler fetches the web pages, extracts links from fetched web pages, adds the links to
the list of pages to be visited and so on. During this process, modules for page fetching,
15 Processed datasets: PIO publications and press releases
16 TMX stands for “Translation Memory eXchange” - see
https://www.maxprograms.com/articles/tmx.html for further details
17 https://elrc-share.eu
content normalization (i.e., conversion to UTF-8), boilerplate removal (i.e. "noisy"
elements like navigation headers, advertisements, disclaimers, etc.) and language
identification are used.</p>
        <p>The content of each page is then compared to a user-provided domain definition,
which consists of term triplets (&lt;relevance weight, (multi-word) term, subdomain&gt;)
that describe the targeted domain. If the page is classified as relevant to the targeted
domain, an XML file is generated containing basic metadata (e.g. title, URL, language,
domain, etc.) and the content split into paragraphs. The set of XML files corresponding
to domain-relevant webpages is then forwarded to further processing.</p>
        <p>As the web contains many near-duplicate documents, a module for (near)
de-duplication is exploited to eliminate the negative effect of duplicates in creating a
representative corpus. Having collected the in-domain, de-duplicated sets of pages in the targeted
languages, the next steps concern the detection of bitexts (i.e. pairs of documents that
could be considered parallel) and the identification of sentence pairs in each document
pair.</p>
        <p>Finally, a battery of criteria is applied with the purpose of filtering out sentence pairs
with potential alignment or translation issues, or of limited use for training MT systems.
As an example, after crawling the “Science in Poland” website18 of the Ministry of
Science and Higher Education of Poland, an EN-PL parallel corpus in the “education”
domain of about 28K TUs (Translation Units)19 was automatically constructed. Fig. 1
below summarizes the focused crawling process using the ILSP-FC within ELRC.</p>
        <p>Besides crawling websites of National Agencies, we also target websites of
International organisations, and broadcast websites, which make their multilingual content
available for use and process. To this end, the modules of the ILSP-FC toolkit (data
18 https://scienceinpoland.pap.pl/en
19 EN-PL “Science in Poland” corpus
acquisition, web page cleaning and normalization, detection of pairs of parallel
documents, identification of sentence pairs) were applied on VoxEurop and constructed a
multilingual (EN, DE, FR, ES, IT, PT, NL, CS, PL, RO) dataset of 927707 TUs in
total20.</p>
        <p>In case of any technical or legal questions around the preparation and/or submission
of language resources, potential contributors can contact the ELRC Helpdesk (email:
help@lr-coordination.eu, Skype: ELRC Helpdesk, phone: +33 970 440522).
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Data Processing and Validation</title>
        <p>To engage as many data holders as possible and to facilitate the sharing of language
data, ELRC encourages contributions of language data in a digital format. The
processing of the contributed data is then taken up by the ELRC language technologists
in order to convert them to machine readable data, ready to be used for training LT
systems. Each LR is analysed and processed by ELRC experts to ensure compliance
with the Language Resources Data Formats Specification agreed with the EC.
According to this specification, resulting parallel data should be provided in the TMX21 format
in UTF-8 encoding, without optional data fields (e.g. translator id, adjacent segments)
and without non-printable control characters. Monolingual corpora are to be delivered
in plain text format without any additional annotation, in UTF-8 encoding, single file
by language and resource, segmented into paragraphs. Terminology resources should
be provided in the TBX format.22</p>
        <p>As indicated earlier, the ELRC Helpdesk team can support providers with a wide
range of processing services, going from some basic data cleaning to more sophisticated
text extraction from problematic formats such as PDFs or OCR-requiring scanned
documents, etc. Processing has been performed for “The Udáras na Gaeltachta Corpus of
bilingual PDFs and Word documents (Processed)”23, which was provided by the
ADAPT Centre, DCU, Dublin, Ireland, and has undergone the following processing
steps:
1. Text extraction from PDF and DOC documents.
2. Automatic document pair detection.</p>
        <p>3. Automatic sentence alignment.</p>
        <p>This has resulted in a clean and aligned Irish-English corpus which can be used by the
community at large. All processing information about a particular resource is detailed
in the corresponding Processing Report that is provided with the processed version of
the corpus (together with the validation report, see below), so that ELRC can keep track
20 VoxEurop corpus
21 TMX stands for „Translation Memory eXchange”. TMX is an XML specification for the
exchange of translation memories.
22 TBX stands for “TermBase eXchange” TBX is an XML-based format for the representation
and exchange of terminology data
23
https://elrc-share.eu/repository/browse/the-udaras-na-gaeltachta-corpus-of-bilingual-pdfsand-word-documents-processed/ed8a4632c35711e8b7d400155d026706a233557ae9d246eb8a7a0dec13f35e9a/
of all data management steps performed and data users can check the processing steps
followed.</p>
        <p>
          ELRC also take care of the validation of their language resources. In the context of
ELRC, validation is understood as the quality control of a LR against a list of relevant
criteria
          <xref ref-type="bibr" rid="ref2">(Schneller et al., 2018)</xref>
          . It is important to note that due to the different processes
of gathering the data and their quality level, the validation may be conducted in two
different ways:
 Quick Content Check (QCC): it can be assumed that some data consist of
high-quality data in terms of content (in particular translations for multilingual data, data
produced by human experts), but require a technical- and legal-oriented evaluation.
Here validation includes:
o checking compliance of data with the ELRC objectives and scope,
o checking the format of provided data,
o checking that the metadata fields have been correctly filled in and are
compliant with the data content, and
o checking whether the legal information provided is compliant with the ELRC
scope.
 Extended content validation: a deeper content validation may be considered
necessary, for instance for data that derive from automatic processing (like crawling). On
a legal point of view, even though crawling already delivers the format that
corresponds to the LR production requirements, the list of crawled URLs was manually
checked to assess if the websites are under the scope of the PSI (Public Sector
Information Directive 2003/98/EC
          <xref ref-type="bibr" rid="ref5">(modified in 2013 by the Directive 2013/37/UE)</xref>
          ), for
details see Fig. 3 below). Content from websites that do not fall under the PSI
Directive, or content that is not explicitly marked as open with a permissive license,
must be excluded. On a technical point of view, the content validation may be
followed along two steps:
o automatic procedure: a series of processing steps may be considered necessary,
such as cleaning the data, removing TUs whose quality can be deemed as poor
by automated means.
o manual procedure: this step may be undertaken only when the Editor(s) deem
it necessary, e.g. for high-priority under-resourced languages where data
quality should compensate for the lack of quantity. Errors in Translation Units (TU)
need to be reported by human annotators. Depending on the acceptance
percentage (eg. 10%) decided by the Editor(s), the LR that is declared as
nonacceptable is discarded. In other particular cases, e.g. for high-volume
highpriority LRs, the corresponding LR may be kept with indicators showing the
probability of finding the same characteristics to help maximize the TU recall
(e.g. by taking TUs marked as “Machine-translated text” or “Free translation”
into account).
24 https://www.microsoft.com/en-us/download/details.aspx?id=52608
25 https://github.com/aboSamoor/pycld2
26 https://www.maxprograms.com/products/tmxvalidator.html
to be provided for each data set, and all available legal related documentation asserts
the quality of the data. ELRC’s Validation Guidelines are available online through the
ELRC website27.
3.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>IPR Clearance</title>
        <p>As shown in Fig. 3, in order to determine the appropriate license for a particular LR,
several questions need to be assessed:
 Does the data fall within the scope of the PSI (Public Sector Information</p>
        <p>
          Directive 2003/98/EC
          <xref ref-type="bibr" rid="ref5">(modified in 2013 by the Directive 2013/37/UE)</xref>
          ?
 Is the data protected by copyright? (National laws may contain rules
excluding certain works from copyright protection)
 If the data is protected by copyright, can I identify the owner of the
copyright or the author of the work? (see IPR Holder field)
 Is the data available under a public license? For example, certain datasets
are made available by the owner of copyright under a license that allows
reuse or redistribution free of charge (e.g., cc licenses, NCGL 1.0, OGL 3.0
etc.)
27
http://www.lr-coordination.eu/sites/default/files/common/ELRC
        </p>
        <p>SHARE%20repository_Guidelines%20for%20generic%20services%20projects_merged%20
-%20FINAL%2020200227.pdf.</p>
        <p>If no public license is clearly marked on the document, you should check
the terms of use or if any documentation may help you determine the
conditions of reuse of the material.</p>
        <p>
          It is important to point out that most of the aforementioned legal issues are debated on
fora organized by the ELDA team in charge of the legal helpdesk. For instance the legal
workshops organized as satellite events of LREC are major sources of knowledge and
input that help share the information about these issues within the community and allow
to get a clear picture of the various IPR contexts in different countries
          <xref ref-type="bibr" rid="ref6">(Choukri et. al,
2020)</xref>
          .
        </p>
        <p>If we go back to the example of the “The Udáras na Gaeltachta Corpus of bilingual
PDFs and Word documents (Processed)” presented in the Data processing and
Validation section, this resource was released under the legal status “Open under PSI” and
was then considered as a valid resource that could be published and shared widely. The
validation report also provides the Attribution text and the pre-existing rights to be
considered.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Summary and Conclusions</title>
      <p>Having received the first resources starting in Spring 2016, ELRC has managed to
collect almost 2,500 LRs and corresponding tools, covering all official EU languages, plus
Icelandic and both varieties of Norwegian, Bokmål and Nynorsk. This amounts to more
than 200 billion words in all EU languages, including bi- or multi-lingual contents in
digital editable formats ranging from reports, publications and other materials for
internal and external use, web contents and brochures, but also terminologies and glossaries.
More than 60 public sector organisations across Europe have shared their language data
with ELRC, including in particular national ministries, governmental bodies and public
services. However, in order to make all this data available and re-usable for the
development of MT systems, a dedicated validation and clearing process is necessary
involving both technical (manual and automatic) and legal evaluation. The ELRC workflows
and infrastructure that are in place facilitate a sustainable language data sharing, storing,
documenting and rendering cycle, thus unlocking data for training Language
Technology systems, for the benefit of the LT community across Europe.
5</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The European Language Resource Coordination (ELRC) is a service contract operating
under the EC’s Connecting Europe Facility SMART programme (starting from
SMART 2014/1074 in April 2015, continued under SMART 2015/1091 LOT 2
“Language Resource coordination and collection with related legal and technical work” and
SMART 2015/1091 LOT 3 “Acquisition of additional Language Resources and related
refinement/processing services and their provision of the Language Resource
Repository of CEF Automated Translation Platform” until end of 2021 within SMART
2019/1083 “Action on Automated Translation Core Service Platform (CSP)”).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>European</given-names>
            <surname>Commission</surname>
          </string-name>
          (
          <year>2017</year>
          ). eTranslation - Making
          <string-name>
            <surname>European Digital Public Services Multilingual</surname>
          </string-name>
          . Available at: https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation [last accessed:
          <volume>22</volume>
          .
          <fpage>02</fpage>
          .
          <year>2018</year>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Schneller</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fernandez-Barrera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mapelli</surname>
            ,
            <given-names>V;</given-names>
          </string-name>
          <string-name>
            <surname>Popescu</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Arranz</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Giagkou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Prokopidis</surname>
            ,
            <given-names>P</given-names>
          </string-name>
          ; Papavassiliou,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Rozis</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          (
          <year>2018</year>
          ). European
          <string-name>
            <surname>Language Resource Coordination - Validation Guidelines</surname>
          </string-name>
          . Available at https://lr-coordination.eu/sites/default/files/common/Validation_guidelines_CEF-AT_
          <year>v6</year>
          .2_20180720.pdf [last accessed:
          <volume>11</volume>
          .
          <fpage>01</fpage>
          .
          <year>2020</year>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. CEF: https://ec.europa.eu/digital-single-market/en/connecting-europe-facility
          <source>[last accessed: 22.02</source>
          .
          <year>2018</year>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. CEF eTranslation: https://ec.europa.eu/cefdigital/wiki/display/CEFDIGITAL/eTranslation</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <source>Directive</source>
          <year>2013</year>
          /37/UE: http://eur-lex.europa.eu/legal-content/EN/TXT/?uri=
          <source>CELEX%3A32013L0037 [last accessed: 22.09</source>
          .
          <year>2017</year>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lindén</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rigault</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Siegert</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          (
          <year>2020</year>
          ):
          <source>Proceedings of the LREC2020 Workshop on Legal and Ethical Issues</source>
          , Marseille, France, European Language Resources Association (ELRC). Available at: https://www.aclweb.org/anthology/
          <year>2020</year>
          .legal2020-
          <fpage>1</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <source>European Parliament: Report on Language Equality in the Digital Age</source>
          (
          <year>2018</year>
          /2018(INI)),
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. ELRC-SHARE: http://www.lr-coordination.eu/resources [last accessed:
          <volume>22</volume>
          .
          <fpage>02</fpage>
          .
          <year>2018</year>
          ]
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Lösch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Mapelli</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vasiļjevs</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Smal</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Schnur</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Choukri</surname>
            , K.; van Genabith,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>: “European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management”</article-title>
          ,
          <source>In: Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ), Miyazaki, Japan,
          <string-name>
            <given-names>European</given-names>
            <surname>Language Resources Association (ELRA).</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Labropoulou,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Deligiannis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Giagkou</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Managing Public Sector Data for Multilingual Applications Development</article-title>
          ,
          <source>In Proceedings of the 11th Language Resources and Evaluation Conference (LREC</source>
          <year>2018</year>
          ), Miyazaki, Japan, May
          <year>2018</year>
          .
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Papavassiliou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prokopidis</surname>
            ,
            <given-names>P. &amp; G. Thurmair.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>A modular open-source focused crawler for mining monolingual and bilingual corpora from the web</article-title>
          .
          <source>In Proceedings of the Sixth Workshop on Building and Using Comparable Corpora</source>
          , pages
          <fpage>43</fpage>
          -
          <lpage>51</lpage>
          . Sofia, Bulgaria: Association for Computational Linguistics
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Papavassiliou</surname>
          </string-name>
          , V.;
          <string-name>
            <surname>Prokopidis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2018</year>
          )
          <article-title>Discovering parallel language resources for training MT engines</article-title>
          .
          <source>In Proceedings of the 11th Language Resources and Evaluation Conference (LREC</source>
          <year>2018</year>
          ), Miyazaki, Japan, May
          <year>2018</year>
          .
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>The META-SHARE Language Resources Sharing Infrastructure: Principles, Challenges, Solutions</article-title>
          .
          <source>In Proceedings of the Eighth International Language Resources and Evaluation (LREC</source>
          <year>2012</year>
          ), Istanbul, Turkey.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Smal</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lösch</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          .; van Genabith,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ; Giagkou,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Busemann</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>2020</year>
          )
          <article-title>: “Language Data Sharing in European Public services - Overcoming Obstacles and Creating Sustainble Data Sharing Infrastructures” in: Nicoletta Calzolari</article-title>
          , Frédéric Béchet, Philippe Blache, Christopher Cieri, Khalid Choukri, Thierry Declerck, ,
          <string-name>
            <surname>Sara</surname>
            <given-names>Goggi</given-names>
          </string-name>
          , Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Asuncion Moreno, Jan Odijk, Stelios Piperidis (eds.):
          <source>Proceedings of the Twelfth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2020</year>
          ), Pages
          <fpage>3443</fpage>
          -3448, Marseille, France, European Language Resources Association (ELRA).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Babych</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Measuring comparability of documents in non-parallel corpora for efficient extraction of (semi-) parallel translation equivalents</article-title>
          .
          <source>In Proceedings of the Joint Workshop on Exploiting Synergies between Information Retrieval and Machine Translation (ESIRMT) and Hybrid Approaches to Machine Translation (HyTra)</source>
          (pp.
          <fpage>10</fpage>
          -
          <lpage>19</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Vasiļjevs</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Rozis</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Kalniņš,
          <string-name>
            <given-names>R.</given-names>
            ;
            <surname>Bērziņš</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Collecting Language Resources from Public Administrations in the Nordic and Baltic Countries</article-title>
          .
          <source>In Proceedings of the 11th Language Resources and Evaluation Conference (LREC</source>
          <year>2018</year>
          ), Miyazaki, Japan, May
          <year>2018</year>
          .
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>