<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Curation Policies for EUDAT Collaborative Data Infrastructure</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>© Vasily Bunakov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Alexia de Casanove</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Pascal Dugénie</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Rene van Horik</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Simon Lambert</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Javier Quinteros</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>© Linda Reijnhoudt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Science</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Technology Facilities Council</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Harwell Campus</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>United Kingdom</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CINES</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Montpellier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Data Archiving</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Networked Services (DANS)</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>The Hague</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Netherlands</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>GFZ German Research Centre for Geoscience</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Potsdam</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany vasily.bunakov@stfc.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>casanove@cines.fr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>dugenie@cines.fr</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rene.van.horik@dans.knaw.nl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>simon.lambert@stfc.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>javier@gfz-potsdam.de</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>linda.reijnhoudt@dans.knaw.nl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Proceedings of the XIX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2017)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>72</fpage>
      <lpage>78</lpage>
      <abstract>
        <p>The work outlines an approach to the development of a data curation framework in the EUDAT Collaborative Data Infrastructure. Practical use cases are described as well as provisional results of defining granular data curation policies with high potential for their machine-executable implementation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        EUDAT Collaborative Data Infrastructure (CDI) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a
European e-infrastructure of data services and
information resources in support of research. This
infrastructure and its services have been developed in close
collaboration with over 50 research communities spanning
across many different scientific disciplines, with more
than 20 major European research organizations, data
centres and computing centres involved. Researchers,
research communities and service providers can use
EUDAT data services to manage research data
according to their own needs.
      </p>
      <p>
        The EUDAT services offering has emerged as a
result of two consecutive FP7 and Horizon 2020 projects,
with the actual services focused on different aspects of
data management and data use, and supported by a
variety of information technology stacks. The major
EUDAT services [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] are:
• B2ACCESS – identity and authorization service;
• B2HANDLE – service for assigning and managing
persistent identifiers;
• B2DROP – service for secure and trusted data
exchange;
• B2SHARE – service for sharing small-scale “long
tail” data;
• B2SAFE – robust, safe and highly available service
for storing large-scale data in community and
departmental repositories;
• B2FIND – service for data discovery across the
EUDAT infrastructure (data catalogue).
      </p>
      <p>Data curation (or digital curation) is the selection,
preservation, maintenance, collection and archiving of
digital assets and hence is the essential part of research
data management. Sensible data curation requires
establishing and developing long-term repositories of digital
assets for their current and future use by researchers and
wider society. Collaborative data infrastructures like
EUDAT that span across the borders should play a
significant role in research data curation.</p>
      <p>Historically, EUDAT services have been built with
only a few considerations for conscious data curation,
with secure and controlled access to data being one of
the major initial goals to achieve. Other aspects of data
curation started playing a more prominent role when
services matured to production stage and became a part
of an operational collaborative infrastructure.
Specifically, operational requirements of B2SAFE service (that
currently offers what long-term digital preservation
projects typically call “bit-level” preservation), as well
as automated data transfers across interrelated
B2DROP, B2SHARE and B2FIND services have made
it essential to systematically explore the topic of data
curation in EUDAT.</p>
      <p>The decision was made to formulate the core
approach to data curation with the involvement of two
prominent unrelated research communities with
substantial amounts of data to manage and then, using these
two use cases as a proof-of-concept for clearly
formulated data curation activities, get other user
communities involved.</p>
      <p>
        Another decision made was to reuse the outputs of
the SCAPE project [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Research Data Alliance
Practical Policy Working Group [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in order to set up a
reasonable data curation framework for EUDAT.
      </p>
      <p>
        The rest of the paper outlines the core use cases,
characterizes the SCAPE and RDA outputs that are
deemed to be applicable in EUDAT context, describes
mapping of SCAPE policy elements [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to granular data
policies in EUDAT, and sets directions for further
works on data policies in EUDAT.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2 HERBADROP use case</title>
      <sec id="sec-2-1">
        <title>2.1 Motivations and relation to EUDAT services</title>
        <p>
          The HERBADROP data pilot [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] aims to offer an
archival service for long-term preservation of herbarium
specimen images and to develop innovative processes
for extracting metadata from those images.
HERBADROP follows the global trend towards
scalable industrial-style digitizing of herbaria specimens. It
is designed as both an archival service for long-term
preservation of herbarium specimen images and a tool
for analysing and extracting information written on the
image, both supported by CINES [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], by using Optical
Character Recognition (OCR) analysis.
        </p>
        <p>Making the specimen images and data available
online from different institutes allows cross domain
research and data analysis for botanists and researchers
with diverse interests (e.g. ecology, social and cultural
history, climate change).</p>
        <p>
          Herbaria hold large numbers of collections:
approximately 22 million herbarium specimens exist as
botanical reference objects in Germany, 20 million in France
and about 500 million worldwide. High resolution
images of these specimens require substantial bandwidth
and disk space. New methods of extracting information
from the specimen labels have been developed using
OCR but using this technology for biological specimens
is particularly complex due to the presence of biological
material in the image with the text, the non-standard
vocabularies, and the variable and ancient fonts. Much
of the information is only available using handwritten
text recognition or botanical pattern recognition which
are less mature technologies than OCR [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>The proposed platform is expected to support or
even substitute costly manual data input as much as
possible. The platform will also curate and enrich
metadata resulting from image analysis using optical
character recognition (OCR) and pattern matching.</p>
        <p>Results are exposed as platform independent Web
services which can be effectively integrated into
herbarium data management systems as well as metadata
capture workflows. Since 2016, five European community
partners1 have been involved. Their contribution to the
1The partners in the HERBADROP data pilot are: Musée
National d’Histoire Naturelle (MNHN) – Paris, France; Royal
Botanic Garden of Edinburgh (RBGE) – United Kingdom;
Botanic Garden and Botanical Museum (BGBM) – Berlin,
pilot represents a business model that can be potentially
replicated by other institutes.</p>
        <p>The EUDAT B2SAFE service is used in the first
step of the ingestion process. Existing images of
herbarium specimens along with the associated metadata are
transmitted to the CINES repository using B2SAFE
transfer service. The ingestion into B2SAFE is carried
out in accordance with the centralized persistent
identifiers (PID) management system used in EUDAT. It is
envisaged that discovery and visualization of the data
objects will be performed with the EUDAT B2FIND
service.</p>
        <p>The data workflow in HERBADROP is represented
by Figure 1.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Data curation scenarios for HERBADROP</title>
        <p>
          The HERBADROP communities have expressed their
wish to implement specific use cases such as identifying
duplicates amongst specimens from the different
museums. This kind of requirement is very useful to improve
EUDAT services. Another example of policy is long
term preservation that involves a number of controls
including file format verification and metadata quality.
Amongst HERBADROP users, two partners of the
community have proposed practical scenarios for data
curation: Digitarium [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and the Royal Botanic Garden
of Edinburgh (RGBE).
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Scenario proposed by Digitarium (Finland)</title>
        <p>
          Digitarium [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] would like to use Optical Character
Recognition (OCR) data to generate metadata based on
the label information available for the herbarium
specimen. Firstly, a Natural Language Processing based
system will be used to do OCR quality check and extract
relevant terms. Then metadata will be either
automatically generated, or manually inserted through the
transcription portal [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] but with the help of OCR data.
        </p>
        <p>More general for EUDAT infrastructure services,
Digitarium would like to utilize and integrate them into
the whole digitisation process of natural history
biological collections. The data flow goes from the beginning
Germany; Digitarium – Finland; Naturalis Biodiversity Center
– Netherlands
of the digitisation process i.e. imaging, to storage, then
to transcription and analysis, until accessing. This
involves data storage, high-performance computing
resources, and web services in EUDAT.</p>
        <p>Firstly, the images from the imaging station can be
transferred into EUDAT storage for long-term
preservation instantly or in batch. After transferring, HPC can
access the images and do OCR to extract label
information to generate preliminary metadata. This metadata
has to be associated with corresponding images. The
data can be openly accessed. However, the access rights
of data have to be set up for different purposes, such as
endangered species protection.</p>
        <p>Secondly, using HTTP APIs, the images and their
metadata can be accessible from EUDAT by data-owner
portals. Therefore, browsing and transcribing are
available. Updated metadata will be transferred back into the
EUDAT B2SAFE service. Different versions of
metadata have to be kept.</p>
        <p>Thirdly, the metadata is indexed. Therefore, the data
can be searched or filtered based on different terms for
further scientific usages. HPC resources can be utilized
also on the data for different researches.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Scenario proposed by RBGE (the Royal Botanic</title>
      </sec>
      <sec id="sec-2-5">
        <title>Garden of Edinburgh) in association with MNHN (Musée National d’Histoire Naturelle) – Paris</title>
        <p>
          The core of the concept of HERBADROP is to harvest
metadata from OCR analysis of the text that is a part of
herbarium images. The choice has been to proceed to a
full text analysis using a Lucene-based engine
Elasticsearch [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The objective of this approach is to
provide a powerful interface for further data curation as
part of the preservation process (identifying duplicates,
or inducing new taxonomic relations, etc.), see [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>Safeguarding long-term data storage is an important
precondition for reliable access to herbarium specimen
information. Thanks to this pilot, it is possible to
envisage long-term storage for herbarium specimen images.
Moreover, the specimens will be discoverable by the
entire scientific community. Thus, undescribed species
stored in herbaria can be examined by experts to aid
identification and discovery of new species.</p>
        <p>Distribution information for species over time can
be evaluated and these data could provide evidence of
the point in time when an invasive species first occurred
in a certain area. Historians could analyse herbarium
data to create itineraries for historical characters. The
data can be used to calibrate predictive models of the
oncoming changes in biodiversity patterns under global
threats. This diverse information will be useful for a
wide user community including conservationists, policy
makers, and politicians.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 GEOFON use case</title>
      <p>
        The second use-case concerns GFZ, the German
Research Centre for Geosciences. GFZ provides valuable
seismological services in the form of a seismological
infrastructure named GEOFON [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to research and
better understand our complex system Earth.
      </p>
      <p>
        GFZ is one of the members of the EPOS initiative
(European Plate Observatory System) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and, in this
context, collaborates with other two seismological data
centres related to EPOS (KNMI, INGV) in the EUDAT
project.
      </p>
      <p>Besides being one of the fastest earthquake
information provider worldwide, GEOFON is also one of the
largest nodes of the European Integrated Data Archive
(EIDA) for seismological data under the ORFEUS2
umbrella, which is a distributed data centre established
to (a) securely archive seismic waveform data and
related metadata, gathered by European research
infrastructures, and (b) provide transparent access to the archives
by the geosciences research communities.</p>
      <p>
        The internal structure of GEOFON is based on three
pillars:
• A global seismic network operated in close
collaboration with many partner institutions with focus on
EuroMed and Indian Ocean regions. The network
consists of ca. 110 high quality stations, which
acquire data in real time [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
• A global earthquake monitoring system which uses
data from GEOFON and partner networks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. It
publishes most timely earthquake information. First
automatic solutions are available few minutes after
the events and mostly manually revised later.
• A comprehensive seismological data archive for GFZ
and partner networks, for permanent networks as
well as for temporary deployments.
      </p>
      <p>For some GEOFON partner networks, GEOFON
acts as a data centre saving a replica of the original copy
and at the same time as a data distribution centre.
Additionally, data from many temporary station deployments
are permanently archived at GEOFON, in particular
passive seismological experiments of the GFZ
Geophysical Instrument Pool Potsdam (GIPP) and the
German Task Force Earthquake.</p>
      <p>Most data are open for public access, as well as
realtime data feeds when available. However, there is a
small amount of data under an embargo period, usually
for a limited amount of time (3–4 years).</p>
      <sec id="sec-3-1">
        <title>3.1 Data workflow in GEOFON</title>
        <p>GEOFON supports two scenarios for the ingestion of
data into its archive: one for permanent networks and
one for temporary (and most probably already finished)
experiments.</p>
        <p>
          Usually, raw data is transmitted to the data centre
with the metadata (technical hardware description) to be
able to operate with it. In the case of permanent
networks raw data is received continuously from the
stations around the world via satellite using a protocol
2 Observatories and Research Facilities for European
Seismology (http://www.orfeus-eu.org/)
called SeedLink [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], a real-time data acquisition
protocol which works on TCP. The packets of each
individual station are always transferred in timely (FIFO) order.
        </p>
        <p>In the case of temporary experiments network
operators provide usually, first, the metadata needed to use
the data, and in a second phase the data to be archived.
Data transmission can be done as in the permanent
networks case (SeedLink protocol), or can also be
transmitted to the data centre by the network operator using
some client-server tools provided by GEOFON, which
will do the first quality check of the data format. In
some cases, both methods could be used.</p>
        <p>A schematic view of the workflow at GEOFON can
be seen in Fig. 2. It should be noted that this workflow
is also valid for many of the seismological data centres
belonging to EIDA/ORFEUS. For instance, the other
two data centres piloting EUDAT services (KNMI and
INGV).</p>
        <p>Figure 2 Data workflow from GEOFON. It also
represents the workflows from a generic seismological data
centre as the ones under the EIDA/ORFEUS initiative.
Boxes in black are generic activities from the data
centre. Blue boxes show activities related to the EUDAT
service B2SAFE, while brown boxes show the tasks
related to B2HANDLE</p>
        <p>In both cases, permanent and temporary networks,
data go through some quality checks after being
received. When data are sent in real-time there is a first
control by sorting the records before actually ingesting
them into the archive (~1 day after reception). After 4–6
weeks, for stations that still have the buffered data, a
gap filling process is started.</p>
        <p>When data have been bulk uploaded to the data
centre by the network operator, it is immediately checked
to exclude overlaps. In this case, as all available data is
copied off-line, there is no need to check for problems
related to real-time transmission, like gaps and proper
order of records, as they are checked by the automatic
archiving tools.</p>
        <p>In the case that the data is under an embargo period,
the access control list is created or updated. After
completion of the last steps, data is opened through standard
access protocols.</p>
        <p>The internal organization of the archive is based on
an approach called SeisComP3 Data Structure (SDS).
This means that files are stored under a predefined
directory structure, which uses the codes from the
network/station/channel used to record the data as well as
the year. The continuous time series are stored in a
standard seismological format called Mini-SEED. The
time series are split in daily files for each recording
sensor and, therefore, files are closed when the day
finishes. At that moment, “new” data (recently closed
files) can be processed to obtain derived products from
them. For instance, quality metrics on the data or
detailed availability information, which are offered to our
users by means of a Web service.</p>
        <p>Once the data is archived users can make use of any
of the services provided by GEOFON to retrieve it.
Considering that there are different services which can
provide the data to the users, the usage statistics is
centralized in one database to be able to analyse the impact
of the data on the community regardless of the method
used to retrieve it.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2 Service hosting environment with the inclusion of</title>
      </sec>
      <sec id="sec-3-3">
        <title>EUDAT services</title>
        <p>Considering the workflow depicted in the previous
section, GEOFON introduced some EUDAT services in
order to automate and/or improve some of the tasks
related to it.</p>
        <p>Many services are being provided at GEOFON (e.g.
interactive web portals, proprietary protocols to get data
or derived products), with two of them (Station-WS and
Dataselect) being particularly important, as they are
international standards and the core services for the
community upon which other services are built.
StationWS serves the information describing the hardware and
everything related to the deployment, while Dataselect
serves the data.</p>
        <p>Two main EUDAT services have been integrated in
the GEOFON workflow; namely, B2SAFE and
B2HANDLE. The former is used to accomplish most of
the Data Management tasks, while the latter is used to
manage/store Persistent Identifiers (PIDs).</p>
        <p>As the archive is stored in a directory structure from
a partition, the B2SAFE service “mounts” the archive as
an external resource in read-only mode.</p>
        <p>One of the main requirements for the Data Policies
at GEOFON is the capability to trigger processes based
on the inclusion of new data. In the context of B2SAFE,
this can be done by means of automatic rules which are
executed under certain conditions (e.g. new data
ingested).</p>
        <p>With the proper rules we can enforce that, after new
data is detected by B2SAFE, a certain set of actions is
executed. For instance, the derived products can be
generated and data can be replicated to a partner data
centre from the EUDAT CDI, the Karlsruhe Institute of
Technology (KIT). Also, as part of this replication
process, persistent identifiers (PIDs) are generated for each
file, so that the PID can be used to globally and
univocally identify the file.</p>
        <p>PIDs are managed and stored by means of the
already mentioned service called B2HANDLE, which is
based on a Handle Server and other libraries developed
within the project. GFZ has a broad expertise in this
type of tools and, therefore, we decided to deploy our
own B2HANDLE server and work with our local
instance.</p>
        <p>Each generated PID is stored with a set of key-value
pairs called “PID Record”. The information in the PID
Record allows, among other things, to track other
copies of the file in different data centres or validate its
integrity by means of pre-calculated checksums.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.3 Data Policies to apply at GEOFON through</title>
      </sec>
      <sec id="sec-3-5">
        <title>EUDAT services</title>
        <p>After the formalization of the internal workflows at
GEOFON, and the inclusion of requirements from the
community and the data centre, we defined a set of Data
Policies to be enforced by means of the tools available
within EUDAT and new developments, which could be
useful for different communities.</p>
        <p>Some of them are related to the Replication process.
For instance:
• replicate every new file in the archive to our internal
backup server;
• if we are the official provider of the data in a file,
replicate it to an off-site partner within the EUDAT
CDI;
• seismological data that does not belong to us but
comes from our earthquake early monitoring system
should be kept for 6 months only; data still need to
be replicated to the internal server;
• file deletion must not be possible in an automated
way. In case that the system detects that a file should
be deleted, an email should be sent to the appropriate
operator.</p>
        <p>Regarding the access control of the files:
• “Restricted data” must be tagged and proper access
control must be applied to them;
• access restrictions can be automatically removed
after a period of time (embargo period);
• data must be able to be accessed via an HTTP API
respecting the ACL (Access Control List);
Regarding automatic metadata extraction:
•</p>
        <p>Metrics derived from the data must be automatically
calculated to populate some of our services when
new data is ingested.
• Detailed statistics related to the data access should be
available for the data owners/creators.
• In case that data are modified (e.g. correcting errors,
filling gaps), this information should be available for
future use (provenance information).</p>
        <p>Regarding the integrity of the stored data:
• a weekly process will select ~2% of the folders in
our archive and verify that the synchronization is
correct; the idea is that every file will be checked at
least once in a year;
• check that the data is stored in SDS format;
• start and end time of network/station operation must
be available and data outside this time span must not
be allowed.</p>
        <p>The identified relevant policies are being gradually
implemented using generic EUDAT services and
GEOFON-specific software.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Mapping of EUDAT data policies to</title>
    </sec>
    <sec id="sec-5">
      <title>SCAPE and RDA policy curation frameworks</title>
      <p>
        For the design and implementation of data curation
actions in EUDAT, the relevant outputs of SCAPE project
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Practical Policy Working Group of the Research
Data Alliance [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] have been identified. SCAPE outputs
are perceived of high quality owing to the advanced
thinking that considered long-term digital preservation
policies at a granular level suitable for the
machineexecutable implementation. RDA Practical Policy
Working Group outputs are a result of a substantial
international collaborative effort including experts in
iRODS platform [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] that is a technological foundation
of the EUDAT B2SAFE service.
      </p>
      <p>
        For SCAPE, we used the catalogue of preservation
policy elements [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that is a systematized compendium
of granular policies with examples of what SCAPE
called “control policies” (granular statements that are
easily translatable to machine-executable functions),
and for the RDA Practical Policy Working Group it was
their practical policy implementations report [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] that
compiled a set of machine-executable functions for
iRODS platform [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>In addition to this top-down retrospective review of
the SCAPE and RDA outputs, a bottom-up analysis of
control policies applicable to the GEOFON and
HERBADROP use case was performed, with a number
of control policies identified as prime candidates for
implementation in EUDAT B2SAFE. These policies are
presented in Table 1.</p>
      <p>
        Then the gap analysis was performed against
SCAPE policy elements, to see whether these
bottomup identified control policies allow enough coverage of
the extensively defined data curation policy landscape
of SCAPE project. SCAPE policy elements catalogue
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is two-level with Guidance Policies on the top level
and Policy Elements on the granular level. An example
of Guidance Policy is Authenticity Policy that breaks
down to Integrity, Reliability and Provenance as policy
elements. Hence control policies in Data Integrity
checks category from Table 1 correspond to Integrity
policy element of Authenticity Policy in the SCAPE
policy elements catalogue.
      </p>
      <p>One noticeable gap discovered through this mapping
exercise is the Digital Object lifecycle which was paid
due attention to in SCAPE policy landscape but is
missing in the current EUDAT considerations. This gap may
be hard to address as EUDAT is a collaborative project
that accumulates data from a large variety of research
communities with a wide range of digital object types
and lifecycles. However, this discovery should inform
the future operation of EUDAT services so that they
could meet all reasonable (and multi-aspect)
requirements for data curation and long-term digital
preservation.</p>
    </sec>
    <sec id="sec-6">
      <title>5 Conclusion and further work</title>
      <p>Analysis of data curation requirements of two use cases:
HERBADROP and GEOFON has been performed,
coupled with the retrospective review of the elaborated
data curation policies from a dedicated EU project
(SCAPE) and practical (machine-executable) policies
that were the output of the dedicated RDA working
group.</p>
      <p>A set of granular control policies have been
identified as candidates for implementation in two use cases,
and a gap analysis of these policies has been performed
against the SCAPE catalogue of policy elements. A
similar gap analysis should be performed against the
RDA practical policies catalogue, in order to see what
existing iRODS implementations can be reused for the
creation of machine-executable policies in EUDAT
B2SAFE service.</p>
      <p>After the set of identified policies is applied in the
two use cases that have been involved in their
formulation, the same policy framework should be applied in a
larger number of research communities associated with
EUDAT through its pilot programme.</p>
      <p>
        The scope of projects and initiatives in data curation
and long-term digital preservation can be extended
beyond SCAPE and RDA working groups; this
specifically applies to popular functional models of digital
preservation like OAIS [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] that we feel have not been
thoroughly evaluated so far for their potential
application in EUDAT.
      </p>
      <p>The major result of these works is going to be a
conceptually and terminologically consistent catalogue
of machine-executable policies for EUDAT services
that will be explicitly mapped to requirements of the
participating research communities, as well as to mature
data policy frameworks developed by EU projects and
international collaborations dedicated to data curation
and long-term digital preservation.</p>
      <p>The EUDAT data policies catalogue will serve then
both as guidance for machine-executable policy
implementations and as a validation tool to ensure the
compliance of EUDAT CDI services to high-level policies
of data curation and long-term digital preservation. This
should allow to promote certain EUDAT platforms such
as B2SAFE from their current status of “bit-level” data
management solutions topolicy-driven services where
the actual set of policies can be configured according to
a particular use case.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This work is supported by EUDAT 2020 project that
receives funding from the European Union’s Horizon
2020 research and innovation programme under the
grant agreement No. 654065. The views expressed are
those of authors and not necessarily of the project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>EUDAT</given-names>
            <surname>Collaborative Data</surname>
          </string-name>
          <article-title>Infrastructure</article-title>
          . https://www.eudat.eu/eudat-cdi
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>[2] SCAPE: Scalable Preservation Environments</article-title>
          . http://scape-project.eu/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Research</given-names>
            <surname>Data Alliance Practical</surname>
          </string-name>
          Policy Working Group. https://www.rd
          <article-title>-alliance.org/groups/ practical-policy-wg</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>SCAPE</surname>
          </string-name>
          <article-title>Catalogue of Preservation Policy Elements</article-title>
          . http://scape-project.eu/wp-content/ uploads/2014/02/SCAPE_D13.
          <article-title>2_KB_V1.0</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>[5] EPOS: European Plates Observing System</article-title>
          . https://www.epos-ip.org/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>[6] CINES: French National IT Center for Higher Education and Research</article-title>
          . https://www.cines.fr/en/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Hanka</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kind</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <source>The GEOFON Program. Annals of Geophysics</source>
          ,
          <volume>37</volume>
          (
          <issue>5</issue>
          ), Nov.
          <year>1994</year>
          .
          <article-title>ISSN 2037-416X</article-title>
          . doi:
          <volume>10</volume>
          .4401/ag-4196
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>GEOFON</given-names>
            <surname>Data Centre</surname>
          </string-name>
          (
          <year>1993</year>
          )
          <article-title>: GEOFON Seismic Network. Deutsches GeoForschungsZentrum GFZ</article-title>
          . Other/Seismic Network. doi:
          <volume>10</volume>
          .14470/TR560404
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Practical</given-names>
            <surname>Policy Implementations Report</surname>
          </string-name>
          . http://dx.doi.org/10.15497/83E1B3F9-7E17
          <string-name>
            <surname>-</surname>
          </string-name>
          484A
          <string-name>
            <surname>-A466-B3E5775121CC</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hanka</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saul</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harjadi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Fauzi and GITEWS Seismology Group:
          <article-title>Real-time Earthquake Monitoring for Tsunami Warning in the Indian Ocean and Beyond, Nat. Hazards Earth Syst</article-title>
          . Sci.,
          <volume>10</volume>
          , pp.
          <fpage>2611</fpage>
          -
          <lpage>2622</lpage>
          (
          <year>2010</year>
          ). doi:
          <volume>10</volume>
          .5194/nhess-10-
          <fpage>2611</fpage>
          -2010
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <article-title>iRODS: Integrated Rule-Oriented Data System</article-title>
          . https://irods.org/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Haston</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chagnoux</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dugénie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Herbadrop - Long-term Preservation of Herbarium Specimen Images</article-title>
          .
          <article-title>Proc. of the second Eudat User Forum</article-title>
          . Rome (
          <year>2016</year>
          ). https://www.eudat.
          <article-title>eu/communities/ long-term-preservation-of-herbarium-specimenimages</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Dugénie</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chagnoux</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>EUDAT Data Pilot Herbadrop. Second Interim Herbadrop Data Pilot report (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <article-title>Digitarium: Service Centre for High Performance digitization</article-title>
          . http://digitarium.fi/en
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <article-title>DigiWeb+digitization platform</article-title>
          . http://digiweb. digitarium.fi/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Elasticsearch</given-names>
            <surname>Search</surname>
          </string-name>
          and Analytics Engine. https:// www.elastic.co
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <article-title>SeedLink Protocol and Tools Overview</article-title>
          . http:// ds.iris.edu/ds/nodes/dmc/services/seedlink/
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <article-title>Reference Model for an Open Archival Information System (OAIS), Recommended Practice</article-title>
          , CCSDS
          <volume>650</volume>
          .
          <article-title>0-M-2 (Magenta Book)</article-title>
          .
          <source>Issue 2</source>
          ,
          <year>June 2012</year>
          .
          <article-title>CCSDS (The Consultative Committee for Space Data Systems)</article-title>
          , Washington DC (
          <year>2012</year>
          ).
          <article-title>EUDAT services</article-title>
          . https://www.eudat. eu/servicessupport
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <article-title>EUDAT services</article-title>
          . https://www.eudat.eu/servicessupport
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>