<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Adding Biodiversity Datasets from Argentinian Patagonia to the Web of Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcos Zarate</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>German Braun</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Fillottrani</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro para el Estudio de Sistemas Marinos, Centro Nacional Patagonico (CESIMAR-CENPAT)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Comision de Investigaciones Cient cas de la provincia de Buenos Aires (CIC)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Consejo Nacional de Invenstigaciones Cient cas y Tecnicas (CONICET)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universidad Nacional de la Patagonia San Juan Bosco (UNPSJB)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universidad Nacional del Comahue (UNCOMA)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Universidad Nacional del Sur (UNS)</institution>
          ,
          <country country="AR">Argentina</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this work we present a framework to publish biodiversity data from Argentinian Patagonia as Linked Open Data (LOD). These datasets contains information of biological species (mammals, plants, parasites, among others) have been collected by researchers from the Centro Nacional Patagonico (CENPAT), and have initially been made available as Darwin Core Archive (DwC-A) les. We introduce and detail a transformation process and explain how to access and exploit them, promoting integration with other repositories.</p>
      </abstract>
      <kwd-group>
        <kwd>Biocollections</kwd>
        <kwd>Darwin Core</kwd>
        <kwd>Linked data</kwd>
        <kwd>RDF</kwd>
        <kwd>SPARQL</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Animal, plant and marine biodiversity comprise the \natural capital" that keeps
our ecosystems functional and economies productive. However, since the world
is experiencing a dramatic loss of biodiversity [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ], an analysis about its impact
is being done by digitising and publishing biological collections [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. To this end,
the biodiversity community has standardised shared common vocabularies such
as Darwin Core (DwC) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] together with platforms as the Integrated Publishing
Toolkit (IPT) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] aiming at publishing and sharing biodiversity data. As a
consequence, the biodiversity community now have hundreds of millions of records
published in common formats and aggregated into centralised portals.
Nevertheless, new challenges emerged from this initiative for e ectively using such a large
volume of data. In particular, as the number of species, geographic regions, and
institutions continue growing, answering questions about the complex
interrelationships among these data become increasingly di cult. The Semantic Web
(SW) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] provides possible solutions to these problems by enabling the Web of
Linked Data (LD) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], where data objects are uniquely identi ed and the
relationships among them are explicitly de ned. LD is a powerful and compelling
approach for spreading and consuming scienti c data. It involves publishing,
sharing and connecting data on the Web, and o ers a new way of data
integration and interoperability. The driving force to implement LD spaces is the RDF
technology. Moreover, there is an increasing recognition of the advantages of LD
technologies in the life sciences [
        <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
        ].
      </p>
      <p>
        In this same direction, CENPAT1 has started to publicly share its data
under Open Data licence.2 Data are available as Darwin Core Archive (DwC-A)
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which are a set of les for describing the structure and relationships of the
raw data along with metadata les conforming the DwC standard. Nevertheless,
the well-known IPT platform focuses on publishing content in unstructured or
semi-structured formats but reducing the possibilities to interoperate with other
datasets and make them accessible for machines. To enhance this approach, we
present a transformation process to publish these data as RDF datasets. This
process uses OpenRe ne [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for generating RDF triples from semi-structured
data and de ne URIs. It also uses GraphDB [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], previously known as OWLIM
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], for storing, browsing, accessing and linking data with external RDF
datasets. Along this process, we follow the stages de ned in the LOD Life-Cycle
proposed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We claim that this work is an opportunity to exploit data from
biodiversity in Argentina because they had been never published as LOD.
      </p>
      <p>This work is structured as follows. Section 2 describes the main features of
the datasets selected and their relationships with DwC. Section 3 describes the
transformation process to RDF, while section 4 presents its publication and its
access. Section 5 shows the framework to discover links to other datasets. Next,
section 6 presents the exploitation of the dataset. Finally, we draw conclusions
and suggest some future improvements.
2</p>
      <p>CENPAT Data Sources
In this section, before describing our datasets, we brie y explain the DwC
standard and DwC-A, which these datasets are based on.
2.1</p>
      <p>
        Darwin Core Terms and Darwin Core Archive
DwC [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is a body of standards for biodiversity informatics. It provides stable
terms and vocabularies for sharing biodiversity data. DwC is maintained by
TDWG3 (Biodiversity Information Standards, formerly The International
Working Group on Taxonomic Databases). Its terms are organised into nine categories
(often referred to as classes ), six of which cover broad aspects of the biodiversity
domain. Occurrence refers to existence of an organism at both particular place
and time. Location is the place where the organism were observed (normally
a geographical region or place). Event is the relationship between Occurrence
and Location and register protocols and methods, dates, time and eld notes.
1 http://www.cenpat-conicet.gob.ar/
2 https://creativecommons.org/licenses/by/4.0/legalcode
3 http://www.tdwg.org/
Finally, Taxon refers to scienti c names, vernacular names, etc. of the organism
observed. The remaining categories cover relationships to other resources,
measurements, and generic information about records. DwC also makes use of Dublin
Core terms [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], for example: type, modi ed, language, rights, rightsHolder,
accessRights, bibliographicCitation, references.
      </p>
      <p>
        In the same direction, Darwin Core Archive (DwC-A) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a biodiversity
informatics data standard that makes use of the DwC terms to produce a
single, self-contained dataset and thus sharing both species-level (taxonomic)
and species-occurrence data. Moreover, each DwC-A includes these les. Firstly,
the core data file (mandatory) consists of a standard set of DwC terms
together with the raw data. This le is formatted as elded text, where data records
are expressed as rows of text, and data elements (columns) are separated with a
standard delimiter such as a tab or comma. Its rst row speci es the headers for
each column. Secondly, the descriptor metafile de nes how the core data le
is organised and maps each data column to a corresponding DwC term. Lastly,
the resource metadata provides information about the dataset itself such as its
description (abstract), agents responsible for authorship, publication and
documentation, bibliographic and citation information, collection method, among
others.
2.2
      </p>
      <p>
        Dataset Features
The datasets analysed belong to CENPAT and are available as DwC-A in an
IPT server from this institution. They include collections of marine, terrestrial,
parasites and plant species mainly registered from several points of the
Argentinian Patagonia. Data are generated in di erent ways: some of them by means
of electronic devices placed in di erent animals to study environmental variables,
while others are observations of species in their natural habitat or species
studied in laboratories. To ensure the quality of these data, the records have been
structured according to the procedure described in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Up to May 2017, CENPAT owns 33 datasets representing about 273.419
occurrence records, where 80% of them have been also georeferenced. Some of
these collections contain unique data never published because of the age of the
records (1970s). As a consequence, making this information available as LOD
is so important for researchers, who are studying species conservation and the
impact of man in biodiversity along the last years [
        <xref ref-type="bibr" rid="ref16 ref17">16,17</xref>
        ].
3
      </p>
      <p>
        Linked Data Creation
Publishing data as LD involves data cleaning, mapping and conversion processes
from DwC-A to RDF triples. The architecture of such a process is shown in Fig. 1
and has been structured as described in the following subsections.
The DwC-A are manually extracted from the IPT repository and their
occurrences les (occurrence.txt) are processed using OpenRe ne tool [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. There,
occurrences are cleaned and converted to standardised data types such as dates,
numerical values, etc. and empty columns are removed. OpenRe ne also allows
adding reconciliation services based on SPARQL endpoints, which return
candidate resources from external datasets to be matched to elds in the local datasets.
In our process, we use DBpedia [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] endpoint4 to reconcile the Country column
with the dbo:country resource in DBpedia, the link between the resources is
made through the property owl:sameAs. After that, if the reconciliation is done,
we create a new column for the corresponding URI of the resource. In particular,
we add the column named dbpediaCountryURI for the original Country.
      </p>
      <p>Another reconciliation service5 used, it was based on a taxonomic database
Encyclopedia of Life (EOL)6 which allows to reconcile accepted names in EOL
database. Speci cally, the reconciliation is applied to the column scientificName
so that we create a new column named EOL page for the EOL page describing
the specie. Unfortunately, this whole process is time-consuming because not all
values are automatically matched and thus ambiguous suggestions must be xed.
Moreover, in this phase only two columns have been possible to reconcile because
the process returns unsuitable results using DBpedia services some columns like
institutionCode or locality.
4 https://dbpedia.org/sparql
5 http://iphylo.org/~rpage/phyloinformatics/services/reconciliation_eol.</p>
      <p>php
6 http://www.eol.org/
3.2</p>
      <p>
        RDF Schema Alignment and URI De nition
After cleaning and reconciling, data are converted to RDF triples using RDF
Re ne7, which is an extension of OpenRe ne tool. RDF Re ne allows users to
go through a graphical interface describing the RDF scheme alignment skeleton
to be shared among di erent datasets. The RDF skeleton speci es the subject,
predicate and the object of the triples to be generated. The next step in the
process is to set up pre xes. Since datasets include localities, locations and research
institutes, we set up pre xes for well-known vocabularies such as the W3C Basic
Geo ontology [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], Geonames [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], DBpedia, FOAF [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], Darwin-SW [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] for
establishing relationships among DwC classes and Taxon Concept.8 Table 1 shows
the pre xes used.
      </p>
      <p>In order to generate URI for each resource, in this approach we used GREL
(General Re ne Expression Language) also provided by OpenRe ne, the general
structure of the URIs is described below:</p>
      <p>http://[base uri]/[DwC class]/[value]
where: [base uri] is the one speci es in Table 1, [DwC class] is the respective
DwC class and [value] is the value of the cells in the le of occurrences. It
is also important to note that the generated URIs are instances of the classes
de ned in the DwC standard. Finally, the resulting RDF triple for an occurrence
is:
SUBJECT : &lt; base_uri / occurrence / f6bbf85d -85 ea -4605 -87 fa - d81aca73a1cd &gt;
PREDICATE : rdf : type
OBJECT : dwc : Occurrence
7 http://refine.deri.ie/
8 http://lod.taxonconcept.org/ontology/txn.owl
f
:n :v
c -o</p>
      <p>n o
a e
m r
e a</p>
      <p>t
E r
v y
e C
n o
t d
D e
a
t
y it
o m
r E
I v
n e
s n
t t
i
t</p>
      <p>D
u a
ito te
n
C
o
d
e
P -1
A 0 i
T -2
-C "2 ^
O^</p>
      <p>^
N^
IC sx s
E :
T d t</p>
      <p>r
a i
.</p>
      <p>x i
d a
d :s l</p>
      <p>d ^
ed :d ^
c e x
m ic s
m :d</p>
      <p>s
a t
l r
i
n
g
h :
a c
s i
E e
O t
L i</p>
      <p>e
en f
g u
N a c e a</p>
      <p>R P c</p>
      <p>N
u l
m C e</p>
      <p>o
b u e r</p>
      <p>m d
e n
r t a
r
k
R co eg am
p p
l e
e c
to i
^" e
^
x ^
sd ^
:
s s
t
irn :s
t
x 8 e
d 6 u
3 s
9 ,
" 1
g ir ^</p>
      <p>7
n ^ 5
g x "8
sd ^
: ^
ts x
r s
i
n d
g :s
t
r
i
n
g
e</p>
      <p>s
o a u l
n t
g it
i u
m C e</p>
      <p>R co
o
b u e r
tu d e n
d e r t a
e rk
s
a
m
m d e
d v t</p>
      <p>e r
ed rb y laL laL
B a
m ta i
l
o i
v r
d r
g u
N a c e e
en f</p>
      <p>p i</p>
      <p>R ag cN
m i
r d c d d aC ind cco ab E s p o k g fa lc C
ceo cw oun ice ce s O ice yh red ign en m a o
r :
u is</p>
      <p>O</p>
      <p>L n l
t
u r d s ly
m
o
m
u i ss lu
m
n
s
u
s
e
d
u
t
i
o
n
C
o
d
e
e
t
ffoa :cd d d d d C ht from lcou abT
w w w w l</p>
      <p>a t
L c c c c s p</p>
      <p>m l
D :E :</p>
      <p>O :T s s</p>
      <p>DO is
/</p>
      <p>c
b/ i
t I e</p>
      <p>s
F . m
m s c</p>
      <p>e l
p c a</p>
      <p>U
R th
T a
h i
e n
o s
nd sse
p c
a o
s s
h p
b o o</p>
      <p>w n
p l
e e s d
n s
_
h th ign</p>
      <p>p to
r ow e
e
fi s</p>
      <p>r
ne on op th
_ l</p>
      <p>y re e
s</p>
      <p>t c
c th ie a
r
i e s t
p</p>
      <p>e
m u g
t s o
s a e r
/ i
r n
d ie
a s
d p nd fo
f
_s ro
p a t</p>
      <p>n h
ek re</p>
      <p>e
l t e
t i
n ,
j. se p</p>
      <p>e x D
o s am w</p>
      <p>C
s e le s
on th o a</p>
      <p>t
e f n
c th ad
om e r</p>
      <p>d
p ilt .</p>
      <p>a o
se ls re
ch o o</p>
      <p>b ev</p>
      <p>N n e s
E o P c r h
c t</p>
      <p>a o
TA se te w</p>
      <p>s
p c
i
t e
h a
e t
r - .</p>
      <p>G t</p>
      <p>I x
t e L</p>
      <p>U AI .</p>
      <p>R - o</p>
      <p>I L r
i i
: :
a l
g o
e c
n t
t/ i
a ta
e c
n u o
s t
o e
c n t f</p>
      <p>/ 6 n M
en f/ dw bb ce ir
p 6
/ o
r n
/ re /
b c f
a
t
-c f
b a
8
o 5
n d a
i
c - e
te 85 ce -a
&gt; e
m d
m 5
8 f
n
p 6
8 f u
5 6 n
b g
b a
85 le
d o</p>
      <p>n
8 i
4</p>
      <p>5 n
0 e a
5 a &gt;
8 -4
a
4 a
6 t
0</p>
      <p>v
5 1
-8 .
7
f
a
d
8
1
&gt; -d -8
8 7
1 f
a a
c
a d
7 8
3 1
a a
1 c
c a
d 7
&gt; 3a
1
c
d
&gt;</p>
      <p>Publishing and Accessing Data
The transformed biodiversity data have been published, and can to be accessed,
through GraphDB. GraphDB is a highly e cient and robust graph database
with RDF and SPARQL support. It allows users to explore the hierarchy of
RDF classes (Class hierarchy), where each class can be browsed to explore
its instances. Similarly, relationships among these classes also can be explored
giving an overview about how many links exist between instances of the two
classes (Class relationship). Each link is a RDF statement where its subject
and object are class instances and its predicate is the link itself. Lastly, users also
can explore resources providing URIs representing any of the subject, predicate
or object of a triple (View resource).</p>
      <p>Finally, Fig. 2 shows the resulting graph for the description of a southern
elephant seal skull, which is part of the CENPAT collection of marine
mammals and contains information about where has been found, who has been
collected for, sex and scienti c name, among others. Another way to access the
same information is to explore the View resource in the GraphDB repository
http://crowd.fi.uncoma.edu.ar:3333/resource/find for the speci c
occurrence f6bbf85d-85ea-4605-87fa-d81aca73a1cd, while the serialization of the
complete graph in Turtle syntax can be consulted in.9
9 https://github.com/cenpat-gilia/CENPAT-GILIA-LOD/blob/master/rdf/graph.</p>
      <p>ttl, accessed at September 2017
Interlinking other datasets in a semi-automated way is crucial aiming at
facilitating data integration. In this context, OpenRe ne reconciliation service
is able to match some links to DBpedia, but since it is still limited, our
process should use more powerful tools to discover links to other datasets. For
this task, our approach preliminarily integrate SILK framework10 that uses
Silk-Link Specification Language (Silk-LSL) to express heuristics for
deciding whether a semantic relationship exists between two entities. For interlinking
species between DBpedia and our dataset, we used Levenshtein distance a
comparison operator that evaluates two inputs and computes the similarity based on
a user-de ned distance measure and a user-de ned threshold. This comparator
receives as input two strings dbp:binomial (Binomial nomenclature in DBpedia)
and the combination of dwc:genus + dwc:specificEpithet (the concatenation
of these two de nes the scienti c name of the species). The Levenshtein distance
comparator was set up with &lt;Thresholds = "0.0" and Weight = "1"&gt;. After
the execution, SILK discovered 15 links to DBpedia with an accuracy of 100%
and 85 link with an accuracy between 65% and 75%. In this case, we permit
only one outgoing owl:sameAs link from each resource. The complete Silk-LSL
script can be downloaded from.11</p>
      <p>However, although a set of links has been successfully generated, users'
feedback is needed to lter some species wrongly matched by the tool. Finally, we
must identify further candidates for interlinking and tests other properties or
classes from our dataset in order to increase the automatic capabilities of the
framework.
6</p>
      <p>Exploitation
This section shows how the di erent types of observations of species can be
retrieved, complemented with information of another datasets and ltered by
submitting SPARQL queries to GraphDB endpoint. Moreover, it provides some
experiments in R by using the SPARQL12 package. Each SPARQL query in
following examples assumes the pre x de ned in Table 1.</p>
      <p>Total Number of Species in the CENPAT Dataset. The following query
retrieves the species of the dataset. To this end, it includes the scienti c name of
the species and also its amount of occurrences, to execute this query in GraphDB
see.13 The Fig. 3 shows only the rst resulting records.
10 http://silkframework.org/
11 https://github.com/cenpat-gilia/CENPAT-GILIA-LOD/blob/master/SILK/
link-spec.xml, accessed at September 2017
12 https://cran.rproject.org/web/packages/SPARQL/SPARQL.pdf
13 http://crowd.fi.uncoma.edu.ar:3333/sparql?savedQueryName=species-count
SELECT ? scname ( COUNT (? s) AS ? observations )
{? s a dwc : Occurrence .
?s dsw : toTaxon ? taxon .</p>
      <p>? taxon dwc : scientificName ? scname }
GROUP BY ? scname
ORDER BY DESC ( COUNT (? s ))</p>
      <p>Occurrences by Year. The following query allows to observe the temporality
of the occurrences and its results are visualised using R as shown the Fig. 4. The
R script is available in.14
SELECT ? year ( COUNT (? s) as ? count )
{? s a dwc : Event .</p>
      <p>?s dwc : verbatimEventDate ? date }
GROUP BY ( year (? date ) AS ? year )
ORDER BY ASC (? year )</p>
      <p>Conservation Status of Species. Conservation status are de ned by The
IUCN Global Species Programme15 and are taken as a global reference.
Information about the state of conservation is missing in CENPAT datasets so that
14 https://github.com/cenpat-gilia/CENPAT-GILIA-LOD/blob/master/
r-scripts/occurrences-by-year.R, accessed at September 2017
15 http://www.iucnredlist.org/
providing these data linking other RDF datasets is highly desirable. To this end,
the following query capture these missing data using the owl:sameAs property.
The results are shown in Fig. 5, to execute this query in GraphDB, see.16
SELECT ? scname ? eol_page ? c_status
WHERE { ?s a dwc : Taxon .</p>
      <p>?s dwc : scientificName ? scname .
?s txn : hasEOLPage ? eol_page .</p>
      <p>?s owl : sameAs ? resource .</p>
      <p>SERVICE &lt; http :// dbpedia . org / sparql &gt; {</p>
      <p>? resource dbo : conservationStatus ? c_status .}
}
picted in Fig. 6 using R, and the script is available in.17
7 Conclusions and Further Works
In this work we have presented a framework to publish biodiversity data from
Argentinian Patagonia as LOD, which have initially been made available as
16 http://crowd.fi.uncoma.edu.ar:3333/sparql?savedQueryName=</p>
      <p>conservation-status
17 https://github.com/cenpat-gilia/CENPAT-GILIA-LOD/blob/master/
r-scripts/positions-ml.R, accessed at September 2017
Darwin Core Archive les. The aim is to facilitate the access of researchers to
important data and thus giving a valuable support to the scienti c analysis of the
biodiversity. In addition, this work is the rst Argentinian initiative to convert
biodiversity data according to the criteria established by LOD.</p>
      <p>We have detailed the transformation process and explained how to access and
exploit them, promoting integration with other repositories. Moreover, we have
depicted this process using queries extracted from the domain of application.
Such RDF repository is hosted at http://crowd.fi.uncoma.edu.ar:3333/
together with an SPARQL endpoint, in this initial stage we store 202.119 triples.</p>
      <p>
        As future works, we plan to automate some tasks of the process and interlink
with more datasets. Moreover, providing easier SPARQL access for non-skilled
users. Finally, we are analysing other ontologies such as ENVO [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], NCBI [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
and OWL Time [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and working on a suite of complementary ontologies for
describing every aspect of semantic biodiversity.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Craig</given-names>
            <surname>Moritz</surname>
          </string-name>
          , James L Patton,
          <string-name>
            <surname>Chris J Conroy</surname>
          </string-name>
          , Juan L Parra,
          <article-title>Gary C White,</article-title>
          and
          <string-name>
            <surname>Steven R Beissinger.</surname>
          </string-name>
          <article-title>Impact of a century of climate change on small-mammal communities in Yosemite National Park</article-title>
          , USA. Science,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Adriana</given-names>
            <surname>Verges</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter D Steinberg</surname>
            ,
            <given-names>Mark E Hay</given-names>
          </string-name>
          , Alistair GB Poore, Alexandra H Campbell, Enric Ballesteros, Kenneth L Heck, David J Booth,
          <article-title>Melinda A Coleman, and</article-title>
          <string-name>
            <surname>Feary.</surname>
          </string-name>
          <article-title>The tropicalization of temperate marine ecosystems: climate-mediated changes in herbivory and community phase shifts</article-title>
          .
          <source>In Proc. R. Soc. B. The Royal Society</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Malcolm</given-names>
            <surname>Scoble</surname>
          </string-name>
          .
          <article-title>Rationale and value of natural history collections digitisation</article-title>
          .
          <source>Biodiversity Informatics</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>John</given-names>
            <surname>Wieczorek</surname>
          </string-name>
          , David Bloom,
          <string-name>
            <given-names>Robert</given-names>
            <surname>Guralnick</surname>
          </string-name>
          , Stan Blum, Markus Doring, Renato Giovanni, Tim Robertson, and David Vieglais.
          <article-title>Darwin core: An evolving community-developed biodiversity data standard</article-title>
          .
          <source>PLoS ONE</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Tim</given-names>
            <surname>Robertson</surname>
          </string-name>
          , Markus Doring, Robert Guralnick, David Bloom, John Wieczorek, Kyle Braak, Javier Otegui, Laura Russell, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Desmet</surname>
          </string-name>
          .
          <article-title>The GBIF integrated publishing toolkit: facilitating the e cient publishing of biodiversity data on the internet</article-title>
          .
          <source>PLoS One</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>James</given-names>
            <surname>Hendler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ora</given-names>
            <surname>Lassila</surname>
          </string-name>
          , et al.
          <source>The Semantic Web. Scienti c American</source>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Tom Heath, and
          <string-name>
            <surname>Tim</surname>
          </string-name>
          Berners-Lee.
          <article-title>Linked data-the story so far</article-title>
          .
          <source>Semantic services, interoperability and web applications: emerging concepts</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Francois</given-names>
            <surname>Belleau</surname>
          </string-name>
          ,
          <string-name>
            <surname>Marc-Alexandre</surname>
            <given-names>Nolin</given-names>
          </string-name>
          , Nicole Tourigny, Philippe Rigault, and Jean Morissette. Bio2rdf:
          <article-title>Towards a mashup to build bioinformatics knowledge systems</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Jouni</given-names>
            <surname>Tuominen</surname>
          </string-name>
          , Nina Laurenne, and
          <article-title>Eero Hyvonen. Biological Names and Taxonomies on the Semantic Web { Managing the Change in Scienti c Conception</article-title>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>K Doring M Robertson T Remsen</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braak. Darwin Core Archive How-To Guide</surname>
          </string-name>
          .
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Ruben</given-names>
            <surname>Verborgh</surname>
          </string-name>
          and Max De Wilde.
          <article-title>Using OpenRe ne</article-title>
          .
          <source>Packt Publishing Ltd</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Barry</surname>
            <given-names>Bishop</given-names>
          </string-name>
          , Atanas Kiryakov, Damyan Ognyano , Ivan Peikov, Zdravko Tashev, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Velkov</surname>
          </string-name>
          .
          <article-title>OWLIM: A family of scalable semantic repositories</article-title>
          .
          <source>Semantic Web</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Soren Auer, Lorenz Buhmann, Christian Dirschl, Orri Erling, Michael Hausenblas, Robert Isele, Jens Lehmann, Michael Martin,
          <string-name>
            <given-names>Pablo N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          , Bert Van Nu elen, Claus Stadler,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Tramp</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hugh</given-names>
            <surname>Williams</surname>
          </string-name>
          .
          <article-title>Managing the Life-Cycle of Linked Data with the LOD2 Stack</article-title>
          . In
          <source>International Semantic Web Conference (2), Lecture Notes in Computer Science</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. Dublin Core Metadata Initiative et al.
          <article-title>Dublin core metadata element set</article-title>
          ,
          <source>version 1.1</source>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mark J Costello and John Wieczorek</surname>
          </string-name>
          .
          <article-title>Best practice for biodiversity data management and publication</article-title>
          .
          <source>Biological Conservation</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Reed S Beaman and Nico Cellinese</surname>
          </string-name>
          .
          <article-title>Mass digitization of scienti c collections: New opportunities to transform the use of biological specimens and underwrite biodiversity science</article-title>
          .
          <source>ZooKeys</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Ana</surname>
            <given-names>Vollmar</given-names>
          </string-name>
          , James Alexander Macklin, and
          <string-name>
            <given-names>Linda</given-names>
            <surname>Ford</surname>
          </string-name>
          .
          <article-title>Natural history specimen digitization: challenges and concerns</article-title>
          .
          <source>Biodiversity Informatics</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. Soren Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives.
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>The Semantic Web</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>D</given-names>
            <surname>Brickley</surname>
          </string-name>
          . W3C Semantic Web Interest Group:
          <article-title>Basic Geo (WGS84 lat/long)</article-title>
          <string-name>
            <surname>Vocabulary</surname>
          </string-name>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>Marc</given-names>
            <surname>Wick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B</given-names>
            <surname>Vatant</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B</given-names>
            <surname>Christophe</surname>
          </string-name>
          .
          <article-title>Geonames ontology</article-title>
          . http: // www. geonames. org/ ontology , accessed
          <source>at Sep</source>
          <year>2017</year>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>Dan</given-names>
            <surname>Brickley</surname>
          </string-name>
          and
          <string-name>
            <given-names>Libby</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>The Friend Of A Friend (FOAF) vocabulary speci cation</article-title>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Steven</surname>
            J Baskauf and
            <given-names>Campbell O</given-names>
          </string-name>
          <string-name>
            <surname>Webb.</surname>
          </string-name>
          Darwin-SW:
          <article-title>Darwin Core-based terms for expressing biodiversity data as RDF</article-title>
          .
          <source>Semantic Web</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. Pier Luigi Buttigieg, Evangelos Pa lis, Suzanna E. Lewis,
          <string-name>
            <given-names>Mark P.</given-names>
            <surname>Schildhauer</surname>
          </string-name>
          , Ramona L.
          <string-name>
            <surname>Walls</surname>
            , and
            <given-names>Christopher J. Mungall.</given-names>
          </string-name>
          <article-title>The environment ontology in 2016: bridging domains with increased scope, semantic density, and interoperation</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>Scott</given-names>
            <surname>Federhen</surname>
          </string-name>
          .
          <article-title>The NCBI Taxonomy database</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25. Time Ontology in
          <string-name>
            <surname>OWL</surname>
          </string-name>
          ,
          <year>2006</year>
          . http://www.w3.org/TR/owl-time,
          <source>accessed at September</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>