<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Oceanographic Data Management: Towards the Publishing of Pampa Azul Oceanographic Campaigns as Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcos Zarate</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Rosales</string-name>
          <email>prosales@unpata.edu.ar</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Fillottrani</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Delrieux</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mirtha Lewis</string-name>
          <email>mirthag@cenpat-conicet.gob.ar</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for the Study of Marine Systems</institution>
          ,
          <addr-line>(CENPAT-CONICET)</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centro de Investigaciones y Transferencia Golfo San Jorge</institution>
          ,
          <addr-line>(CONICET)</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Computer Science and Engineering Department</institution>
          ,
          <addr-line>(UNS)</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Electric and Computer Engineering Department</institution>
          ,
          <addr-line>(UNS)</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universidad Nacional de la Patagonia San Juan Bosco</institution>
          ,
          <addr-line>(UNPSJB)</addr-line>
          ,
          <country country="AR">Argentina</country>
        </aff>
      </contrib-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Pampa Azul is a governmental initiative in Argentina that supports research
activities through oceanographic campaigns and promotes interdisciplinary
cooperation between institutes of marine research in areas of national jurisdiction.
The Argentine Continental shelf houses commercial sheries, biodiversity,
hydrocarbon basins and mineral deposits of great economic and ecologic importance.
In particular, the San Jorge gulf6 has ecologic and geographic features that
brought together broad and sustained oceanographic research activities. The
gulf is located in the central region of the Patagonian ocean litoral, where oil
and shing industries coexist along with tourism, which gives rise to ongoing
and acute environmental risks. For this reason, accurate and frequent oceanic
sampling and measurement is of critical social, economic and ecologic
importance. Oceanographic campaigns funded by the Pampa Azul initiative are the
basis of scienti c research at sea, yielding huge amounts of data, highly
heterogeneous in types and formats, and scattered across distributed data repositories.</p>
      <p>Oceanographic research and e cient management of the collected data often
appear to be two widely separated worlds. Data managers consider the careful
collection, management and dissemination of research data as essential for the
e ective use, while researchers consider data management as a merely technical
issue, of little relevance for their interests. Consequently, data management is
often insu ciently planned, if at all, and receives very low priority and budget.
As part of governmental policies towards ocean management, researchers from
Argentine scienti c institutions are required to disseminate their activities and
accomplishments to give greater visibility to the national e orts in this eld.
6 http://www.pampazul.gob.ar/areas-prioritarias/golfo-san-jorge/
For individual researchers, this situation presents a di cult challenge regarding
discovery, access, and integration of data which they need to conduct scienti c
inquiries. As speci c cases of these problems, we can mention (i) data level
conicts caused by di erences that arise in data domains due to multiple possible
representations and similar data interpretations, and (ii) lack of an integrated
knowledge infrastructure which hinders the ability of researchers to analyze
potential discovery scenarios if more than one repository is involved. For example,
they may be interested in knowing if in a marine region there are CTD data
available7, provided by oceanographic campaigns in the region of interest performed
by research vessels not belonging to Pampa Azul.</p>
      <p>
        The Semantic Web (SW) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provides possible solutions to these and other
problems by enabling the web of Linked Data (LD) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which is a methodology
for publishing data and metadata in a structured format in a way such that links
may be created and exploited between objects. The key enabling components are
URIs, HTTP, the Resource Description Framework (RDF), and the SPARQL
Protocol and RDF Query Language (SPARQL).
      </p>
      <p>
        In this short paper we present the initial steps in the creation of a
oceanographic linked dataset using information from the oceanographic campaigns
of Pampa Azul. To achieve consistency, discoverability and make the
datasets readable by machines and humans, we use di erent controlled vocabularies
among them NERC Vocabulary Server8 together with the geospatial standard
for the semantic web GeoSPARQL [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the reuse of the ontological design
pattern (ODP) for oceanographic cruises [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In addition we complement the
Pampa Azul information with the oceanographic linked dataset Rolling Deck
to Repository (R2R9).
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Creating the RDF Dataset</title>
      <p>The publication of the dataset involves di erent steps that are described below
and the architecture of such a process can be consulted in10.</p>
      <p>{ Imput data: National Marine Data System (NMDS11) is a web platform
that allows publishing datasets of oceanographic campaigns that was sampled
in the Argentine sea. These datasets are composed of (i) metadata of the
oceanographic campaigns (name of the campaign, vessel, dates, people and
institutions involved, geographical coverage among others.) in XML format,
and (ii) data recorded by the vessel in its trajectory, which contains the
information of the measured variables (pressure, salinity, temperature, depth,
7 CTD: instrument used to measure the conductivity, temperature, and pressure of
seawater.
8 https://www.bodc.ac.uk/resources/vocabularies/vocabulary_search/P01/
9 http://data.rvdata.us/
10 https://github.com/cenpat/pa-lod/blob/master/images/framework-PA.png
accessed at April 2018
11 http://www.datosdelmar.mincyt.gob.ar/index.php
positions where the variable was sampled among others) and additionally
contains information of the equipment that was used to sample (CTD,
Termosalinometer, etc.).
{ Data extraction and cleaning: Metadata and data of campaigns are
manually extracted from the NMDS repository and their content are
processed using OpenRe ne tool12. There, the columns are cleaned and
converted to standardised data types such as dates, numerical values, etc. and
empty columns are removed.
{ URI strategy: Currently URIs for the resources belonging to oceanographic
campaigns follow the pattern:</p>
      <p>http://data.pa.gob.ar/lod/ftypeg/fconceptg/fIDg
The domain only be used for the publication of pampa azul information
and not include the name of any organization, as they may evolve over time.
ftypeg can take any of the following values: resource for the HTTP URI of
a resource, and page and data for that resource's HTML and RDF
documents respectively. fconceptg gives a hint as to what this resource is
about by referring to the class to which that resource belongs, for example,
Cruise, Dataset, Person etc. fIDg for the unique identi ers we use the
ones provided in the original datasets, normally identi ed with a Universally
Unique Identi er (UUID).
{ Conversion to RDF: Data are converted to RDF triples using RDF
Rene13 that allows users to go through a graphical interface describing the
RDF scheme skeleton which speci es the subject, predicate and the object
of the triples to be generated. The next step in the process is to set up
pre xes. Since datasets include localities, locations and research institutes,
we set up pre xes for well-known vocabularies such as FOAF, Dublin Core,
NERC parameter codes and GeoSPARQL. To see the resulting graph after
the conversion see the link14.
{ RDF storage: The transformed data have been published, and can to be
visualised through GraphDB which is a highly e cient and robust graph
database with RDF and GeoSPARQL support. It allows users to explore the
hierarchy of RDF classes, relationships among these classes, etc.
3</p>
      <p>Use Case: Complementing Pampa Azul Information
With R2R
The following use case explores the R2R oceanographic linked dataset using a
federated query to retrieve all the trajectories that exist in the R2R dataset and
that are within a polygon de ned in the FILTER clause, this polygon de nes the
exclusive Argentine economic area, so this query is interesting since it allows to
know which cruises traveled that region at some time.
12 http://openrefine.org/
13 http://refine.deri.ie/
14 https://github.com/cenpat/pa-lod/blob/master/images/pa-graph.png
cessed at April 2018
acPREFIX geosparql : &lt;http :// www . opengis . net / ont / geosparql #&gt;
PREFIX geof : &lt;http :// www . opengis . net / def / geosparql / function /&gt;
PREFIX sf : &lt;http :// www . opengis . net / ont / sf #&gt;
SELECT ? track
WHERE {
SERVICE &lt;http :// data . rvdata . us / sparql &gt; {
? track a sf : LineString .
? track geosparql : asWKT ? gWKT
FILTER ( geof : sfWithin (? gWKT , " POLYGON (( -63.80859375 -41.45713698292349 ,
-47.109375 -41.457136982923494 , -47.109375 -50.9151558824997 , -63.80859375
-50.9151558824997 , -63.80859375 -41.457136982923494)) " ^^ sf : WktLiteral )) }
}
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion and Future Work</title>
      <p>In this short paper we presented an overview of our initial e orts to create a
Linked Oceanographic Dataset, reusing Ontological Design Patterns and speci c
vocabularies of this domain. In order to test our dataset we extracted metadata
in XML format from NMDS. In this initial stage (April 2018) our platform stored
618K RDF triples with a total of ten classes instantiated. Also for the user to
exploit the dataset we de ne SPARQL queries that can be accessed through the
link15. Finally, the user can visually explore the dataset, accessing the
following link to the GraphDB interface16 (user: pauser password: pauser). We have
shown that RDF can be used to represent metadata about oceanographic
campaigns of Pampa Azul initiative in a useful way, but there is still a lot of fruitful
work to be done. Particularly as future work, we need to develop information
visualization interfaces that allow non-expert users to explore the data. For this
we have explored map4rdf17 a mapping and faceted browsing tool for exploring
and visualizing RDF datasets enhanced with geometrical information.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Tim</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>James</given-names>
            <surname>Hendler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ora</given-names>
            <surname>Lassila</surname>
          </string-name>
          , et al.
          <source>The Semantic Web. Scienti c American</source>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Bizer</surname>
          </string-name>
          , Tom Heath, and
          <string-name>
            <surname>Tim</surname>
          </string-name>
          Berners-Lee.
          <article-title>Linked data-the story so far</article-title>
          .
          <source>Semantic services, interoperability and web applications: emerging concepts</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Robert</given-names>
            <surname>Battle</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dave</given-names>
            <surname>Kolas</surname>
          </string-name>
          .
          <article-title>Enabling the geospatial semantic web with parliament and geosparql</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>3</volume>
          (
          <issue>4</issue>
          ):
          <volume>355</volume>
          {
          <fpage>370</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Adila</given-names>
            <surname>Krisnadhi</surname>
          </string-name>
          , Robert Arko, Suzanne Carbotte, Cynthia Chandler, Michelle Cheatham, Timothy Finin, Pascal Hitzler, Krzysztof Janowicz, Thomas Narock,
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Raymond</surname>
          </string-name>
          , et al.
          <article-title>An ontology pattern for oceanographic cruises: Towards an oceanographer's dream of integrated knowledge discovery</article-title>
          .
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>15 https://github.com/cenpat/pa-lod/tree/master/SPARQL accessed at April 2018</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>16 http://web.cenpat-conicet.gob.ar:7200/login</mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>17 http://oegdev.dia.fi.upm.es/projects/map4rdf/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>