<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Normalizing Resource Identifiers using Lexicons in the Global Change Information System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Linking Earth Science Identifiers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Communities</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian Duggan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>bduggan@usgcrp.gov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert E. Wolfe</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rewolfe@usgcrp.gov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Linked Data</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Co-reference</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Gerald Manipon</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Justin C. Goldstein</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>NASA Goddard Space Flight Center 8800</institution>
          <addr-line>Greenbelt Rd Greenbelt, MD 20771</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>NASA/Jet Propulsion Laboratory 4800 Oak Grove Dr Pasadena</institution>
          ,
          <addr-line>CA 91011</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Steven Aulenbach</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>US Global Change Research Program 1717 Pennsylvania Ave NW Washington</institution>
          ,
          <addr-line>DC 20006</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University Corporation for Atmospheric Research P.</institution>
          <addr-line>O. Box 3000 Boulder, CO 80307</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Earth Science informatics involves collaboration between multiple groups of people with diverse specializations and goals, often using variations in terminology to refer to common resources. The uniformity of the resource identifiers often does not cross organizational boundaries. Because of this, permanent, widely used, unambiguous identifiers for resources are elusive. We examine real world cases of changing and inconsistent identifiers which inherently work against persistence and uniformity. We also present a solution which mediates factors in these situations; namely the creation of lexicons: mappings of sets of terms to URIs which are curated within the Global Change Information System (GCIS).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>We discuss aspects of the GCIS which facilitate the use
of lexicons: an information model which disambiguates
resources, a RESTful API which provides metadata through
content-negotiation, and a strategy for long term curation of
URIs, including mechanisms for handling changes to URIs
and variations in terms used by different communities while
providing persistent URIs and preserving relationships
between resources.</p>
      <p>We provide working definitions of terms, context s, and
lexicons, and relate them to the practical challenges of
disambiguation and curation. We also discuss the mechanisms
employed and architecture of the GCIS, and how these choices
facilitate representation of persistent identifiers and
mappings of them to identifiers used colloquially within various
earth science communities of practice.
1.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
    </sec>
    <sec id="sec-3">
      <title>Background</title>
      <p>
        The U.S. Global Change Research Program (USGCRP) was
established in 1989 by Presidential Initiative and mandated
by the U.S. Congress in the Global Change Research Act
(GCRA) of 1990 to “assist the Nation and the world to
understand, assess, predict, and respond to human-induced
and natural processes of global change.”[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] The USGCRP
has recently sponsored the creation of the Global Change
Information System (GCIS) to better coordinate and
integrate the use of federal information products on changes in
the global environment and the implications of those changes
for society.
      </p>
      <p>In May, 2014, the USGCRP released the Third National
Climate Assessment (NCA3). This 800 page document,
authored by 300 people, each of which are affiliated with
multiple organizations, has 30 chapters, 161 findings, 290
figures and 3,395 references. The references refer to
publications, including government reports, peer-reviewed scientific
journal articles, and books. The publications are often
supported by datasets that could be based on observations,
measurements, processed or derived data, or model projections.
Model projection datasets are created by runs of models
constrained by scenarios. Observations and measurements
are taken using instruments on platforms. The information
in the NCA3 provided a starting point for the contents of
the GCIS. The GCIS information model includes
representations of reports, chapters, findings, figures, people,
organizations, references, publications, platforms, instruments,
models and scenarios. The GCIS was used to support the
production of the report and the dissemination of the web
version of the report.</p>
    </sec>
    <sec id="sec-4">
      <title>1.2 Motivation</title>
      <p>A key design goal of the GCIS was to disambiguate and
identify distinct resources, and provide references to sources
of information. Other goals included: supporting scientific
traceability and reproducibility, facilitating the creation of
reports such as the NCA3, providing a scalable backend for
rich web versions of the NCA3 and other reports, providing
an API for other types of applications, providing the
ability to run structured queries about disparate types of earth
science information, and facilitating the discovery and
representation of connections between earth science information
managed by independent organizations.</p>
      <p>The GCIS has been implemented as a RESTful API, whose
endpoints are URIs which are part of a knowledge base
formalized by an ontology and distributed using a SPARQL
endpoint to a triple store. The API supports content
negotiation, and the HTML representations form a navigable
web site.</p>
    </sec>
    <sec id="sec-5">
      <title>1.3 Lexicons</title>
      <p>This paper focuses on the concept of lexicons and the
processes and techniques involved in the creation and
maintenance of mappings from persistent URIs to pre-existing
Earth Science identifiers. In particular, we discuss
challenges and techniques for dealing with colloquial identifiers
(terms) which are often specific to communities of practice.
We also discuss our techniques for maintaining long term
persistent identifiers, and working with changing or
inconsistent terms.</p>
    </sec>
    <sec id="sec-6">
      <title>2. RELATED WORK</title>
      <p>
        The general problem of disambiguation of resources has been
known for some time, dating back at least to Leibnitz’s
formulation of the Identity of Indiscernibles in 1686 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
In 2006, the WWW’s Technical Architecture Group addressed
issues surrounding identification and URIs by distinguishing
between resources and information resources. An HTTP
request for a resource can return a 303 (“See Other”) response
which directs a user to an information resource [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The
former in general does not have a representation which can
be transmitted over HTTP, whereas the latter does.
As noted in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a service which provides sufficiently
descriptive information about resources can provide disambiguation
services.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] describes some of the difficulties of using automated
techniques to discover equivalent identifiers. Despite these
difficulties, various sophisticated attempts have been made, such
as [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Besides automatic classification, an
alternative technique has been the creation of a declarative
language for resolving identifier ambiguities [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Once there are unambiguous URIs, if these can be mapped
between RDF datasets, the mapping forms a “linkset” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
which can be distributed, harvested and used. The
representation of linksets can be further refined with careful use
of “owl:sameAs” [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and other relationships.
      </p>
      <p>CrossRef</p>
      <p>Journal
Article</p>
      <p>NOAA
Graphic
USGCRP
Report</p>
      <p>NASA/CNES</p>
      <p>Joint Mission
NASA/JPL</p>
      <p>Instrument
Various phrases have been used to describe this problem: the
co-reference problem, the identity problem, disambiguation,
and more. We chose the phrase “normalization” following
its usage with relational databases and character encodings.</p>
    </sec>
    <sec id="sec-7">
      <title>3. EXAMPLES</title>
    </sec>
    <sec id="sec-8">
      <title>3.1 Traceability</title>
      <p>Figure 2.26 of the Third National Climate Assessment1
depicts Past and Projected Changes to Global Sea Level Rise
(see figure 1). A sequence of resolvable URIs within GCIS
traces this figure to a journal article which used a dataset
which used another dataset, which was captured by an
instrument on a platform. Each step in this sequence has a
permanent URI within GCIS, but also refers to identifiers
outside of GCIS; the journal article has a DOI managed by
a publisher and resolvable using CrossRef, the first dataset
has a URL managed by scientists, the second dataset has an
identifier managed by a NASA Data Archive, the instrument
and platform have identifiers created by the Committee on
Earth Observing Satellites (CEOS). While all these
organizations do provide machine-readable versions of their data,
the identifiers are often curated independently. However,
within each organization, there are terms which
unambiguously identify a particular resource.</p>
    </sec>
    <sec id="sec-9">
      <title>3.2 Identification</title>
      <p>In some situations, there may be intentional uses of different
names by different organizations. For instance, the
“SACD/Aquarius” mission may also be called “Scientific
Application Satellite-D”</p>
      <p>
        SAC-D/Aquarius is a cooperative international
mission between CONAE (Comisi´on Nacional de
1http://data.globalchange.gov/report/nca3/chapter/
2/figure/26
Actividades Espaciales), Argentina, and NASA,
USA. NASA uses the term [SAC-D/Aquarius] for
the mission [..] At CONAE, which provides the
spacecraft, the mission is referred to as Scientific
Application Satellite-D [..] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
This is a clear case in which two different communities
refer to the same resource differently. Such differences can
propagate from narrative descriptions into identifiers found
in serializations of information. When this happens,
techniques for reconciling the identifiers become necessary.
      </p>
    </sec>
    <sec id="sec-10">
      <title>3.3 Synchronization</title>
      <p>Organizations in the domain of Earth Science distribute
data, metadata, and information using a variety of
serializations and interfaces, including:</p>
      <sec id="sec-10-1">
        <title>ECHO 10 ISO 19115 FGDC DIF</title>
        <p>NASA Earth Observing System
(EOS) Clearing House (ECHO)
Geographic Information
Metadata
Federal Geographic Data
Committee Content Standard for
Digital Geospatial Metadata
Directory Interchange Format,
NASA’s Global Change Master
Directory (GCMD)
W3C Data Catalog Vocabulary
Open Archives Initiative Protocol
for Metadata Harvesting
Miscellaneous Serializations</p>
      </sec>
      <sec id="sec-10-2">
        <title>CSV, JSON, YAML</title>
        <p>Information conveyed varies by the standard and by the
implementation. Despite these differences, the defining
characteristics of resources generally remain consistent across
representations. This enables a subset of information to be
brought into the GCIS for the purpose of identification and
disambiguation. In other words, when deciding what to
harvest from other sources, a critical question is: does this piece
of information help to distinguish this resource from other
similar resources?</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>3.4 Communities</title>
      <p>Science teams working on remote sensing missions produce
scientific data and send them to data archives. A primary
concern of archives is to maintain the fidelity of the science
data as they are received. Because of this division of
responsibilities, it’s not clear which community would be better
suited to take on the task of harmonizing identifiers across
data producers. Instead differences in choices of identifiers
may be passed on to end users of the data.</p>
    </sec>
    <sec id="sec-12">
      <title>4. CONCEPTS</title>
    </sec>
    <sec id="sec-13">
      <title>4.1 Terms as Identifiers</title>
      <p>Issues surrounding characters, case folding, encoding, and
strings are often omitted when resources are identified in
narrative situations. This creates ambiguities in
representations and postpones normalization issues. This motivates
our definition of the word “term” which is based on the
Universal Character Set (UCS).</p>
      <p>
        Definition 1. We define a term to be a sequence of
characters from the Universal Character Set (UCS) which is used
as an identifier for a resource by a group of people.
The linked data glossary [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] defines a “term” using the notion
of a controlled vocabulary, and also defines a controlled
vocabulary using the notion of a term. We explicitly make use
of Unicode characters in order to avoid circular definitions
like this.
      </p>
      <p>
        Within the GCIS, terms are encoded in UTF-8 and
potentially normalized with Normalization Form C (NFC). We
note that W3C recommendations for string matching are
still evolving [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
    </sec>
    <sec id="sec-14">
      <title>4.2 Communities of Practice</title>
      <p>Within communities of practice, terms are created and used
as part of the activities and communication between
members of the community. Because of this, disambiguation
of terms across communities is a secondary consideration.
Within a community, a resource can be identified clearly
when it is being referenced in a manner which is consistent
with other similar resources. With this in mind, we group
together terms using the type of resource.</p>
      <p>Definition 2. We define a context to be a set of terms used
to identify resources of the same type.</p>
      <p>In example 3.1, ”mission” and ”instrument” are contexts.
We then define a lexicon by putting together the contexts.</p>
      <p>Definition 3. We define a lexicon to be a set of contexts
used by a particular community.</p>
      <p>Returning to example 3.1, the CEOS lexicon uses ”Mission”
and ”Instrument” contexts to group together terms.</p>
    </sec>
    <sec id="sec-15">
      <title>4.3 GCIDs</title>
      <p>As previously noted, entities in the GCIS are identified uniquely
using a URI. The URI for a particular entity is called a GCIS
Identifier, or GCID.</p>
      <p>Definition 4. The GCID is the URI for an entity in the
Global Change Information System.</p>
      <p>Note that any GCID can be used in SPARQL queries, and
can also be resolved using the GCIS as an endpoint and
using content-negotiation.</p>
      <p>One organization’s “instrument” may be another one’s
“sensor“. One organization’s “platform” may be another one’s
“mission” or a third one’s “source”. We use lexicons to
represent the way NASA’s Physical Oceanography Active Data
Archive Center (PODAAC), ECHO, GCMD and CEOS all
refer to the same resource:
Lexicon | Context | Term | GCID (*) 5.2 System Architecture
------------------------------------------------------- The GCIS architecture incorporates elements of relational
podaac | Source | JASON-1 | /platform/jason-1 and semantic systems: cascading updates, referential
ingccemods || pMriesfsLiaobneIld || J2A8S6ON-1 || //ppllaattffoorrmm//jjaassoonn--11 tegrity, strict type checking and other well-established
feaecho | ShortName | JASON-1 | /platform/jason-1 tures of relational databases are all valuable in maintaining
podaac | Sensor | POSEIDON-2 | /instrument/poseidon-2 the quality of the URIs and their relationships.
ceos | InstrumentId | 182 | /instrument/poseidon-2
(*) under http://data.globalchange.gov
See also: http://data.globalchange.gov/lexicon</p>
    </sec>
    <sec id="sec-16">
      <title>4.4 Linksets</title>
      <p>When a term is an identifier which is part of another triple
store, we can use ”owl:sameAs” to connect the two identifiers
and form a linkset. We treat dbpedia as a lexicon with a
single context (”resource”). An application of this is
writing a federated SPARQL query to compare crowdsourced
information in dbpedia (or wikidata) to authoritative
information from CEOS. This can belp improve the quality of
the data in both places.</p>
    </sec>
    <sec id="sec-17">
      <title>5. IMPLEMENTATION</title>
    </sec>
    <sec id="sec-18">
      <title>5.1 Lexicon Interface</title>
      <sec id="sec-18-1">
        <title>5.1.1 Creating, Updating</title>
        <p>Creating a lexicon involves several steps:</p>
        <sec id="sec-18-1-1">
          <title>1. Identifying a distinct set of terms. 2. Choosing a context. 3. Choosing an existing lexicon or adding a new lexicon. 4. Associating the terms with GCIDs.</title>
          <p>An example of performing step 3 in an HTTP transaction
follows; in this example we are associating “Aqua” a term
used by CEOS, with the GCID “/platform/aqua” and the
context “Mission”:</p>
          <p>PUT /lexicon/ceos/Mission/Aqua
Host: data.globalchange.gov
Content-Type: application/json
{ "gcid" : "/platform/aqua" }
An alternative interface is available which allows the terms
and context to be sent in the payload, rather than the URI.</p>
        </sec>
      </sec>
      <sec id="sec-18-2">
        <title>5.1.2 Querying</title>
        <p>Looking up a term using a lexicon involves sending a GET
request for an IRI containing the context and the term, and
receiving a status code of 303 and a Location header, to
indicate the corresponding GCIS URI, if one exists.
Request:</p>
        <p>GET /lexicon/ceos/Mission/Aqua HTTP/1.1</p>
        <p>Host: data.globalchange.gov
Response:
303 See Other
Location: /platform/aqua
HTTP requests to the GCIS RESTful interface are handled
by querying this relational database. The database contains
simple explicit tables for resources that use natural
identifiers as primary keys. JSON structures are formed using the
names of the columns as names of the keys in the JSON
objects. There is also a generic table which may be thought
of as a parent table for many of the tables (i.e. using table
inheritance).</p>
        <p>Relationships between resources are stored in one of two
ways: 1. Foreign keys between base tables. 2. Relationships
between two entries in the generic table, via a mapping table,
which may be annotated with a semantic relationship.
Triples are generated data from the tables to fill in text
templates which output turtle. The turtle is parsed and
used to populate a triple store. The triple store is rebuilt
weekly; there are no incremental updates.</p>
      </sec>
    </sec>
    <sec id="sec-19">
      <title>5.3 Terms</title>
      <p>When new terms appear, they are captured in the GCIS.
Scripts run periodically and pull information from various
sources. The mechanism for assimilating new terms is to
provide operators with notifications of unmatched terms;
operators then manually associate the new terms with GCIS
resources.</p>
      <p>Example 1. A new satellite achieves orbit. The
information in CEOS reflects this new information. It is pulled into
GCIS, and a new entry is created. A new default GCIS
identifier is created from a descriptive field (but it may be
updated, as described below). The relational database is
populated with information that reflects that data source.
Since no term yet exists for this in the ceos lexicon, a new
one is created which maps the term used by CEOS to the
newly created identifier in the GCIS.</p>
      <p>Example 2. An instrument on a spacecraft begins to
produce new data. A science team processes the data and sends
them to an archive. The archive makes the data available.
In this case, a new term has been created, and it must be
matched to a GCID manually. The new data is ingested,
no match occurs. An operator notices this and manually
associates the new term with an existing GCID.
Example 3. Two organizations use different terms for one
satellite. Both organizations distribute data collected by an
instrument on board the satellite. In this case, there are
two lexicons, one for each organization, and each one has a
context which has terms which map to the same GCID.
5.4 URIs</p>
      <sec id="sec-19-1">
        <title>5.4.1 Creation</title>
        <p>URIs are formed using primary keys in the relational database.
The values of the columns comprising the primary key are
assembled along with the name of the table, into a URI
which corresponds to a row of data in a table. The
representation of a resource may involve joining to other related
tables.</p>
      </sec>
      <sec id="sec-19-2">
        <title>5.4.2 Persistence</title>
        <p>When the value of a primary key column is changed, a
cascading foreign key update will change the values in any
related tables. Also, triggers will change the columns in the
generic (parent) table. Because the templates are based on
information in the database, these changes will
automatically propagate to the semantic representation of the
resource.</p>
        <p>Also, for any changes, a note is written to an audit table
which contains the old identifier and the new one.
When the GCIS receives a request for a resource that does
not exist, it first checks the above audit log. If an entry exists
for the requested identifier, this entry is used to construct a
URL to which a redirect is then returned.</p>
        <p>This mechanism allows for changes to identifiers in the GCIS
while continuing to provide consistent endpoints in the API.</p>
      </sec>
      <sec id="sec-19-3">
        <title>5.4.3 Preserving Mappings</title>
        <p>Changes to the parent table described above also trigger
changes to the lexicon tables. So the complete flow for an
identifier change is:</p>
        <p>API or web form
-&gt; primary key of base table
-&gt; primary key of generic table (cascading update)
-&gt; gcid entry in list of terms (trigger)
-&gt; turtle template (which uses database queries)
-&gt; Triple store (by ingesting the rendered template)
-&gt; SPARQL endpoint</p>
      </sec>
      <sec id="sec-19-4">
        <title>5.4.4 Validation of Existing Information</title>
        <p>Identifying changes to terms is important in order to effect
changes like the ones above. This requires continuous
validation. In order to perform this validation, scripts must
periodically check to see if the terms are still valid. If there
is a mapping from terms to URLs, this can be accomplished
through a simple HEAD request. If not, more advanced
techniques may be necessary.</p>
      </sec>
    </sec>
    <sec id="sec-20">
      <title>6. CONCLUSIONS AND FUTURE WORK</title>
      <p>Lexicons are a practical way of mapping identifiers within
the earth science community to each other, when uniform
identifiers for resources do not exist. The GCIS provides
resolvable URIs which serve to fill the gap between
organizations with varying terms for the same resource. Providing
both a semantic and relational mechanism for storage, and
both a semantic and RESTful API allows the GCIS to have
the advantages of both architectures, namely referential
integrity and backwards compatibility, as well as flexibility
and adaptability. While we have seen some success in using
cross comparisons of data to improve quality of individual
data sources, more work can be done in this area. There
is also work to be done in scaling up the number and type
of data sources. Another future improvement is to provide
useful user interfaces for scalable human disambiguation.</p>
    </sec>
    <sec id="sec-21">
      <title>7. ACKNOWLEDGMENTS</title>
      <p>Thanks to the various people and organizations who have
contributed to the GCIS and the development of its
information model, including Andrew Buddenberg and the
Technical Support Unit at NOAA’s National Climatic Data
Center, Xiaogang Ma and the group at the RPI’s Tetherless
World Constellation, and the University Corporation for
Atmospheric Research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>U.S.</given-names>
            <surname>Public Law</surname>
          </string-name>
          101-
          <volume>606</volume>
          (
          <issue>11</issue>
          /16/90) 104 Stat.
          <fpage>3096</fpage>
          -
          <lpage>3104</lpage>
          , 1990 Global Change Research Act of 1990
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Akar</surname>
          </string-name>
          , T. G. Hala¸c,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dikenelli</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. E.</given-names>
            <surname>Ekinci</surname>
          </string-name>
          .
          <article-title>Querying the web of interlinked datasets using VOID descriptions</article-title>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Booth</surname>
          </string-name>
          .
          <article-title>URIs and the myth of resource identity</article-title>
          .
          <source>In Identity, Reference, and the Web Workshop at the WWW Conference</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dimou</surname>
          </string-name>
          , M. V, E. S, P. Colpaert,
          <string-name>
            <given-names>R.</given-names>
            <surname>Verborgh</surname>
          </string-name>
          , E. Mannens, and
          <string-name>
            <given-names>R. V. D.</given-names>
            <surname>Walle</surname>
          </string-name>
          .
          <article-title>RDF mapping language (rml) a generic language for integrated RDF mappings of heterogeneous data</article-title>
          .
          <source>In Linked Data on the Web, WWW</source>
          <year>2014</year>
          , Seoul, South Korea,
          <year>2014</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Halpin</surname>
          </string-name>
          .
          <article-title>Identity, reference, and meaning on the web</article-title>
          .
          <source>In Proceedings of the Workshop on Identity, Meaning and the Web (IMW06)</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Halpin</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Hayes</surname>
          </string-name>
          .
          <article-title>When owl:sameas isn't the same: an analysis of identity links on the semantic web</article-title>
          .
          <source>In Linked Data on the Web, WWW</source>
          <year>2010</year>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Hyland</surname>
          </string-name>
          , G. Atemezing,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pendleton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Srivatstava</surname>
          </string-name>
          .
          <article-title>Linked data glossary</article-title>
          . Working group note, W3C,
          <year>June 2013</year>
          . http://www.w3.org/TR/2013/ NOTE-ld-glossary-
          <volume>20130627</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaffri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Glaser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I. C.</given-names>
            <surname>Millard</surname>
          </string-name>
          .
          <article-title>Managing URI synonymity to enable consistent reference on the semantic web</article-title>
          .
          <source>In Workshop on Identity, Reference, and the Web (IRSW) at ESWC2008</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kramer. SAC-D (Sat´elite de Aplicaciones</surname>
          </string-name>
          <article-title>Cient´ıficas-D)/Aquarius Mission</article-title>
          . https://directory.eoportal.org/web/eoportal/ satellite-missions/s/sac-d,
          <year>2015</year>
          . [Online; accessed 12-March-2015].
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Lewis</surname>
          </string-name>
          .
          <article-title>Dereferencing HTTP URIs</article-title>
          .
          <article-title>Draft tag finding, W3C</article-title>
          , May
          <year>2007</year>
          . http://www.w3.org/2001/ tag/doc/httpRange-14/
          <fpage>2007</fpage>
          -05-31/HttpRange-14.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Maali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Peristeras</surname>
          </string-name>
          .
          <article-title>Re-using cool URIs: Entity reconciliation against LOD hubs</article-title>
          .
          <source>In Linked Data on the Web, WWW</source>
          <year>2011</year>
          , Seoul, South Korea,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>D. B. Nguyen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hoffart</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Theobald</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          .
          <article-title>Aida-light: High-throughput named-entity disambiguation</article-title>
          .
          <source>In Linked Data on the Web, WWW</source>
          <year>2014</year>
          , Seoul, South Korea,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Phillips</surname>
          </string-name>
          .
          <article-title>Character model for the world wide web: String matching and searching</article-title>
          .
          <source>W3C working draft, W3C</source>
          ,
          <year>July 2014</year>
          . http://www.w3.org/TR/charmod-norm/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>