<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Adding eScience Assets to the Data Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pete Johnston Eduserv Foundation Bath UK</string-name>
          <email>pete.johnston@eduserv.org.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carl Lagoze Cornell University Ithaca</institution>
          ,
          <addr-line>NY</addr-line>
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Design</institution>
          ,
          <addr-line>Standardization</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Herbert Van de Sompel Los Alamos National Laboratory Los Alamos</institution>
          ,
          <addr-line>NM</addr-line>
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Michael L. Nelson Old Dominion University Norfolk, VA</institution>
          <country country="US">USA</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Robert Sanderson University of Liverpool Liverpool</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Simeon Warner Cornell University Ithaca</institution>
          ,
          <addr-line>NY</addr-line>
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <volume>20</volume>
      <issue>2009</issue>
      <abstract>
        <p>Aggregations of Web resources are increasingly important in scholarship as it adopts new methods that are data-centric, collaborative, and networked-based. The same notion of aggregations of resources is common to the mashed-up, socially networked information environment of Web 2.0. We present a mechanism to identify and describe aggregations of Web resources that has resulted from the Open Archives Initiative - Object Reuse and Exchange (OAI-ORE) project. The OAI-ORE speci cations are based on the principles of the Architecture of the World Wide Web, the Semantic Web, and the Linked Data e ort. Therefore, their incorporation into the cyberinfrastructure that supports eScholarship will ensure the integration of the products of scholarly research into the Data Web.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.5.4 [Information Systems]: Hypertext/Hypermedia</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>
        The rapid evolution of computing, networking, and data
capturing technologies, along with advances in data mining
and analysis, are fundamentally changing the way scholarly
research is conducted [
        <xref ref-type="bibr" rid="ref2 ref5">2, 5</xref>
        ]. Although there are di erences
amongst disciplines in their receptivity to change [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], an
increasing number of scholars in the natural sciences, social
sciences, and humanities have adopted new research
methods that are network-based, highly collaborative, and
dataintensive. Because of the central role of vast amounts of data
in these new research methods, there has been increased
attention to sustainable infrastructures for registering,
preserving, and sharing datasets [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        In parallel with this change in research methodology there
has been substantial change in the way that research results
are communicated. With the emergence of the Web,
scholarly publishers, both commercial and learned societies,
almost universally deliver journal papers, conference
proceedings, and monographs via the Web. While Web delivery of
research results has improved their accessibility and
searchability, it represents an evolution of traditional publication
practices rather than a fundamental change in the scholarly
communication paradigm. Even in their digital
manifestations, scholarly publications are mostly textually-based and
static. To date, there are few examples of scholarly
communication that move beyond the dissemination of these
traditional artifacts into a more data-centric,
semanticallylinked, and social network-embedded scholarly
communication model that resembles the profound changes in social,
political, and economic discourse characteristic of Web 2.0.
This radically di erent model would expose process as well
as product [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ], improving opportunities to verify the
reproducibility of research results, and making the full spectrum
of artifacts generated in the scholarly value chain available
for reuse [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ].
      </p>
      <p>
        The deployment of radically new models depends on the
development of basic technical infrastructure, so-called
cyberinfrastructure. This cyberinfrastructure must include a
number of components. These include a means to identify
and cite datasets in the scholarly discourse (e.g., [
        <xref ref-type="bibr" rid="ref1 ref38">38, 1</xref>
        ]),
a standard for identifying scholarly authors to
unambiguously tie them to their creations and improve the quality of
scientometric information (e.g., ResearcherID1 and Digital
Author Identi er2), and standards to allow machine
readability of the products of scholarly process thereby
facilitating computational analysis and extraction of secondary and
tertiary knowledge products. Semantic technologies are an
important component of this cyberinfrastructure, providing
a foundation for open agreements on data formats, metadata
frameworks to describe data, and ontology-based solutions
for formal representation of scienti c knowledge, all of which
are important components of promoting a machine-readable
scholarly record.
      </p>
      <p>
        This paper focuses on one aspect of this
cyberinfrastruc1http://www.researcherid.com/
2http://www.surffoundation.nl/smartsite.dws?ch=
eng&amp;id=13480
ture that arises from the changing nature of publications
that are characteristic of collaborative, data-centric
scholarship. These emerging publications are aggregations of
multiple resources. Such aggregations are already prevalent in
existing scholarly repositories, which commonly o er access to
textual documents in multiple formats, each available from
a di erent network location. But, the changes in scholarship
described above, and especially the need to include data in
the publication process, increases the complexity of these
aggregations and calls for the adoption of a common
approach to handle them. In the remainder of this paper,
we describe our work within Open Archives Initiative -
Object Reuse and Exchange (OAI-ORE), a two-year project
to investigate common methods to handle aggregations of
Web resources that culminated in October 2008 with the
release of the OAI-ORE speci cations [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. These speci
cations were motivated by the resource aggregations common
to scholarly communication. We believe that their generic,
Web-centric approach makes them applicable to use cases in
the Web at large, providing the basis for improved search
results, improved information navigation, and richer services
within browsers for a large class of Web applications.
      </p>
      <p>The OAI-ORE speci cations leverage the principles of the
Architecture of the World Wide Web, the Semantic Web,
and the Linked Data e ort. As a result, future
developments in cyberinfrastructure and scholarly communication
that are based on OAI-ORE will integrate well with the
Web and with the tools, agents and applications that
operate within it. This will make it possible to embed or mash up
the products of scholarship into cyber-learning e orts,
cooperative reference tools such as Wikipedia, and the larger
social discourse that is now characteristic of Web 2.0. The
essence of the OAI-ORE solution to the resource aggregation
problem can be summarized is as follows:</p>
      <p>
        The data model is expressed in terms of the
primitives of Web Architecture and the Semantic Web:
Resources, Representations, URIs and RDF triples.
The central entity in the data model, the Aggregation,
is a Resource that stands for a set of other Resources.
An Aggregation is a Resource with a URI but without
a Representation (we refer to this as a non-document
Resource from now on). This approach is aligned with
the manner in which real-world entities or concepts are
included in the Web via the mechanisms proposed by
the Linked Data e ort [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Another Resource, the Resource Map, has a
Representation that is a description of the Aggregation. The
Resource Map is accessible via the URI of the
Aggregation using the mechanisms de ned for Cool URIs for
the Semantic Web [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ].
      </p>
      <p>The Representation of a Resource Map is a
serialization of the triples that describe the Aggregation. The
speci cation describes RDF/XML, RDFa, and Atom
serialization syntaxes.</p>
    </sec>
    <sec id="sec-3">
      <title>AGGREGATIONS 2. 2.1</title>
    </sec>
    <sec id="sec-4">
      <title>Aggregations in Scholarly Communication</title>
      <p>
        Most institutional repositories [
        <xref ref-type="bibr" rid="ref24 ref31">24, 31</xref>
        ] routinely store and
disseminate relatively simple aggregations, consisting of
multiple access formats (e.g., PDF, HTML, LaTeX) for the same
document. In addition, prototypes exist of applications that
allow authoring, storing, and disseminating more complex
scholarly publications in the form of aggregations [
        <xref ref-type="bibr" rid="ref33 ref42 ref8">8, 33,
42</xref>
        ]. These more complex aggregations may consist of a
textual article, one or more datasets that led to the discoveries
reported in the article, perhaps a visualization of a speci c
state of the dataset, and the software used to generate the
visualization. All constituents of such an aggregation are
distributed on the Web. One notable aspect of these more
complex visions of an aggregate scholarly publication is the
importance of semantic relationships among constituents of
the aggregation. These relationships include citation,
versioning, provenance, commentary, and the like.
      </p>
      <p>Some characteristics of the aggregations that are already
common in scholarship can be illustrated by means of a
document from arXiv.org, a well-known repository of physics,
mathematics, and computer science research results. The
human start page, or \splash page", for this document is
shown in Figure 1. Some aspects of the page relevant to the
resource aggregation problem are highlighted in red
rectangles, each with a number. The meanings of the highlighted
areas are as follows:
1. The URI http://arxiv.org/abs/astro-ph/0601007
of the human start page for the arXiv document.
2. The formats in which the document is available, i.e.</p>
      <p>PostScript, PDF, etc. These are e ectively the
constituents of the aggregation that is the arXiv
document.
3. The title of the arXiv document.
4. The authors of the arXiv document.
5. The creation and last modi cation date of the arXiv
document.
6. Identi ers of resources that are in some manner
comparable to this arXiv document. For example, a version
of this document was later published as an article in a
peer-reviewed journal, and the Digital Object
Identier of that article is shown.
7. The versions of this arXiv document.
8. Links to other arXiv documents in the same collection
(i.e., astro-ph).
9. Citations made by this arXiv document, and citations
it received from other documents.</p>
      <p>This rather simple example highlights the core issues that
OAI-ORE addresses. First, although the URI of the
human start page is commonly used as the URI for the entire
arXiv document, within the Web Architecture that URI only
identi es the page itself, and not the aggregation that is the
arXiv document. The ability to cite, annotate, version, and
associate properties with the aggregation itself relies on it
having a unique identity, distinct from the splash page or
the resources linked from it.</p>
      <p>Second, without the use of (frequently imperfect)
heuristics unique to the speci c human start page, it is not
readable by machines and agents. Because the HTML of this
human start page usually leaves the semantics of hyperlinks
unde ned, a machine agent cannot unambiguously
distinguish between links to constituents (e.g. the PostScript,</p>
      <p>PDF, etc.) of the document and links that point at
information that is clearly outside of the document such as the
navigational aids shown as (8) in Figure 1. Similarly, agents
can not interpret relationships of the document to other
documents, identi ers related to this document, versions of this
document, etc.</p>
      <p>In essence, the problem is that there is no standard way
to describe the constituents or boundary of an aggregation,
or to qualify and identify a resource as being an aggregation.
While a robot could learn the semantics implied by arXiv's
HTML in Figure 1, such \screen scraping" is brittle and not
scalable for applications accessing aggregations in thousands
of di erent repositories, each with their own presentation
idiom.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Integrating Aggregations into the Web</title>
      <p>
        A number of early e orts in cyberinfrastructure, for
example the initial grid architecture [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ] and technologies for
digital libraries, leveraged aspects of the Web infrastructure but
often failed to fully conform with Web Architecture
principles. For example, institutional repositories frequently have
identi er schemes and access protocols distinct from those
existing on the Web at large. As a result, much of their
content is accessible on the Web, but it poorly integrates
with mainstream Web applications and may even be
overlooked by major search engines, unless the search engines
make special accommodations for their protocols and access
schemes.
      </p>
      <p>
        Our prior work on the Open Archives Initiative
Protocol For Metadata Harvesting (OAI-PMH) [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] demonstrates
this problem. OAI-PMH is an interoperability speci cation
released in 2001 aimed at streamlining the process of
incrementally collecting XML metadata (typically bibliographic
metadata) from information systems. It shares many
design characteristics with Atom [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] and is widely adopted in
its targeted community of scholarly repositories. But,
OAIPMH, in contrast to Atom, has not gained broader adoption,
mainly because its architecture is not well aligned with the
Resource/URI/Representation foundations of the Web
Architecture. For example, OAI-PMH clients must construct
a request URI by combining a repository speci c base URI,
the identi er of the item of interest, and a format tag in an
OAI-PMH speci c manner, often preventing general Web
clients that are unaware of the protocol from accessing the
available metadata [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>The Web-centric, resource-centric approach of OAI-ORE
recti es this architectural shortcoming and thereby provides
the foundation for full accessibility of the products of eScience
in the general Web environment. Furthermore, it makes the
solution available to a broader class of Web applications in
which the practice of aggregating resources is quite
common. For example, we accumulate URLs in bookmarks or
favorites lists in our browser, collect photos into sets in
popular sites like Flickr, browse over multiple page documents
that are linked together through \prev" and \next" tags,
and talk about Web sites as if they had some real existence
beyond the set of pages of which they consist. Despite our
frequent use of these aggregations, their existence on the
Web is quite ephemeral because there is no common way
to identify, describe, and hence handle them. This is what
OAI-ORE provides.
3.</p>
    </sec>
    <sec id="sec-6">
      <title>THE OAI-ORE SOLUTION</title>
      <p>
        In this section we describe the various elements of the
OAI-ORE solution to the resource aggregation problem
outlined above. It encompasses an RDF-based data model,
syntaxes for serializing instances of the data model, and
mechanisms for providing HTTP access to those serializations.
Complete details are available through the OAI-ORE
documentation suite [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
      </p>
      <p>
        As noted earlier, this solution is based on the primitives
de ned in the Architecture of the World Wide Web [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] that
de nes a Resource as an item of interest; a URI as a global
identi er for a Resource; and a Representation as a
datastream corresponding to the state of a Resource at the time
its URI is dereferenced via some protocol (e.g. HTTP). In
addition, the solution is grounded in the principles
introduced by the Semantic Web, in which URIs are also used
to identify non-document Resources, such as real-world
entities (e.g. people or cars), or even abstract entities (e.g. ideas
or classes). These non-document Resources have no
Representation to indicate their meaning. OAI-ORE adopts the
following approach, proposed by the Linked Data e ort [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
for obtaining information about those Resources:
Use of HTTP URIs to identify those non-document
Resources;
Publication of another Resource with a Representation
that provides information about the non-document
Resource at a HTTP URI other than the HTTP URI of
the non-document Resource;
Leverage of HTTP mechanisms to allow discovery of
the HTTP URI of the published resource from the
HTTP URI of the non-document resource.
3.1
      </p>
    </sec>
    <sec id="sec-7">
      <title>Data Model</title>
      <p>
        The essence of the RDF-based data model is described
here and is illustrated in Figure 2. The full details are
available in the OAI-ORE Abstract Data Model speci
cation [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>In order to be able to unambiguously refer to an
aggregation of Web resources, a new Resource is introduced that
stands for a set or collection of other Resources. This new
Resource, named an Aggregation, has a URI just like any
other Resource on the Web. And, since an Aggregation is
a conceptual construct, it is a non-document Resource that
does not have a Representation.</p>
      <p>Following the Linked Data guidelines, another Resource
is introduced to make information about the Aggregation
available. This new Resource, named a Resource Map, has
a URI and a machine-readable Representation that provides
details about the Aggregation. In essence, a Resource Map
expresses which Aggregation it describes (the ore:describes
relationship in Figure 2), and it lists the Aggregated
Resources that are part of the Aggregation (the ore:aggregates
relationship in Figure 2, a subproperty of
dcterms:hasPart). But, a Resource Map can also express
relationships and properties pertaining to all these Resources,
as well as metadata pertaining to the Resource Map itself,
e.g. who published it and when it was most recently
modied (the dcterms:creator and dcterms:modified
relationships in Figure 2). A Resource Map can also express
relationships of the Aggregation, Aggregated Resources, and
the Resource Map itself with any arbitrary other Resource,
as long as the resulting RDF graph is connected.</p>
      <p>In addition, for discovery purposes, the data model allows
a Resource Map to express that an Aggregated Resource of
a speci c Aggregation is also part of another Aggregation.
This is achieved by means of the ore:isAggregatedBy
relationship (the inverse of ore:aggregates) between the
Aggregated Resource and that other Aggregation. Also
stating that an Aggregated Resource is itself an Aggregation
(nesting Aggregations) is supported. To that purpose, an
ore:isDescribedBy relationship (the inverse of
ore:describes, and a subproperty of rdfs:seeAlso) is
expressed between the Aggregated Resource and a Resource
Map that describes it as being itself an Aggregation.
Furthermore, the use of non-protocol-based identi ers (such
as DOIs) that can be expressed as URIs is quite common
for referencing scholarly assets. In order to support this
practice, the ore:similarTo relationship between an
Aggregation and a somehow equivalent resource identi ed by
a non-protocol-based URI is expressed. The speci city of
ore:similarTo is situated between rdfs:seeAlso and
owl:sameAs.
3.2</p>
    </sec>
    <sec id="sec-8">
      <title>Proxies: Aggregated Resources in Context</title>
      <p>We note that the URI asserted in a Resource Map to
denote an Aggregated Resource of a particular Aggregation is
no di erent than the URI that denotes that Resource
independent of the Aggregation. However, it is important in
scholarly communication, among others for the purpose of
citing and expressing provenance, that a resource such as a
dataset included in some context, for example a speci c
article, be distinct from the same dataset outside the context
of that article, or in the context of another article.</p>
      <p>
        To accomplish this di erentiation, OAI-ORE introduces
the notion of a Proxy. A Proxy is a Resource that stands for
an Aggregated Resource in the context of a speci c
Aggregation. The URI of a Proxy provides a mechanism for
denoting a Resource in context. Figure 3 shows the ore:ProxyFor
and ore:ProxyIn relationships between a Proxy and an
Aggregated Resource and an Aggregation, respectively. It also
illustrates how citing the Aggregated Resource is di erent
from citing its Proxy: the former cites a Resource \as is",
the latter cites that Resource as it exists in the context of
a speci c Aggregation. In order to work seamlessly in the
Web and to provide context information to OAI-ORE aware
clients, resolution of HTTP URIs assigned to Proxies must
lead to the Aggregated Resource, and the response must
include a HTTP Link Header [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] that points to the
Aggregation.
3.3
      </p>
    </sec>
    <sec id="sec-9">
      <title>Resource Map Serializations</title>
      <p>A Resource Map has a Representation that describes an
Aggregation in some serialization syntax. OAI-ORE
explicitly speci es three serialization syntaxes, Atom XML,
RDF/XML, and RDFa, while other serialization syntaxes
are possible. Which one to choose will largely depend on
the use case and on the technical environment available to a
Resource Map publisher. For example, in cases where an
expressive HTML splash page exists an RDFa approach might
be attractive. Note that multiple Resource Maps, each
using a di erent serialization syntax can describe the same
Aggregation, and that these may di er in expressiveness3.</p>
      <p>
        Although the data model is based on RDF, we were
committed to also specify a serialization based on Atom, to
allow Aggregations to become the subject of Web 2.0 reuse
scenarios and of work ows based on the Atom Publishing
Protocol [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The Atom Publishing Protocol adds a
uniform read/write approach to Web 2.0, which could be of
signi cant bene t in scholarly communication scenarios.
      </p>
      <p>However, the task of reconciling the data model with the
Atom model proved to be non-trivial due to tensions
between the RDF model and the XML-oriented Atom
speci cation. The former is graph-based, with precise
semantics that are global rather than local to a speci c document.
The latter is hierarchical, (XML) document-centric, and has
intentionally loose element de nitions. It took several,
dramatically di erent iterations of the Atom serialization to
arrive at an acceptable solution.</p>
      <p>The resulting approach expresses an Aggregation by means
of an Atom entry, and makes use of Atom's extensibility
mechanisms in much the same way as Google Data does. For
example, Atom's link element with an OAI-ORE-speci c
value for the rel attribute is used to aggregate resources.
And, awaiting a solution from the Atom community to deal
express triples, an ore:triples element was introduced to
act as a wrapper for RDF descriptions. To support
unambiguous interpretation of Atom serializations of Resource
Maps, a GRDDL transform was implemented that extracts
all contained triples that pertain to the OAI-ORE data model,
both from the native Atom elements and from the ore:triples
extension element, and expresses them in RDF/XML4.
3.4</p>
    </sec>
    <sec id="sec-10">
      <title>Leveraging HTTP</title>
      <p>
        In order to make OAI-ORE work in the HTTP-based
Web, both the Aggregation and the Resource Map are
assigned HTTP URIs, and the Cool URIs for the Semantic
Web guidelines [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] are adopted to support discovery of the
HTTP URI of a Resource Map given the HTTP URI of an
Aggregation. Figure 4 illustrates a situation in which the
arXiv Aggregation is described by both an Atom XML and
an RDF/XML Resource Map, and in which a client is led
to the Atom version via an HTTP 303 redirect and Content
Negotiation.
3.5
      </p>
    </sec>
    <sec id="sec-11">
      <title>Authoritative Resource Maps</title>
      <p>After one party has published a Resource Map that
contains a description and a URI for a new Aggregation, any
other party can publish competing or even con icting
Resource Maps that describe the same Aggregation. To
ad3See http://www.openarchives.org/ore/atom for detailed
Atom and RDF/XML versions of Resources Maps
corresponding to Figure 1.
4http://www.openarchives.org/ore/atom-grddl
dress this we distinguish between Authoritative and
NonAuthoritative Resource Maps in the same way as the Linked
Data guidelines. An Authoritative Resource Map is one
that is accessible by dereferencing the URI of the
Aggregation that it describes, for example using the aforementioned
Cool URI mechanisms. A Non-Authoritative Resource Map
is one not reachable in this manner. The rationale for this
approach is that the party that introduces a new
Aggregation simultaneously mints URIs for both the Aggregation
and the Resource Map, and actually controls both.
4.</p>
    </sec>
    <sec id="sec-12">
      <title>EARLY DEMONSTRATORS</title>
      <p>Since the OAI-ORE speci cations have only been released
recently, an in-depth evaluation of functionality, adoption,
and impact is premature. Still, in this section we give an
insight in e orts by early adopters to leverage the speci
cations. Four use cases are described below. Additional
illustrations of its application are provided by the submissions
to the ORE Challenge at RepoCamp 20085.
4.1</p>
    </sec>
    <sec id="sec-13">
      <title>Foresite: Revealing Aggregations</title>
      <p>In order to provide feedback on the evolving OAI-ORE
speci cation, the UK's Joint Information Systems
Committee (JISC)6 funded an experiment to investigate applying it
to an extensive scholarly collection: the approximately four
million articles that are part of the JSTOR7 collection. By
developing open source OAI-ORE libraries8 and applying
them to produce interlinked Resource Maps, the Foresite
project e ectively demonstrated the feasibility of exposing
common scholarly artifacts to the Data Web in the manner
proposed by OAI-ORE. The project provided valuable
feedback that helped re ne the OAI-ORE speci cations, and
had a signi cant impact on the aforementioned discussions
regarding the Atom serialization of Resource Maps.</p>
      <p>The overall structure of the Aggregations, and associated
Resource Maps, produced for the JSTOR collection mirrors
the journal - issue - article hierarchy of the JSTOR content.
Each journal is modeled as an Aggregation of journal issues;
5http://www.openarchives.org/ore/RepoCamp2008/
6http://www.jisc.ac.uk/
7http://www.jstor.org/
8http://foresite-toolkit.googlecode.com/
each issue is an Aggregation of articles; and each article is an
Aggregation of individual page images and a PDF-formatted
version of the entire article (Figure 5). The Aggregated
Resources at each level are also the subject and/or object
of a fst:followedBy relationship introduced to preserve
the page-turning order for pages within an article, articles
within an issue and so forth. Because fst:followedBy is not
a global relationship, but rather only applies within the
context of a speci c Aggregation, Proxies for these Aggregated
Resources were introduced. The article Aggregations
interlink via dcterms:references relationships for citations,
further con rming the necessity of the graph-based nature
of the OAI-ORE date model, even though the main JSTOR
content hierarchy is tree-shaped. The Resource Maps were
published on a Web server at the University of Liverpool.</p>
      <p>The resulting OAI-ORE descriptions are of immediate
business importance to JSTOR. While JSTOR stores the
OCR-ed full-text of each article, it is only able to openly
expose this kind of topological metadata, and would lose
its market advantage (and the participation of contributing
publishers) if the full-text were exposed. Having the
topology of their collection available in a standardized format that
provides links back to their protected full-text documents
and images, facilitates reuse in third party applications that
can help drive tra c to the JSTOR site and increase its
customer base.</p>
      <p>In order to provide a value-added service on the basis of
the generated Resource Maps without requiring JSTOR to
integrate prototype code into their production portal, the
Foresite Explorer { a visualization application9, was
developed using GreaseMonkey10 and its cross-site capable
XmlHttpRequest. This one-click-install plug-in for Firefox11
extracts the URI of the resource that is currently being viewed
in the JSTOR Web interface and retrieves the associated
RDF/XML Resource Map that describes the Aggregation
9http://foresite.cheshire3.org/explorer/
10http://www.greasespot.net/
11http://www.mozilla.com/firefox/
to which the Web resource corresponds from the Liverpool
Web server. The plug-in then parses and displays the
Resource Map graph via dynamic SVG. Nodes in the display
represent Aggregations, Aggregated Resources, and related
Resources. Nodes for Aggregations can be clicked to expand
or contract the visualization; in case of expansion, new
Resource Maps are obtained, parsed, and again visualized.</p>
      <p>Further experiments using the same approach were
carried out on mainstream Web portals, leveraging the
provided Web service APIs to obtain metadata, and to express
it according to the ORE data model. Flickr12 and Amazon13
were selected, and wrapper services were built to generate
Resource Maps on demand through REST interactions, and
to publish them on the Liverpool server. Flickr provides a
rich dataset with photos, photo sets, users, groups, favorites
and even comments and tags that can all be modeled as
Aggregations. Figure 6 shows a visualization of the
structure of the Flickr Set \Glaciers" that consists of ve
photographs. In the Foresite Explorer, this set is represented
with an Aggregation visualized as the top right node within
the OAI-ORE logo (left bottom of Figure 6), emitting a red
dcterms:creator arc and a white ore:aggregates arc. The
latter leads to the ve photographs. The third photograph
is selected, and another white ore:aggregates arc reaches
out to the available image les (di ering image resolutions)
represented as black nodes. The purple nodes indicate other
aggregations in which the selected photo is aggregated.</p>
      <p>Amazon o ers fewer constructs that readily map to the
OAI-ORE data model, but the user wishlists is a compelling
one. The mapping to the data model is as follows: a
wishlist becomes an Aggregation, and wished-for items become
Aggregated Resources. Interestingly, each item in an
Amazon wishlist has a unique identi er by which it is purchased.
That identi er is only valid within that speci c wishlist to
allow tracking of individual items, once purchased. These
wishlist speci c constructs map directly the Proxies of the
OAI-ORE model. The GreaseMonkey script was updated to
discover these identi ers that are necessary to interact with
the Amazon Web services, and Proxy-based relationships
12http://www.flickr.com/
13http://www.amazon.com/
were added to the visualization.</p>
      <p>Overall, the Foresite experiment has illustrated the
applicability of the OAI-ORE resource aggregation model as
well as the feasibility to leverage it to create a value-added
service. It has demonstrated this for both common
scholarly communication artifacts and speci c constructs used
by popular Web portals. The Foresite experiment will be
described in more detail in a dedicated, future publication.
4.2</p>
    </sec>
    <sec id="sec-14">
      <title>Astronomy Publication Workflow</title>
      <p>Datasets are of fundamental importance in observational
sciences such as astronomy. The astronomy community has
developed sophisticated repositories and data standards,
exempli ed by the Sloan Digital Sky Survey14 and the
National Virtual Observatory15, which provide excellent
facilities for registering and accessing large datasets. However,
when submitting an article, both new datasets that were
created to arrive at ndings reported in an article, and data
citation information that reveals the reuse of existing datasets
are often lost, \left behind" on the personal computer of the
author.</p>
      <p>
        A team at Johns Hopkins University is collaborating with
the American Astronomical Society to capture datasets as
part of the publication work ow [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In the newly devised
publication work ows, OAI-ORE Aggregations are used to
glue an article and its associated datasets together, and
Resource Maps that describe these Aggregations are the tokens
that move around between author, publisher and dataset
repository as the publication process proceeds [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. At each
stage of the publication work ow, the Resource Map is used
to convey the current state of the Aggregation, and is then
updated to re ect the new state that is then passed on to
the next work ow phase. For example, as a Resource Map
is passed from the publisher to the dataset repository and
back again, it is updated to contain the URIs of datasets
that are registered in the repository, and that were used for
the article. This allows the publisher to link to the datasets
that were used for a speci c article, and the repository to
link to papers that used a speci c dataset.
      </p>
      <p>Generally, the availability of these Aggregations enables
new services to be built on both the publishing platform and
the data repository. If the practices proposed by this novel
publication work ow became commonplace, it would
represent a signi cant improvement in the e ciency of scienti c
communication.
4.3</p>
    </sec>
    <sec id="sec-15">
      <title>Authoring, Editing and Reusing</title>
      <p>The success of OAI-ORE depends on the ease with which
Aggregations and Resource Maps are authored and
disseminated on the Web. In many cases, they will be generated
automatically based on information that is available in an
information system. For example, the arXiv.org database
contains all information that is necessary to automatically
generate Aggregations and their associated Resource Maps,
as shown in the Appendices. And, in the astronomy project
described above, the ability to create Resource Maps is built
into familiar authoring environments in a manner that makes
it a side-e ect of the authoring process and thus minimizes
the burden on authors.</p>
      <p>
        Like all cyberinfrastructure, the success of such authoring
environments depends on the manner in which assembling
14http://www.sdss.org/
15http://www.us-vo.org/
all resources that relate to a particular research task or
publication ts into the normal scholarly work ow. Two
authoring environments that demonstrate this are the Literature
Object Reuse and Exchange (LORE) tool created by Gerber
et al.16, and by the SCOPE work of Cheung et al. [
        <xref ref-type="bibr" rid="ref21 ref8">8, 21</xref>
        ].
LORE is a Firefox extension that communicates via Ajax
with a Sesame2 data store for maintaining the OAI-ORE
graphs that are generated. LORE allows for the generation
of ne-grained metadata and relationships, for example,
allowing indicating that a certain resource is contextual
information about the literature work that is being studied.
The SCOPE work led to the development of the Provenance
Explorer, a stand-alone Java application with functionalities
similar to those of LORE, but aimed at the creation, editing
and publication of scienti c compound objects.
4.4
      </p>
    </sec>
    <sec id="sec-16">
      <title>Enhanced Publications</title>
      <p>The Dutch SURFshare program17 and the European
DRIVER II project18 are collaborating on
cyberinfrastructure to join a multitude of scienti c repositories that hold
publications and research data. The goal is to give
researchers better means to share and access scienti c
materials through innovative services. One of the envisioned
services relates to enhanced publications, composites of textual
publications and supporting resources such as research-data,
visualizations, annotations, related websites, etc. To ensure
the integrity and usability of such enhanced publications it
is important that all its components and their interrelations
are being preserved.</p>
      <p>
        A study into object models suitable for the
representation of enhanced publications recommended the use of
OAIORE. As a result, a demonstrator project [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] was launched
in which enhanced publications for multiple scienti c
disciplines ranging from engineering to journalism were modeled
according to OAI-ORE, and in which approaches to meet
a variety of requirements were explored, including
presentation, navigation, persistent identi cation, granularity of
referencing, handling of sequentially ordered resources,
visualization of interrelationships, etc. The results are available
at the project site19. The project chose RDF/XML to
express Resource Maps and uses an XSLT-based approach to
dynamically generate an HTML \splash page" from them.
In each splash page, a Content tab (Figure 7) lists all
crucial metadata about the enhanced publication, prominently
shows its textual component and associated metadata, and
neatly lists additional resources again with metadata. Many
of these resources are themselves modeled as Aggregations,
and hence also have their own splash page. To support an
understanding of the relationships among resources of an
Aggregation and of nested Aggregations, a Relations tab
that loads a Java applet fueled by Resource Map content
is introduced. Overall, the demonstrator is remarkable
because of the elegance and simplicity of the ORE
implementation. It clearly illustrates that ORE can be used as a basic
model for enhanced publications, and points at the need for
community-de ned vocabularies to convey expressive
relationships among scienti c resources.
16http://www.openarchives.org/ore/RepoCamp2008/
#LORE
17http://www.surffoundation.nl/en/
18http://www.driver-community.eu/
19http://driver2.dans.knaw.nl/demonstrator/html/
      </p>
    </sec>
    <sec id="sec-17">
      <title>RELATED WORK</title>
      <p>Given the widespread use of aggregations in both the
physical and the Web world, it comes as no surprise that
other e orts have investigated this domain. Prior work in
the Web realm can be grouped in two main categories
depending on the party that introduces aggregations. In one
case, that is the Web navigator (agent or reader), in the
other case it is the administrator of a Web-based information
system. We look at a number of e orts in both categories,
and evaluate their capabilities to identify aggregations, to
enumerate the constituent resources of an aggregation, to
express relationships among resources, and to accommodate
resources that are distributed on the Web.</p>
      <p>
        In the Web navigator case, either an interactive user groups
resources based on some intent, or a robot tries to infer the
implicitly de ned members of an aggregation. The robotic
approaches range from heuristics [
        <xref ref-type="bibr" rid="ref14 ref30">30, 14</xref>
        ] to machine-learning
[
        <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
        ]. While these approaches are useful, they are
imperfect and dependent on the perception of those encoding the
heuristics or training set and they do not necessarily re ect
the intention of the original authors of the Web resources.
And, while these approaches may succeed at selecting the
distributed resources that are part of an implicitly de ned
aggregation, they are not capable of inferring the
relationships between those resources, nor do they propose a way to
unambiguously describe the aggregation.
      </p>
      <p>The approaches that involve an interactive user include
tools such as GroupMe!20 and LinkBunch21. LinkBunch
lets users submit several URIs that are then assigned a new
HTTP URI that, when dereferenced, returns an HTML page
that lists and links to the originally submitted URIs. The
20http://groupme.org/
21http://linkbun.ch/
\bunch" has a new HTTP URI identity, it enumerates its
members, and it readily handles distributed Web resources.
However, the identity of the bunch is the same as that of the
HTML page that describes it, and expressing relationships
between the bunched resources is not supported. GroupMe!
is similar, with the addition of social tagging capabilities,
but has the same problems as LinkBunch.</p>
      <p>
        Some Web navigator approaches work in an opposite
granular direction, supporting disaggregation of a single Web
resource (i.e., an HTML page) into multiple resources. This
can be done automatically, such as for segmented display
on limited devices such as PDAs [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] or for recovering
structured records from Web pages [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Decomposition can also
be done manually, such as for reuse and sharing of parts of
a Web page (e.g., ClipMarks22). All these approaches,
manually or automatically, can be thought of as adding (or
inferring) HTML anchors where none exist. These approaches
assign identities to the newly created resources (fragments
of the original resource), but they provide no approach to
describe the original resource as an aggregation of these new
resources, nor do they allow expressing relationships among
them.
      </p>
      <p>
        In approaches that have the administrator of a Web
information system in the diver seat, several technologies exist to
deal with resource aggregations. Sitemaps were brie y
considered as a serialization option for Resource Maps. Google,
Yahoo and Microsoft support the Sitemap Protocol [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], a
simple XML le format that allows Web sites to list the URIs
they want crawled by robots. Sitemaps provide for minimal
metadata (e.g., last modi cation date, update frequency and
crawl priority), but no attempt is made to provide semantic
typing, and handling arbitrary distributed resources is not
supported. Indeed, in the interest of trust, the Sitemap
Protocol speci es a signi cant limitation on URI paths that can
be listed in a Sitemap le. For example, a Sitemap at level
www.foo.com/a/b can list URIs at level a/b and below, but
it cannot list URIs at www.foo.com/a/c, www.foo.com/d/ or
www.bar.com/.
      </p>
      <p>
        We made a deliberate decision to avoid the many
existing packaging formats, such as MPEG-21 DIDL [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], METS
[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], FOXML [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], IMS-CP [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], and BagIt [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. First,
packaging base64-encoded content in a wrapper document does
not resonate well with the Resource/URI/Representation
paradigm of the Web Architecture. Still, most of these
formats also support a by-reference mechanism to deliver
content, in which URIs can be used. However, although these
formats are prominent in their respective communities, they
have not gained an adoption comparable to that of Atom or
RDF/XML. And while these approaches can address
identi cation, and enumeration of distributed resources, they
have uneven capabilities to express the graph-based
OAIORE model, due to their hierarchical perspective.
      </p>
      <p>
        In the course of the OAI-ORE e ort, we also attempted to
model aggregations as Atom feeds, not entries [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. We
ultimately decided that was the wrong granularity, especially
since common Web 2.0 reuse scenarios, including use with
the Atom Publishing Protocol, work at the level of Atom
entries. The Atom Syndication Format was preferred over
the various RSS formats in anticipation of using the Atom
Publishing Protocol [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        Some elements of the POWDER [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] speci cations that
22http://clipmarks.com/
were developed in the same timeframe as OAI-ORE
address a problem space similar to that of OAI-ORE. However,
POWDER's focus is signi cantly broader, and it approaches
the problem from the opposite perspective,
      </p>
      <p>focusing on capabilities to assert (via \Description
Resources") that a group of resources share certain properties
(e.g. access rights), rather than asserting arbitrary
properties about resources that, for some reason, are grouped
into an aggregation. That is, in POWDER the notion of
shared properties de nes an aggregation, whereas in
OAIORE an aggregation can be created for any reason deemed
important by its creator. Also, while POWDER provides
capabilities to describe a group of resources using a
variety of approaches including regular expressions, it does not
introduce an identity for the aggregation.</p>
    </sec>
    <sec id="sec-18">
      <title>CONCLUSIONS</title>
      <p>This paper has introduced the OAI-ORE solution to the
resource aggregation problem, which we argue meets a
critical need in the development of cyberinfrastructure and the
next generation scholarly communication infrastructure. By
aligning the solution with the Web Architecture, and by
leveraging the practices of the Semantic Web and Linked
Data e ort, it will facilitate better integration of scholarly
communication with the mainstream Web, it will make
scholarly artifacts more readily usable with common Web tools
and applications, and it will bene t the broader community
by making research materials more visible, veri able, and
by facilitating unexpected reuse.</p>
      <p>While OAI-ORE was motivated by scholarly
communication, we believe that the proposed solution has broader
applicability. Aggregations, sets, and collections are as
common on the Web as they are in the everyday physical world.
In many situations it would bene t agents and services if
aggregations were unambiguously enumerated and described,
essentially layering an addition level of resource granularity
upon the Web.</p>
      <p>Evaluation of the OAI-ORE work depends on its
adoption and evolution over time. The work has so far
bene ted from signi cant community involvement throughout
the speci cation process, and the international team that
developed the solution includes representatives with
backgrounds in scholarly publishing, eScience, repository
infrastructure, digital libraries, Web search engines, linked data,
and information interoperability. Work by early adopters,
such as the Foresite project and John's Hopkins
publication work ow project, are promising indicators that these
community contributions have led to a solution that stands
realistic chances for signi cant adoption.</p>
    </sec>
    <sec id="sec-19">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was supported by the National Science
Foundation Divisions of Information and Intelligent Systems and
Undergraduate Education through grant numbers IIS-0430906,
IIS-0643784 and DUE-0840744, the Andrew W. Mellon
Foundation, Microsoft, and the Coalition for Networked
Information. Development of OAI-ORE was based on input from
the OAI-ORE Technical Committee, the OAI-ORE Liaison
Group, the OAI-ORE Advisory Committee, contributors to
the OAI-ORE Google discussion group, and members of
the Digital Library Research &amp; Prototyping Team of the
Los Alamos National Laboratory. Individuals are listed at
http://www.openarchives.org/ore/.
8.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Altman</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>King</surname>
          </string-name>
          .
          <article-title>A proposed standard for the scholarly citation of quantitative data</article-title>
          .
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          /4),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Atkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. K.</given-names>
            <surname>Droegemeier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. I.</given-names>
            <surname>Feldman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Garcia-Molina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Messerschmitt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Messina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Ostriker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Wright</surname>
          </string-name>
          .
          <article-title>Revolutionizing science</article-title>
          and engineering through cyberinfrastructure,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bekaert</surname>
          </string-name>
          , E. De Kooning, and H. Van de Sompel.
          <article-title>Representing digital objects using MPEG-21 Digital Item Declaration</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <volume>159</volume>
          {
          <fpage>173</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          .
          <article-title>How to publish linked data on the web</article-title>
          ,
          <year>2007</year>
          . http://sites.wiwiss.fuberlin.de/bizer/pub/LinkedDataTutorial/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Borgman</surname>
          </string-name>
          .
          <article-title>Scholarship in the digital age : information, infrastructure, and the Internet</article-title>
          . MIT Press, Cambridge, Mass.,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Boyko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kunze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Littman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Madden</surname>
          </string-name>
          .
          <source>The bagit le package format (v0.95)</source>
          , Internet Draft,
          <year>July 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Chakrabarti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Punera</surname>
          </string-name>
          .
          <article-title>A graph-theoretic approach to webpage segmentation</article-title>
          .
          <source>In WWW '08: Proceedings of the 17th international conference on World Wide Web</source>
          , pages
          <volume>377</volume>
          {
          <fpage>386</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hunter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lashtabeg</surname>
          </string-name>
          , and
          <string-name>
            <surname>D. J. SCOPE -</surname>
          </string-name>
          <article-title>a scienti c compound object publishing and editing system</article-title>
          .
          <source>In 3rd International Digital Curation Conference</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choudhury</surname>
          </string-name>
          , T. DiLauro,
          <string-name>
            <given-names>A.</given-names>
            <surname>Szalay</surname>
          </string-name>
          , E. Vishniac,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hanisch</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Ste en</surname>
          </string-name>
          , R. Milkey,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ehling</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Plante</surname>
          </string-name>
          .
          <article-title>Digital data preservation for scholarly publications in astronomy</article-title>
          .
          <source>International Journal of Digital Curation</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>DiLauro</surname>
          </string-name>
          .
          <article-title>OAI-ORE for publishing work ows: Data archiving for journals of the American Astronomical Society</article-title>
          .
          <source>In Open Repositories</source>
          <year>2008</year>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Dmitriev</surname>
          </string-name>
          .
          <article-title>As we may perceive: nding the boundaries of compound documents on the web</article-title>
          .
          <source>In WWW '08: Proceedings of the 17th international conference on World Wide Web</source>
          , pages
          <volume>1029</volume>
          {
          <fpage>1030</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Dmitriev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Suchkov</surname>
          </string-name>
          .
          <article-title>As we may perceive: inferring logical documents from hypertext</article-title>
          .
          <source>In Proceedings of the sixteenth ACM conference on Hypertext and Hypermedia</source>
          , pages
          <volume>66</volume>
          {
          <fpage>74</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J</given-names>
            .
            <surname>Jackson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Bowker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Knobel</surname>
          </string-name>
          .
          <article-title>Understanding infrastructure: Dynamics, tensions, and design</article-title>
          .
          <source>Technical report, National Science Foundation</source>
          ,
          <year>January 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Eiron</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>McCurley</surname>
          </string-name>
          .
          <article-title>Untangling compound documents on the web</article-title>
          .
          <source>In Proceedings of the fourteenth ACM conference on Hypertext and Hypermedia</source>
          , pages
          <volume>85</volume>
          {
          <fpage>94</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Embley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ng</surname>
          </string-name>
          .
          <article-title>Record-boundary discovery in Web documents</article-title>
          .
          <source>In Proceedings of the 1999 ACM SIGMOD international conference on Management of data</source>
          , pages
          <volume>467</volume>
          {
          <fpage>478</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Google</surname>
          </string-name>
          , Microsoft, and
          <string-name>
            <surname>Yahoo</surname>
          </string-name>
          .
          <source>Sitemaps XML format</source>
          ,
          <year>2008</year>
          . http://www.sitemaps.org/protocol.php.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Szalay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Thakar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stoughton</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. vandenBerg.</surname>
          </string-name>
          <article-title>Online scienti c data curation, publication, and archiving</article-title>
          .
          <source>Technical Report arXiv cs.DL/0208012</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gregorio</surname>
          </string-name>
          and B. de hOra.
          <source>The Atom publishing protocol</source>
          ,
          <source>Internet RFC-5023</source>
          ,
          <year>December 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Haslhofer</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schandl</surname>
          </string-name>
          .
          <article-title>The OAI2LOD Server: Exposing OAI-PMH Metadata as Linked Data</article-title>
          .
          <source>In Proceedings of WWW 2008 Workshop Linked Data on the Web (LDOW2008)</source>
          , Beijing,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hoogerwerf</surname>
          </string-name>
          .
          <article-title>Durable enhanced publications</article-title>
          .
          <source>In Proceedings of African Digital Scholarship &amp; Curation</source>
          <year>2009</year>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>L. Hunter J.</given-names>
            ,
            <surname>Chueng</surname>
          </string-name>
          . Provenance explorer
          <article-title>- a graphical interface for constructing scienti c publication pack ages from provenance trails</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          -2):
          <volume>99</volume>
          {
          <fpage>107</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <article-title>IMS Global Learning Consortium</article-title>
          .
          <article-title>IMS content packaging XML binding speci cation version 1.1.3</article-title>
          . http://www.imsglobal.org/content/packaging/,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>I.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <article-title>Architecture of the world wide web, volume one</article-title>
          .
          <source>Technical Report W3C Recommendation 15 December</source>
          <year>2004</year>
          ,
          <fpage>W3C</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Johnson</surname>
          </string-name>
          .
          <article-title>Institutional repositories: Partnering with faculty to enhance scholarly communication</article-title>
          .
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <volume>8</volume>
          (
          <issue>11</issue>
          ),
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Payette</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Shin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wilper</surname>
          </string-name>
          .
          <article-title>Fedora: an architecture for complex objects and their relationships</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <volume>124</volume>
          {
          <fpage>138</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze and H. Van de Sompel</surname>
          </string-name>
          .
          <article-title>The Open Archives Initiative: building a low-barrier interoperability framework</article-title>
          .
          <source>In JCDL '01: Proceedings of the 1st ACM/IEEE-CS Joint Conference on Digital Libraries</source>
          , pages
          <volume>54</volume>
          {
          <fpage>62</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze</surname>
          </string-name>
          , H. Van de Sompel, P. Johnston,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Warner. ORE Speci cation - Abstract Data Model</surname>
          </string-name>
          ,
          <year>2008</year>
          . http://www.openarchives.org/ore/datamodel.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze</surname>
          </string-name>
          , H. Van de Sompel, P. Johnston,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. Warner. ORE</surname>
          </string-name>
          <article-title>Speci cation</article-title>
          and User Guide - Table of Contents,
          <year>2008</year>
          . http://www.openarchives.
          <source>org/ore/1</source>
          .0/toc.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lagoze</surname>
          </string-name>
          , H. Van de Sompel, P. Johnston,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Warner. Object</surname>
          </string-name>
          Re-Use &amp;
          <article-title>Exchange: A Resource-Centric Approach</article-title>
          .
          <source>Technical Report arXiv:0804.2273</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kolak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Vu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Takano</surname>
          </string-name>
          .
          <article-title>De ning logical domains in a web site</article-title>
          .
          <source>In Proceedings of the eleventh ACM on Hypertext and Hypermedia</source>
          , pages
          <volume>123</volume>
          {
          <fpage>132</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Lynch</surname>
          </string-name>
          .
          <article-title>Institutional repositories: Essential infrastructure for scholarship in the digital age</article-title>
          .
          <source>ARL: A Bimonthly Report</source>
          , (
          <volume>226</volume>
          ),
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McDonough. METS:</surname>
          </string-name>
          <article-title>Standardized encoding for digital library objects</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <volume>148</volume>
          {
          <fpage>158</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>P.</given-names>
            <surname>Murray-Rust</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Rzepa</surname>
          </string-name>
          .
          <article-title>The next big thing: From hypermedia to datuments</article-title>
          .
          <source>Journal of Digital Information</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nottingham</surname>
          </string-name>
          .
          <article-title>HTTP header linking</article-title>
          , Internet Draft,
          <year>March 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nottingham</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Sayre</surname>
          </string-name>
          .
          <article-title>The Atom syndication format</article-title>
          ,
          <source>Internet RFC-4287</source>
          ,
          <year>December 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sauermann</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          .
          <article-title>Cool URIs for the semantic web</article-title>
          .
          <source>Technical Report W3C Interest Group Note 31 March</source>
          <year>2008</year>
          , W3C,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>K.</given-names>
            <surname>Scheppe</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Pentecost</surname>
          </string-name>
          .
          <article-title>Protocol for Web Description Resources (POWDER): Primer</article-title>
          .
          <source>Technical Report W3C Working Draft { 14 November</source>
          <year>2008</year>
          ,
          <fpage>W3C</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Sieber</surname>
          </string-name>
          and
          <string-name>
            <surname>B. E. Trumbo.</surname>
          </string-name>
          <article-title>(not) giving credit where credit is due: Citation of data sets</article-title>
          .
          <source>Science and Engineering Ethics</source>
          ,
          <volume>1</volume>
          :
          <fpage>11</fpage>
          {
          <fpage>20</fpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>A.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>The research library in the 21st century: collecting, preserving, and making it accessible resources for scholarship</article-title>
          .
          <source>In No Brief Candle: Reconceiving Research Libraries for the 21st Century. Council on Library and Information Resources</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>S.</given-names>
            <surname>Tuecke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Czajkowski</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Foster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Frey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Graham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kesselman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Maquire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sandholm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Snelling</surname>
          </string-name>
          , and
          <string-name>
            <surname>Vanderbilt</surname>
          </string-name>
          .
          <source>Open Grid Services Infrastructure (OGSI): Version 1.0. Technical Report draft-ggf-ogsi-gridservice-33, Global Grid Forum, January</source>
          <volume>27</volume>
          2003.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>H. Van de Sompel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Payette</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Erickson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lagoze</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Warner</surname>
          </string-name>
          .
          <article-title>Rethinking scholarly communication: Building the system that scholars deserve</article-title>
          .
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <volume>10</volume>
          (
          <issue>9</issue>
          ),
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>R.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Moore</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Hanisch</surname>
          </string-name>
          .
          <article-title>A virtual observatory vision based on publishing and virtual data</article-title>
          .
          <source>Technical report, US National Virtual Observatory</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>