<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linked Data for Libraries: A Project Update</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dean B. Krafft</string-name>
          <email>dean.krafft@cornell.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cornell University Library</institution>
          ,
          <addr-line>Ithaca, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>This poster reports on the first eighteen months of the Mellon-funded two-year Linked Data for Libraries (LD4L) project [1], a partnership of Cornell University Library, Stanford University Libraries, and the Harvard Library Innovation Lab. The goal of the project is to use Linked Open Data to leverage the intellectual value that librarians and other domain experts and scholars add to information resources when they describe, annotate, organize, select, and use those resources, together with the social value evident from patterns of usage. The project is producing an ontology, architecture, and set of tools that work both within and across individual institutions in an extensible network.</p>
      </abstract>
      <kwd-group>
        <kwd>Ontology</kwd>
        <kwd>Linked Data</kwd>
        <kwd>Libraries</kwd>
        <kwd>VIVO</kwd>
        <kwd>Hydra Framework</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The Cornell University Library, the Harvard Library Innovation Lab, and the Stanford
University Libraries have all been exploring new approaches to dramatically improve
the discovery experience for users seeking scholarly information resources, such as
traditional monograph and journal publications, archival materials, research datasets,
images, recordings, cultural artifacts, newspapers and magazines, web archives, and much
more. All three institutions have been looking at ways to gather context and
relationships about these resources that go far beyond traditional metadata approaches. The
goal of this project is to create a Linked Data for Libraries (LD4L) model that works
both within individual institutions and through a coordinated, extensible network of
Linked Open Data (LOD). This LOD will capture the intellectual value that librarians
and other domain experts add to information resources when they describe, annotate,
organize, select, and use those resources, together with the social value evident from
patterns of usage.</p>
      <p>
        To achieve this goal, the project team will:
• Create a set of use cases that specify how LD4L information can enhance user
discovery and understanding of scholarly information resources
• Assemble and, where necessary, create an LD4L ontology to represent the required
bibliographic, person, curation, and usage information as linked data
• Hold a two-day workshop with library, archive, and museum linked data experts
from a variety of institutions to gather feedback on the LD4L use cases, ontology,
and work plan
• Create linked open data sources at each institution providing bibliographic, person,
curation, and usage data for the scholarly information resources of the institution
using the LD4L ontology
• Create and release open-source software for creating institutional LD4L instances
and using LD4L data as part of the Hydra Framework [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
• Create a demonstration search across the combined LD4L linked data from all three
institutions
2
      </p>
      <p>Progress Report
The poster to be presented will summarize progress on three focus areas for the overall
project: use case development, the LD4L ontology, and outcomes from the LD4L
workshop. The poster should be of interest to those who:
1. Are interested in understanding use cases for applying linked data techniques to
describing, discovering, and understanding scholarly information resources;
2. Want to understand the specific ontology choices that the project has made to address
these library use cases; and
3. Want to hear about the feedback from linked data experts on the use cases, ontology,
and demonstration systems that were presented at the LD4L workshop.</p>
      <p>The sections below briefly summarize the material to be presented in each of these
areas.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>LD4L Use Cases</title>
      <p>The work of LD4L has been heavily influenced by use cases; if the LD4L ontology,
any consuming applications, or linked data in general are going to be fit for purpose,
the purpose and criteria for success need to be defined. For the first half of Year 1 of
the LD4L project, partners invested heavily in an extensive process of articulating
what they wanted to accomplish via linked data, for whom, and why it would be
beneficial to realizing the mission of a library. This multi-stage effort used a classic approach
borrowed from agile software development methodologies to articulate functional
requirements or use cases in the form of "stories": "As a &lt;type of user&gt;, I want to
&lt;perform an action&gt;, so that I can &lt;realize a benefit&gt;".</p>
      <p>
        Partners at Harvard, Cornell and Stanford generated a total of 42 raw use cases in
this form. After reviewing for overlap, applicability to linked data, feasibility for
engineering, and availability of data, the project team reduced this suite of use cases to 12
use cases in 6 distinct clusters [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The 6 clusters are links between: 1) bibliographic
and curation data; 2) bibliographic and person data; 3) leveraging external data
including authorities; 4) leveraging the deeper graph (via queries or patterns); 5) leveraging
usage data; and 6) cross-site services.
      </p>
      <p>As an example, Use Case 2.1 (in cluster 2, bibliographic and person data) is: See and
search on works by people to discover more works, and better understand people. An
example story in this use case is: “As a researcher, I'd like to see / search on works &lt;by,
about, cited by, collected, taught&gt; by University faculty &lt;in an OPAC, profiles
system&gt;, to discover works of interest based on connection to people, and to understand
people based on their relation to works.”</p>
      <p>The project team continues to consult the use cases as an ongoing guide. In Year 1,
the work around Use Case Cluster 1, for example, focused on using linked data allow
users to build virtual collections of scholarly information resources drawn from a
variety of source (e.g., items described by library catalog records or items cited in faculty
research profiles). Cornell and Stanford both developed and designed systems to
exercise these capabilities, and demonstrated them in versions of their Blacklight-based
catalogs.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>LD4L Ontology</title>
      <p>One of the major outputs of the project is the LD4L Ontology, which is used to share
information about scholarly resources among the project participants and to
interconnect those resources with the broader web of Linked Open Data.</p>
      <p>
        The group early on confirmed that it makes eminent sense for a project focused on
linked data to draw as much as possible on existing ontologies that have already
achieved significant adoption or show promise for doing so, rather than creating a new
self-contained ontology. Elements of the Bibliographic Ontology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and FaBiO [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] had
already been incorporated into the VIVO-ISF Ontology [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and were familiar to team
members from previous work. The BIBFRAME initiative [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] at the Library of Congress
addresses the representation of MARC metadata in RDF, while OCLC has worked to
extend the Schema.org ontology [
        <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
        ] as a bridge between the library community and
the Web. Use cases 1.1. and 1.2 use the Open Annotation Data Model [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to represent
collection annotations and the Open Archives Initiative Object Reuse and Exchange
(OAI-ORE) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for ordering. For provenance, the Provenance, Authoring and
Versioning (PAV) ontology [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] provides a solid starting point while the W3C Provenance
Ontology (PROV-O) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] offers additional granularity when data are available to
support more nuanced attribution.
      </p>
      <p>
        The team has also prioritized the ability to convert references within library metadata
records from "strings" to "things," reducing reliance on the lexical form of a name by
adopting URI-based identifiers as the primary means of disambiguation. Whenever
possible we seek out persistent global identifiers for the entities being represented –
identifiers from established international efforts such as the Open Researcher and
Contributor ID [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the Virtual International Authority File [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], or the International
Standard Name Identifier [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] for people; global identifier systems are also emerging for
organizations (VIAF, the Ringgold Identify Database [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], and others).
      </p>
      <p>This approach does not preclude the creation or reuse of URIs in a local institutional
namespace as identifiers as part of publishing linked data. When a metadata record only
references an external standard – a subject term in the Getty Art and Architecture
Thesaurus, for example – no local URI is necessary, but when the metadata includes
additional original statements about that entity, a local URI supplemented by an
owl:sameAs assertion to the external entity will be necessary to allow those
locallyasserted statements to be retrievable as linked data.</p>
      <p>Following these principles, the project has now assembled an LD4L Ontology,
which will be used to implement demonstration systems for the use cases during the
final phase of the project. The poster will present a summary of this ontology.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>LD4L Workshop</title>
      <p>
        The poster will also summarize outcomes of the LD4L workshop, which brought
together fifty linked data experts at Stanford in late February 2015, who provided
extensive feedback on the use cases, ontology design, and engineering work to date. Input
from the workshop informed both the ontology and the specific demonstration systems
to be built in the final phase of the project. A full agenda for the workshop, as well as
session notes and slides from the presentations, is available at [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>Acknowledgements. The work described in this poster includes contributions from the
entire LD4L project team, including from Cornell: Dean B. Krafft, Jon Corson-Rikert,
Brian J. Lowe, E. Lynette Rayle, Rebecca Younes, Simeon Warner, Chew Chiat Naun,
Steven Folsom, Jason Kovari, and Jim Blake; from Harvard: David Weinberger, Paul
Deschner, Paolo Ciccarese, Jonathan Kennedy, and Randy Stern; and from Stanford:
Tom Cramer, Philip Schreur, Rob Sanderson, Lynn McRae, Naomi Dushay, Nancy
Lorimer, Darren Weber, and Joshua Greben. The project would like to thank the
Andrew W. Mellon Foundation for its generous support of this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <article-title>Linked Data for Libraries (LD4L</article-title>
          ), http://ld4l.org
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Project</given-names>
            <surname>Hydra</surname>
          </string-name>
          , http://projecthydra.org
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3. https://wiki.duraspace.org/display/ld4l/LD4L+Use+Cases</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Bibliographic</given-names>
            <surname>Ontology</surname>
          </string-name>
          , http://bibliontology.com/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>5. FaBiO, http://vocab.ox.ac.uk/fabio</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>VIVO-ISF Ontology</surname>
          </string-name>
          , https://wiki.duraspace.org/display/VIVO/VIVOISF+Ontology
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. BIBFRAME, http://www.loc.gov/bibframe/</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. http://blog.schema.org/
          <year>2014</year>
          /09/schemaorg-support
          <article-title>-forbibliographic_2</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>9. Schema.org, https://schema.org/</mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>10. Open Annotation Data Model, http://www.openannotation.org/spec/core/</mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>OAI-ORE</surname>
          </string-name>
          , https://www.openarchives.org/ore/
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>12. PAV, http://www.jbiomedsem.com/content/4/1/37</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>13. PROV-O, http://www.w3.org/TR/prov-o/</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>14. ORCID, http://orcid.org/</mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>15. VIAF, http://viaf.org/</mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>16. ISNI, http://www.isni.org</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>17. Ringgold Identify Database, http://www.ringgold.com/identify</mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>18. https://wiki.duraspace.org/display/ld4l/LD4L+Workshop+Overvi ew</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>