<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Conversion of EAD into EDM Linked Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ste en Hennicke</string-name>
          <email>steffen.hennicke@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marlies Olensky</string-name>
          <email>marlies.olensky@ibi.hu-berlin.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victor de Boer</string-name>
          <email>v.de.boer@cs.vu.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antoine Isaac</string-name>
          <email>aisaac@cs.vu.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Wielemaker</string-name>
          <email>j.wielemaker@cs.vu.nl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Europeana</institution>
          ,
          <addr-line>Koninklijke Bibliotheek Prins Willem-Alexanderhof 5, 2509 LK Den Haag</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Humboldt-Universitat zu Berlin, Institut fr Bibliotheksund Informationswissenschaft Dorotheenstrae 26</institution>
          ,
          <addr-line>10117 Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vrije Universiteit Amsterdam, Department of Computer Science De Boelelaan 1081a</institution>
          ,
          <addr-line>1081 HV Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>82</fpage>
      <lpage>88</lpage>
      <abstract>
        <p>We report on ongoing work in Europeana on the conversion of EAD-XML based archival data to an RDF-based representation using the newly developed "Europeana Data Model" (EDM) ontology. This short paper is based on [4]. The project Europeana4 was set up as part of the EU policy framework for the information society and media (i2010 strategy) aiming at the establishment of a single access point to the distributed European cultural heritage. Today, Europeana o ers access to millions of objects from all kinds of cultural heritage communities. The aim of the current agenda of Europeana is to provide semantically contextualized object representations and new functionality based on an API approach and on integration into the Linked Data context [1]. To enable this vision the "Europeana Data Model" (EDM) has been developed. A large portion of the cultural heritage metadata that is prospected to be accessible through Europeana is currently described in archival nding aids used by archives across Europe. The "Encoded Archival Description" (EAD) is an XML standard for encoding such nding aids. As such, a method for converting EAD-compliant metadata to EDM will greatly bene t Europeana's goals. In this paper, we will describe the basic functionality of the EDM and explain some pivotal principles of archival description embodied in EAD-encoded archival nding aids. After this, we elaborate on the EDM-RDF representation of a concrete EAD encoded nding aid. We conclude with the perceived advantages of such a new data representation.</p>
      </abstract>
      <kwd-group>
        <kwd>archive</kwd>
        <kwd>EAD</kwd>
        <kwd>semantic web</kwd>
        <kwd>nding aid</kwd>
        <kwd>RDF</kwd>
        <kwd>EDM</kwd>
        <kwd>Europeana</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>2</p>
    </sec>
    <sec id="sec-2">
      <title>Europeana Data Model (EDM)</title>
      <p>
        The EDM has been speci cally designed to enable Linked Data integration and
to solve the problem of cross-domain data interoperability. The EDM builds on
the reuse of existing standards from the Semantic Web environment but does not
specialize in any community standard [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It acts as a top-level ontology
consisting of elements from standards like OAI-ORE5, Dublin Core Terms6 and SKOS7
and allows for specializations of these elements. Thus, richer metadata can be
expressed through specializations of classes and properties. Some elements were
de ned in the Europeana namespace, yet contain referrals to other metadata
standards. This allows for correct mappings and cross-domain interoperability.
      </p>
      <p>RDF(S)8 is used as an overall meta-model to represent the data. The ORE
approach is used to structure the di erent information snippets belonging to an
object and its representation. It follows the concept of aggregations
(ore:Aggregation): This concept allows to distinguish between digital
representations which are accessible on the Web and thus are modeled as
edm:WebResource9 and the provided object, represented as a edm:ProvidedCHO.</p>
      <p>
        Furthermore, di erent, possibly con icting views from more than one provider
on the same object can be handled in EDM by using the proxy mechanism
(ore:Proxy). The Dublin Core Terms describe the objects. SKOS is used to
model controlled vocabularies which annotate the digital objects [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Archival Description and Finding Aids</title>
      <p>Archival nding aids are guides into the archival material that an archive holds
in the form of archival collections. Typically, printed versions of nding aids
serve archival users in the reading room and archivists in the reference service
as the means to identify relevant archival materials.</p>
      <p>The archival material in an archival collection is organized into records. A
record denotes a group of documents from the archival collection. It does not
describe a single information object like a single book in the library domain.</p>
      <p>According to the principle respect des fonds the description of the internal
structure (original hierarchy and ordering) and the external structure
(provenance) of an archival collection provides information necessary to understand
context and content of the records and to guarantee their authenticity.</p>
      <p>A nding aid contains such information in the form of an archival
description. This archival description typically consists of several parts arranged in a
5 "Open Archives Initiative Protocol - Object Exchange and Reuse": http://www.</p>
      <p>openarchives.org/ore/ [7.10.2010]
6 "Dublin Core": http://dublincore.org/ [7.10.2010]
7 "Simpli ed Knowledge Organization System": http://www.w3.org/2004/02/skos/
[7.10.2010]
8 "Resource Description Framework (Schema)": http://www.w3.org/TR/rdf-primer/
[5.09.2011] and http://www.w3.org/TR/rdf-schema/ [5.09.2011]
9 The namespace pre x "edm" stands for the Europeana namespace
"http://www.europeana.eu/schemas/edm/".
multilevel hierarchy. The top-most part describes the archival collection as a
whole. The following descendant parts describe sub-parts of the previous parts
with increasing detail.</p>
      <p>The leaves of the descriptive tree are about di erent kinds of unit of records
which constitute the smallest parts within the archival description. The smallest
parts of the description do not necessarily correspond to the smallest parts of
the archival collection. The unit of a record can be, for example, an item which
corresponds to one record, or a le with one or more folders of records.</p>
      <p>A call number for the unit of records is used to order one or more physical
boxes with archival documents (photographs, legal documents, letters, et cetera)
from the archive's depot. Typically, searching for archival material in an archival
nding aid means identifying call numbers for units of records whose potential
relevancy for one's purpose is judged by the contextual descriptions.</p>
      <p>
        The archival descriptions we nd in archival nding aids contain huge and rich
amounts of contextual and implicit information (especially through inheritance)
in order to enable archival users and archivists to e ciently and e ectively locate
and discover archival material [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>EAD-encoded Finding Aids</title>
      <p>The "Encoded Archival Description"10 (EAD) standard is the latest and most
promising standardization e ort for encoding archival nding aids for a
digital environment. It provides the infrastructure to accommodate most designs of
nding aids. Typically, an institution uses a subset of the full EAD model. Our
conversion method is speci cally designed for and tested with APEnet-EAD,
which is currently developed by the APEnet project11 within the context of
Europeana. We expect, however, that our method is applicable to other EAD
dialects with slight modi cations to the script. An EAD-document typically
contains the description of one archival collection in the form of a nding aid. The
&lt;eadheader&gt; element12 contains bibliographic and descriptive information to
identify a nding aid document. Its sibling element &lt;archdesc&gt; holds
information about the archival collection as a whole. In our example, the &lt;archdesc&gt;
element contains several descriptive metadata elds which hold information about
the title of the whole archival fond (&lt;unittitle&gt;), the time span the material
covers (&lt;unitdate&gt;), a call number (&lt;unitid&gt;), the name of the repository
where the material is kept (&lt;repository&gt;), and a summary of the contents
(&lt;scopecontent&gt;).</p>
      <p>Within the &lt;archdesc&gt; element, &lt;c&gt; elements of di erent types (classes,
series, subseries, les, or items) represent the multilevel hierarchy of the archival
description providing the intermediate structure and context for the archival
material described in a nding aid. In our example we nd a series which contains
10 "Encoded Archival Description": http://www.loc.gov/ead/ [7.10.2010]
11 The "Archives Portal Europe" (http://www.apenet.eu/) is a data aggregator for
the European archives.
12 The element has been omitted in gure 1.
a le which holds two items. All these levels have a call number and a title which
are constitutive parts of the contextual description. The two items also link to
digital representations &lt;dao&gt;, e.g. digital images, of their contents. To suit the
original purpose of the nding aids, it is crucial to retain all descriptive and
contextual information about records when transforming this structure to RDF.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conversion of EAD-XML to EDM-RDF</title>
      <p>Archdesc and each c-level are represented as a "edm:ProvidedCHO
ore:Aggregation" cluster (cf. gure 2). Both resources are connected via the
property ore:aggregates. Their URIs are constructed by concatenating the
apenet namespace pre x, the resource type (aggregation-, cho-), the type of the
EAD-level (archdesc, series, le, item), and a guaranteed unique identi er, in
this case the unitid of the respective c-level. By having a uniform URI creation
scheme, objects referring to other objects can be easily represented in RDF by
using URIs as objects. This way, if objects are added or metadata is updated, we
ensure that existing objects receive their old unique URI while added objects
receive a new unique URI. The metadata describing the cultural heritage resource
itself (in our case, those are the di erent parts of the archival description
describing and contextualizing records) can be either attached to the edm:ProvidedCHO
directly or to a third resource called ore:Proxy. The proxy mechanism allows
having di erent views, i.e. descriptions, of one and the same object. In such
case a data provider can create proxies and attach the di erent views to it and
thereby keep them distinct. Europeana itself, when doing semantic enrichment,
will create proxies in order to retain the original structure and provenance
information for the metadata. In our example, the proxy is not necessary as there are
no con icting descriptions. We only created a proxy for the rst level (archdesc)
to demonstrate the functionality.</p>
      <p>The fourth EDM-class shown in gure 2 is edm:WebResource which
represents associated web pages, thumbnail images or any other web resources and is
attached to the aggregation. The URI for such a WebResource is typically the
URL provided for the digital representation.</p>
      <p>The descriptive metadata elds can be represented in EDM in two ways:
In the case where an original eld exactly matches a DC Terms property (for
example &lt;unittitle&gt; and dcterms:title), the DC Terms property is used
directly. In the case where the match is not exact, an APEnet-EAD property is
created in RDF which is speci ed as being a sub-property of the appropriate DC
Terms property: For instance, apenet:callnumber is a rdfs:subPropertyOf of
dcterms:identifier, as shown at the two leaves in gure 2. Interoperability at
the EDM level is ensured through RDFS semantics by using the sub-property
method. The language of the content of descriptive metadata elds can be
speci ed by adding a language tagged-RDF literal as value.</p>
      <p>The edm:ProvidedCHO resource carries not only descriptive metadata but
also properties which are used to relate other objects. During conversion the EAD
hierarchy has been translated into a hierarchy relation between the
edm:ProvidedCHO resources which are connected by dct:isPartOf properties.
This hierarchy mirrors the XML-structure of the multilevel archival description
of the EAD le. At the same time these relations represent, on a more abstract
level, the di erent levels of generality of digital object "packages" submitted via
the EAD le to Europeana. The dct:isPartOf properties conceptually re ect
the documented structure of the archival material, i.e. the archival collection
(archdesc) incorporates a series which has a le which holds two items as parts.</p>
      <p>The two c-levels of type item at the bottom in the XML structure are in
an intentional and meaningful sequential order. This sequence is expressed by
asserting an edm:isNextInSequence statement between the resource with title
"Pagina 1" and the resource with title "Pagina 2".
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion: EAD in a Linked Open Data Environment</title>
      <p>We demonstrated how an EAD-XML encoded archival nding aid can be
modeled in a RDF-based representation. The representation in an RDF graph makes
implicit information explicit, for example, the hierarchical and sequential
relations between the di erent parts of the archival description. We aimed at a
conversion which produces a RDF representation which stays as close as
possible to the original structure of the APEnet-EAD model. This way, we have
a conversion template which is feasible for many di erent variations of EAD
encoded nding aids. The method also entails, however, that not all implicit
information pieces in the descriptive metadata have been made explicit or have
been connected: for example, the unit titles on the di erent levels remain only
indirectly connected to each other. A user being on the level of one of the leafs of
the descriptive tree, most certainly needs to know that he is on page 1 (Pagina
1) of the book "Remissorium Philippi (...)" of the "counts of Holland". In order
to use this data one needs a special reasoner or a Linked Data browser which
brings those information pieces from di erent levels together.</p>
      <p>Another option is to merge intermediate levels of the archival description
without digital representations and the descriptive metadata we nd there into
the leafs of the descriptive tree already during the conversion. In the context of
Europeana, each edm:ProvidedCHO is an object which can be found via searches.
If those objects have no digital representations or only partial descriptions then
their information value in the context of Europeana can be questioned. Data
providers creating mappings to EDM, on the one hand, have to consider how
they want to represent their data in the context of the Europeana information
space, but, on the other hand, enjoy great exibility regarding data modeling
with EDM.</p>
      <p>
        At the same time we showed the capability of the EDM to accommodate such
a particular archival domain model. The EDM is able to accommodate EAD and
other di erent standards as we demonstrated elsewhere [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. One of the main
reasons to use EDM as the ontology (instead of some speci c EAD-RDF model) for
an EAD conversion is, that the archival data can now be connected to museum,
library and other archival data within the Europeana information space.
Contextualization through external resources is now possible. For example, person
names can be linked to concepts in a controlled vocabulary like VIAF13. Such
contextualization allows, for example, to disambiguate meaning or to relate the
original object to other cultural heritage objects annotated with the same person
name. Europeana is planning to do enrichments for a number of elds like, for
example, person or place names.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Concordia</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gradmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siebinga</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Not just another portal, not just another digital library: A portrait of Europeana as an application program interface</article-title>
          .
          <source>IFLA Journal</source>
          <volume>36</volume>
          (
          <issue>1</issue>
          ),
          <volume>61</volume>
          {
          <fpage>69</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Doerr</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gradmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hennicke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meghini</surname>
          </string-name>
          , C., van de Sompel, H.:
          <article-title>The Europeana Data Model (EDM) (</article-title>
          <year>2010</year>
          ), http://www.ifla.org/files/hq/ papers/ifla76/149-doerr-en.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Haworth</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          :
          <article-title>Archival Description: Content and Context in Search of Structure</article-title>
          . In: Pitti,
          <string-name>
            <given-names>D.V.</given-names>
            ,
            <surname>Du</surname>
          </string-name>
          , W.M. (eds.)
          <source>Encoded Archival Description on the Internet</source>
          , pp.
          <volume>7</volume>
          {
          <fpage>26</fpage>
          . Haworth Information Press, Binghamton NY (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hennicke</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olensky</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boer</surname>
          </string-name>
          , V.d.,
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wielemaker</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A data model for cross-domain representation: The "Europeana Data Model" in the case of archival and museum data</article-title>
          . In: Griesbaum,
          <string-name>
            <surname>J.</surname>
          </string-name>
          ,
          <source>Internationales Symposium fur Informationswissenschaft</source>
          <volume>12</volume>
          , .H., Hochschulverband fur Informationswissenschaft (eds.) Information und Wissen: global,
          <source>sozial und frei?</source>
          , vol.
          <volume>58</volume>
          , pp.
          <volume>136</volume>
          {
          <fpage>147</fpage>
          . vwh Hulsbusch,
          <string-name>
            <surname>Boizenburg</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Isaac</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Europeana Data Model Primer (</article-title>
          <year>2010</year>
          ), http://version1.europeana.eu/ web/europeana-project/technicaldocuments/ 13 "Virtual International Authority File": http://viaf.org/ [
          <volume>16</volume>
          .
          <fpage>07</fpage>
          .
          <year>2011</year>
          ]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>