<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Self-contained Information Retention Format For Future Semantic Interoperability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simona Rabinovici-Cohen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roger Cummings</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sam Fineberg</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IBM Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haifa simona@il.ibm.com</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antesignanus roger@antesignanus.com</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>HP Storage fineberg@hp.com</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>4</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Long term preservation of digital information, including machine generated large data sets, is a growing necessity in many domains. A key challenge to this need is the creation of vendor-neutral storage containers that can be interpreted over time. We describe SIRF, the Self-contained Information Retention Format, which is being developed by the Storage Networking Industry Association (SNIA) to support this challenge. We de ne the SIRF components, its metadata, categories and elements, along with some security guidelines. SIRF metadata includes the semantic information as well as schema and ontological information needed to preserve the physical integrity and logical meaning of preservation objects. We also describe how the SIRF logical format is serialized for storage containers in the cloud and for tape based containers. Aspects of SIRF serialization for the cloud are being experimented with OpenStack Swift object storage in the ForgetIT EU project.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Generating and collecting very large data sets is becoming a necessity in many
domains that also need to keep that data for long periods. Examples include
genomics, medical records, astronomy, atmospheric science, photographic archives,
video archives, and large-scale e-commerce. While this presents signi cant
opportunities, a key challenge is providing economically scalable storage systems
to e ciently store and preserve the data. This includes both the data itself as
well as semantic metadata necessary to enable search, access, and analytics on
that data in the far future.</p>
      <p>
        The Storage Networking Industry Association (SNIA) conducted a "100 year
archive" survey. It found that 83% of the organizations surveyed have digital
assets they need to retain for over 50 years, and 53% have information they need
to retain "permanently". Recognizing these challenges, SNIA formed the Long
Term Retention (LTR) group [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to address storage aspects of digital retention.
LTR is working on the Self-contained Information Retention Format (SIRF),
to create a standardized vendor neutral storage format that will help its users
2
interpret preservation objects in the future even by systems and applications
that do not exist today. SIRF provides strong encapsulation of large quantities
of metadata with the data at the storage level, and enables easy migration of
the preserved data across storage devices.
      </p>
      <p>
        Both cloud storage and tape technologies are viable alternatives for storage
of data for the long term. Cloud technology is emerging as an infrastructure
suitable for building large and complex systems, presenting a scalable and
coste ective alternative to the traditional storage systems. Thus, the cloud is clearly
an attractive platform for long term preservation solutions, and in particular,
cloud storage can be leveraged for preservation-aware storage [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Tapes are attractive for long term data retention as their expected lifetime
is higher than that of other types of media and their cost is considerably lower.
Moreover, The SNIA Linear Tape File System (LTFS) takes advantage of a new
generation of tape hardware to provide e cient access to tape using standard,
familiar system tools and interfaces. This paper combines SIRF with cloud
technology, as well as separately combines it with tape technology.</p>
      <p>
        A core standard for digital preservation systems is the Open Archival
Information System (OAIS)4, an ISO standard since 2003 (ISO 14721:2003 OAIS).
OAIS metadata can also include semantic metadata [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to facilitate the
preservation of schemas and ontological information. However, OAIS is a high-level
reference model, which means it is exible enough to be used in a wide variety
of environments. More detailed steps and work ow stages need to be developed
for the implementation of an OAIS based system. SIRF adds more detail to the
metadata needed in the storage container.
      </p>
      <p>
        SIRF uses cases and functional requirements were described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] along with
the substantial di erences from other formats. Our main contribution in this
paper includes the de nition of the SIRF format for long term storage containers.
We de ne the SIRF catalog metadata, its categories and elements along with the
rationale behind them. To show that SIRF can be combined with di erent types
of underlying storage containers, we describe SIRF serialization for the cloud and
SIRF serialization for tapes. We also provide some implementation overview of
SIRF aspects in OpenStack cloud object storage5 that is being examined in the
context of the ForgetIT6 European Union integrated research project.
      </p>
      <p>The rest of this paper is organized as follows. In section 2, we discuss the
business need of storage containers for long term retention. In section 3, we
introduce the SIRF container format, its components and metadata. Section
4 de nes the serialization for cloud and for tapes. Section 5 describes some
aspects of experimental usage of SIRF in ForgetIT project for concise managed
preservation of personal data and organizational web sites. In section 6, we review
related work and conclude with a summary and some future work.</p>
    </sec>
    <sec id="sec-2">
      <title>4 http://public.ccsds.org/publications/archive/650x0m2.pdf</title>
    </sec>
    <sec id="sec-3">
      <title>5 http://www.openstack.org/software/openstack-storage</title>
    </sec>
    <sec id="sec-4">
      <title>6 http://www.forgetit-project.eu</title>
      <p>SIRF
3
2</p>
      <sec id="sec-4-1">
        <title>Business Need for Long Term Retention</title>
        <p>While no one wants to lose their digital content, the cost of maintaining integrity
and access is signi cant, in both money and e ort. And unlike paper based
content, the lifespan of digital content can be very short unless if proactive steps
are being taken to protect it. The use of a storage container format like SIRF
adds little expense and greatly increases the sustainability of data. However, this
is not adequate unless if the cost of preserving content is less than the (potential)
cost of losing it.</p>
        <p>In a business context, there are three major reasons why content is preserved.
These are: to preserve history, to mitigate risk or meet a legal mandate, and for
future value of information. One or more of these may apply, and the amount
an entity is willing to spend will di er depending on how well these reasons are
aligned with the business goals of an organization.</p>
        <p>One of the main reasons why people and organizations preserve content is to
preserve history. In the case of an individual, it may be photos, videos, and other
content preserving one's life history. In a business context, libraries, national
archives, historians, and others have a primary mission to preserve history.</p>
        <p>Another often cited reasons for preserving data is for "risk mitigation", or
in some cases for "legal mandate". These are closely related reasons because
legal mandate is often looked at through the lens of legal risk. For example,
an often cited legal mandate is in healthcare, where medical organizations are
required to retain information for the lifetime of a patient. This seems like a
di cult requirement, especially since records are often maintained in private
doctors' o ces and other places that may not exist 50 or 75 years into the
future. Anecdotal evidence shows that medical records are not maintained that
long. So, why is this happening? It is because records retention is expensive, and
there are no penalties for losing information. That is not to say that doctors and
hospitals don't try, rather they won't spend the necessary money.</p>
        <p>Regarding future value of information, one obvious example is in the
entertainment industry. Movies, TV shows, music, and other content can be re-sold
and repurposed decades after its creation. This can result in many dollars in
revenue. So not surprisingly, organizations like the Motion Picture Expert's Group
are at the leading edge of digital preservation. Entertainment companies spend
signi cant amounts of money retaining their content so that they will have it
available to repurpose. However, this does not mean they can retain everything.
With the advent of digital movie production, the amount of data that can be
generated during the creation of a single lm is immense. Therefore, even here
where future value is tangible, some hard choices need to be made.</p>
        <p>So, how does SIRF help? SIRF brings down the expense of preservation,
because data can remain accessible even if the software that created the data
no longer exists. SIRF reduces the complexity of logical and physical
migration, making it easier for businesses to justify. By using SIRF today, it becomes
possible to retain more information, and to retain information with a lower
perceived future value. This is unlike proprietary and undocumented formats, which
become useless soon after a business stops paying for support.
4
3</p>
      </sec>
      <sec id="sec-4-2">
        <title>The SIRF Format</title>
        <p>Archivists and records managers of physical items such as documents, objects,
records, etc., avoid processing each item individually. Instead, they gather
together a group of items that are related in some manner - by usage, by association
with a speci c event, by timing, and so on - and then perform all of the processing
on that group as a unit. Once assembled, an archivist will place the collection in
a physical container (e.g. a le folder or a ling box of standard dimensions), and
that container is attached with a label that gives an overview of the container
content e.g. name and reference number, date, contents description, destroy date.</p>
        <p>We propose an approach to digital content preservation that leverages the
knowledge of the archival profession and helps archivists remain comfortable
with the digital domain. We de ne a digital equivalent to the physical container
- the archival box or le folder - that de nes a collection, and which can be
labeled with standard information in a de ned format to allow retrieval when
needed. SIRF is intended to be that equivalent - a storage container format for a
set of (digital) preservation objects that includes a catalog with metadata related
to the entire contents of the container as well as to the individual objects and
their interrelationship. This logical container makes it easier and more e cient
to provide many of the processes that will be needed to address threats to the
digital content.</p>
        <p>SIRF is a logical container format for the storage subsystem, appropriate
for the long-term storage of digital information. It is a logical data format of a
mountable unit e.g. a lesystem, a cloud container, an object store, a tape, etc.
It assumes the mountable unit includes an object interface layer that constructs
objects out of the sectors and blocks.</p>
        <p>Figure 1 illustrates the SIRF container, which includes the following
components:
{ A magic object that identi es whether this is a SIRF container and gives its
version. The magic object is independent of the media and has an agreed
de ned name and a xed size. It also includes the means to access the SIRF
catalog (for example, the catalog's location).
{ Preservation objects that contain the actual data to be preserved. An
example preservation object can be the OAIS Archival Information Package
(AIP). The container may include multiple versions of a preservation object
and multiple copies of each version, but each speci c preservation object is
generally immutable.
{ A catalog that is updateable and contains semantically enriched metadata
needed to make the container and its preservation objects portable,
accessible, and understandable into the future without relying on metadata external
to the storage subsystem.</p>
        <p>While traditional storage systems include only limited standardized metadata
about each object, SIRF provides the semantically rich metadata needed for long
SIRF
5
term preservation and interpretation of information, and ensures its grouping
with the data. This rich metadata is de ned in the catalog in a logical format to
allow its serialization for di erent storage technologies. We show its mapping to
some of today's storage containers (cloud storage and tapes), but as new storage
technologies become prevalent in the future, additional mappings will need to
be de ned.</p>
        <p>Fig. 1: SIRF Components
3.2</p>
        <sec id="sec-4-2-1">
          <title>SIRF Catalog Metadata Schema</title>
          <p>The SIRF catalog is an object that includes metadata about the preservation
objects (POs) in the container and their schema and interelationships. It has a
well-de ned standardized format so it can be understandable in the future. The
SIRF catalog is separated from the metadata contained in the POs themselves
because a strict standardized format is di cult to impose on the POs that are
generated by di erent applications and domains. Additionally, the SIRF catalog
includes some metadata that is not included in the PO e.g. xity value of the
whole PO. Including this metadata within the PO changes the xity value of
the PO making this metadata inherently incorrect.</p>
          <p>The SIRF catalog includes metadata related to the whole container as well as
metadata related to each preservation object within the container. Both types
of metadata are divided into categories, elements and attributes organized in
a hierarchical representation. The full metadata de nitions and the rationale
behind them are de ned in SIRF draft speci cation7. Here we provide some
example categories for the whole container metadata in subsection 3.2.1 and for
each preservation object within the container in subsection 3.2.2 below.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>3.2.1 Container Information Metadata Schema. The metadata for the</title>
          <p>whole container includes the categories Speci cation, Container ID, State, and
Container Provenance.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7 http://www.snia.org/tech activities/publicreview, to appear</title>
      <p>S. Rabinovici-Cohen et al.</p>
      <p>The Speci cation category includes information about the speci cation used.
As the speci cation may evolve over time and distinct storage containers may
use di erent SIRF speci cations, it's important to include the exact version of
the speci cation in the SIRF catalog including speci cation ID and speci cation
version.</p>
      <p>The Container ID category includes the container unique identi er such as
the tape ID for tape based storage containers or cloud container ID in case of
cloud storage.</p>
      <p>The State category is an indication of the progress of any activities that are
to be carried out against a container. For example, if a container holds many
preservation objects, state may indicate whether all of the objects intended for
a container have been included or not. Or, state may indicate an in-process
migration of a container. Multiple state entries are allowed in case if there are
multiple pending activities.</p>
      <p>The Container Provenance category is metadata describing the history of
the information in a SIRF container (e.g., its origins, chain of custody,
preservation actions and e ects). The Provenance information may vary depending
on the type of information being preserved or its intended audience and it may
be large. Therefore, it is included in the catalog by reference, and the actual
information is stored in another preservation object. The container provenance
information stored in SIRF may be in the W3C-PROV format, or any other well
known provenance format. Regardless of the perspective from which provenance
metadata is derived, it is critical for understanding the container, its history, its
context and meaning.</p>
      <sec id="sec-5-1">
        <title>3.2.2 Object Information Metadata Schema. The metadata for each preser</title>
        <p>vation object includes several categories; from which we'll describe here: Object
IDs, Fixity, and Audit log.</p>
        <p>The Object Identi ers (IDs) category is used to identify a PO and to link
to other POs. Managing identi ers over the long term raises issues such as: how
to ensure uniqueness of identi ers over long term, how to handle evolution of
identi ers over time, how to ensure scalability of identi ers.</p>
        <p>SIRF helps to address these issues by enabling redundancy in identi ers and
registering the evolution (genealogy) of POs. Hence, a PO in a SIRF container
can have multiple identi ers as redundant identi ers. This increases the chances
that at least one of the identi ers will survive for the long term. Nevertheless,
at any time, at least one of the identi ers should be persistent and unique.</p>
        <p>Fixity is used to demonstrate that the content information has not been
altered in an undocumented or unauthorized manner. The xity information can
be seen as an integrity check value. Fixity is sometimes computed via simple
cheap functions such as a CRC, or it can include a stronger and more
expensive (in execution time and space) cryptographic hash function such as MD5 or
SHA-512. No matter how strong the xity computation functions are, they are
likely to become obsolete in the far future when larger amounts of storage and
stronger computing power are available. Thus, the preservation system should
SIRF
7
be allowed to update xity functions in the future, as existing ones become
obsolete. Consequently, the SIRF catalog allows for multiple xity algorithms and
values for a given PO.</p>
        <p>The audit log category is provided as a place for preserving any important
information about how an object has been accessed or modi ed. The extent and
contents of an audit log depend on the needs of the speci c preservation data
store and its use case. Distinct domains have di erent audit logs regulations e.g.,
SEC is for the US nancial market domain, FDA is for the US medical domain.
In SIRF, audit logs are stored in the catalog as links to preservation objects.
3.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>SIRF Container Security Guidelines</title>
        <p>Some of the legal mandates for information retention also incorporate
requirements for privacy and access protection. Where such security-based requirements
exist, they add another level of complexity to long-term retention of the SIRF
container. Much of this additional complexity results from the fact that the
security-based requirements tend to mitigate against other retention
requirements. For instance, while retention generally seeks to make information widely
available and usable, security tends to restrict access to ensure that information
privacy is maintained.</p>
        <p>Information security also adds signi cantly to the amount of metadata that
must be maintained within the container to ensure future usability of the
information. Most obvious is the need to identify the encryption scheme used, and
the need to maintain information about the di erent types of access that should
be granted to the information. All access information needs to be based on the
de nition of abstract roles rather than speci c people because, given the time
periods being addressed by long-term retention, people will change job
functions, organizations will grow, merge, or disappear, and uses for the information
may signi cantly alter. A long-term retention system must be able to
continuously add new users and associate them with existing roles, and change the roles
assigned to existing users.</p>
        <p>The management of keying information, whether related to the encryption
of information or to the authentication of the roles assigned to speci c users,
presents a speci c challenge in terms of long-term retention. Clearly such
information cannot directly be located within the container itself, but su cient
metadata must be included in the container to allow the keying information to
be located, validated, and veri ed.</p>
        <p>ISO/IEC 27040 draft is being created to address the security of both local
and cloud-based security systems. It emphasizes that there are integrity,
authentication, and privacy threats that are particular to long-term storage systems.
It also notes that the long lifetime of information within such systems enables
attacks that require a large amount of access to the information but which can
be disguised as many small requests over an extended period of time. It
highlights the importance of maintaining a log of attack attempts, compromises, and
system and user changes, and notes that such a log must also be maintained for
the long-term. In the current version of SIRF, we support some initial security
guidelines via e.g., the Fixity and the Audit Log categories.
4</p>
        <sec id="sec-5-2-1">
          <title>SIRF Serialization for Cloud and for Tape</title>
          <p>The SIRF serialization for cloud/tape speci es how a cloud container or a tape
container becomes SIRF-compliant. A SIRF-compliant cloud container or tape
container enables future's cloud/tape clients to "understand" containers created
by today's cloud/tape clients even though the properties of the future client is
unknown today. By "understand", we mean we can identify the preservation
objects in the container, the packaging format of each object, its xity values,
etc. (as de ned in the SIRF catalog).</p>
          <p>For the concrete serialization we chose speci c standard based storage
containers. For the cloud, we chose CDMI8 and OpenStack object storage while for
tapes we chose LTFS9 based tapes. No single technology will be usable over the
time spans mandated by current digital preservation needs. SNIA CDMI and
LTFS technologies are among best current choices, but are good for perhaps
1020 years. SIRF provides a vehicle for collecting all of the information that will be
needed to transition to new technologies in the future, and it can be serialized
for future technologies as they emerge.</p>
          <p>For the serialization step, we classify the preservation objects as either simple
preservation object or composite preservation object. A simple PO contains just
one element and is mapped to one object in the CDMI cloud or one le in the
LTFS tape. A simple PO can be for example a jpg photo or a tar le. A composite
PO contains several elements and a manifest that combines the elements. The
composite PO is mapped to several objects in the CDMI cloud or a number of
les in the LTFS tape.
4.1</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>Serialization For Cloud Storage</title>
        <p>The Cloud Data Management Interface (CDMI) is an ISO/IEC 17826:2012
standard created by SNIA that de nes an interoperable format for moving data
and associated metadata between cloud providers. CDMI has several
implementations including an open source implementation for OpenStack Swift10 cloud
storage.</p>
        <p>A CDMI cloud container can be quali ed as a SIRF container when:
{ The SIRF magic object is mapped to the CDMI container metadata.
{ The SIRF catalog is an object in the CDMI container formatted in JSON
(self-describing) that includes one containerInformation section and multiple
objectInformation sections - one for each PO within the container
(selfcontained). This object should be indexed (if possible). There is a CDMI
extension to support indexing with object granularity.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>8 Cloud Data Management Interface - http://www.snia.org/cdmi</title>
    </sec>
    <sec id="sec-7">
      <title>9 Linear Tape File System - http://www.snia.org/ltfs 10 Swift - https://wiki.openstack.org/wiki/Swift</title>
      <p>SIRF
9
{ A SIRF PO that is a simple object (contains one element) is mapped to a</p>
      <p>CDMI data object.
{ A SIRF PO that is a composite object is mapped to a set of data objects
(one for each element) and a manifest data object that includes information
about the elements.</p>
      <p>The interface to the SIRF-compliant CDMI container is the ordinary CDMI
Application Program Interface (CDMI API). In addition, the CDMI API can be
used to store and access the various preservation objects and the catalog object.</p>
      <p>For example, gure 2 depicts a CDMI container named "Patient Container"
that is SIRF-compliant and includes medical encounters and images for the
patient. Assume each encounter is a simple preservation object; each image is a
composite preservation object; and since the container is SIRF-compliant, it also
includes a catalog object.
The Linear Tape File System (LTFS) format speci cation de nes LTFS Volumes.
An LTFS Volume holds data les and corresponding metadata to completely
describe the directory and le structures stored on the volume. Files can be written
to, and read from, an LTFS Volume using standard POSIX le operations. The
LTFS Volume includes an index in XML that contains metadata similar to
information in disk-based le systems such as le name, dates, extent pointers,
extended attributes, etc. LTFS is becoming the standard for linear tape and is
being formalized through SNIA.</p>
      <p>An LTFS volume is comprised of a pair of LTFS partitions: a data partition
(DP) and an index partition (IP). Each partition contains a Label Construct
followed by a Content Area. As depicted in gure 3, a LTFS tape container can
be quali ed also as a SIRF container when the volume format is as follows:
{ The SIRF magic object is mapped to extended attributes of the LTFS index
root directory.
{ The SIRF catalog resides in the index partition and formatted in XML
(self-describing) that includes one containerInformation section and
multiple objectInformation sections - one for each PO within the container
(selfcontained). LTFS application has rules to indicate what to store in the index
partition. That method can be used to indicate to store the SIRF catalog in
the index partition. Alternatively, the index partition can include a reference
to the SIRF catalog that will reside in the data partition.
{ A preservation object (PO) is mapped to an LTFS le or set of les. In case
the PO is a simple object composed of one element, it is mapped to a LTFS
le. In case the PO is a composite object composed of several elements, it
is mapped to a set of LTFS les (one for each element) and a manifest le
that its content includes information about the elements.
The European Union integrated project ForgetIT investigates ways for concise
long term digital preservation and its adoption for personal data and
organizational web sites. It combines three new concepts: managed digital forgetting
inspired from human brain and cognitive psychology; smooth transition between
data active use and its preservation; contextualized remembering keeping the
archive understandable and useful.</p>
      <p>
        The ForgetIT Preserve-or-Forget framework uses the DSpace open source
as its preservation system, where the archival storage is Preservation
DataStores (PDS) in the Cloud [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] that provides preservation-aware storage services
based on the OAIS model. PDS includes the Preservation Engine and the Storlet
Engine. The Preservation Engine transforms the logical OAIS functions and
information objects into processes and physical storage objects. The Preservation
Engine sometimes requires performing data-intensive computational tasks, such
as transformation, migration, xity checks, and data analysis. When the
Preservation Engine requires performing such tasks, it uses storlets - computational
SIRF
11
modules running in a sandbox close to the data. O oading OAIS-based
functionality to the storage decreases probability of data loss, simpli es the applications
and supports automation of preservation processes.
      </p>
      <p>
        The Storlet Engine [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] provides the cloud storage with a capability to
include storlets that run within the storage in a sandbox that provides isolation.
It is plugged into a private cloud or object storage such as OpenStack Swift
and provides a powerful extension mechanism that makes the storage exible,
customizable and extensible. By using storlets, the client bene ts of reduced
bandwidth (reduce the number of bytes transferred over the WAN), enhanced
security (reduce exposure of sensitive data), cost saving (reduce infrastructure
at the client side), and compliance support (improve provenance tracking).
      </p>
      <p>PDS in ForgetIT implements some aspects of SIRF. It creates the various
identi ers used for maintaining the evolution of POs, which can be stored in the
Object IDs category in the SIRF catalog.</p>
      <p>Regarding the xity category in the SIRF catalog, PDS developed a xity
storlet that can compute multiple xity values for each PO, and new hash
functions can be uploaded to the storage as older ones become too weak or even
obsolete.</p>
      <p>While the ForgetIT POs are generated by di erent applications and domains
(personal and organizational use cases), the SIRF catalog presents a standardized
format that can be interpreted in the future.
6
6.1</p>
      <sec id="sec-7-1">
        <title>Discussion and Conclusions</title>
        <sec id="sec-7-1-1">
          <title>Related Work</title>
          <p>
            Storage aspects of archiving and preservation systems have been the focus of a
growing number of studies. You et al. [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] present PRESIDIO, a scalable archival
storage system that e ciently stores diverse data. Adams et al. [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] studied
scienti c and historical archives, covering a mixture of purposes, media types, and
access models. Based on this study, they identify areas for improving the e
ciency and performance of archival storage systems.
          </p>
          <p>
            Long-term preservation systems di er from traditional storage applications
with respect to goals, characteristics, threats, and requirements. Baker et al. [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]
examine these di erences and suggest bit preservation guidelines and alternative
architectural solutions that focus on replication across autonomous sites. Storer
et al. [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] discuss security threats that arise when storing data for long periods of
time. This includes common threats such as loss of integrity, failure of
authentication and compromise of privacy, as well as new speci c threats such as slow
attacks.
          </p>
          <p>
            Dappert and Enders [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] discuss the importance of metadata in a long term
preservation solution. The authors identify several categories of metadata,
including descriptive, preservation related, and structural, arguing that no single
existing metadata schema accommodates the representation of all categories.
The work surveys metadata speci cations contributing to long-term
preservation.
          </p>
          <p>S. Rabinovici-Cohen et al.</p>
        </sec>
        <sec id="sec-7-1-2">
          <title>Conclusions and Future Work</title>
          <p>Moving forward, digital content preservation will have many technical and
cultural challenges. As digital technologies continue to replace physical ones, these
challenges must be solved to prevent us from losing a generation of content.</p>
          <p>SIRF, the Self-contained Information Retention Format, was developed to
address the growing necessity to preserve digital information over long periods
of time. SIRF does this by acting as the digital equivalent of an archivist's
"box". SIRF preserves data and metadata as a single unit and provides a catalog
containing the basic metadata needed to access and preserve content. This aids
in the future understanding of data, and in the migration to new storage devices
and formats.</p>
          <p>We have shown that the SIRF can be serialized for a variety of storage
technologies including LTFS based tape and CDMI cloud containers. This should
provide a means for preserving information for the next years, and a vehicle for
migrating to whatever new storage technologies become prevalent in the future.</p>
          <p>In future work, we would like to improve support for the security guidelines
developed in ISO/IEC 27040. Also, we would like to experiment SIRF in other
projects and serialize it for additional storage containers.</p>
          <p>Acknowledgments. This work was partially funded by the European
Commission in the context of the FP7 ICT project ForgetIT (under grant no: 600826).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>1. SNIA Long Term Retention (LTR) group</article-title>
          . URL: http://www.snia.org/ltr
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rabinovici-Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pease</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>PDS Cloud: Long Term Digital Preservation in the Cloud</article-title>
          .
          <source>In: IC2E 2013: Proceedings of the IEEE International Conference on Cloud Engineering</source>
          , San Francisco, CA (
          <year>March 2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Brunsmann</surname>
          </string-name>
          , J.:
          <article-title>Product Lifecycle Metadata Harmonization with the Future in OAIS Archives</article-title>
          .
          <source>In: DC 2011: Proceedings of the International Conference on Dublin Core and Metadata Applications</source>
          , Hague, The
          <string-name>
            <surname>Netherlands</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Rabinovici-Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cummings</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fineberg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marberg</surname>
          </string-name>
          , J.:
          <string-name>
            <surname>Towards</surname>
            <given-names>SIRF</given-names>
          </string-name>
          :
          <article-title>Self-contained Information Retention Format</article-title>
          .
          <source>In: SYSTOR 2011: Proceedings of the International Systems and Storage Conference</source>
          , Israel (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Rabinovici-Cohen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henis</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Storlet Engine: Performing Computations in Cloud Storage</article-title>
          .
          <source>IBM Technical Report</source>
          H-
          <volume>0320</volume>
          (
          <year>August 2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>You</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pollack</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gopinath</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>PRESIDIO: A Framework for E - cient Archival Data Storage</article-title>
          .
          <source>ACM Transactions on Storage</source>
          <volume>7</volume>
          (
          <issue>2</issue>
          )
          <issue>(</issue>
          <year>July 2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>I.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Storer</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Analysis of Workload Behavior in Scienti c and Historical Long-Term Data Repositories</article-title>
          .
          <source>TOS</source>
          <volume>8</volume>
          (
          <issue>2</issue>
          ) (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthak</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roussopoulos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maniatis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giuli</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bungale</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>A Fresh Look at the Reliability of Long-Term Digital Storage</article-title>
          .
          <source>In: Proceedings of the 1st ACM SIGOPS European Systems Conference</source>
          . (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Storer</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenan</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voruganti</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <string-name>
            <surname>POTSHARDS - A Secure</surname>
          </string-name>
          ,
          <article-title>Recoverable, Long-Term Archival Storage System</article-title>
          .
          <source>TOS</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ) (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dappert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enders</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Digital Perservation Metadata Standards</article-title>
          . Information Standards Quarterly,
          <source>Special Issue on Digital Preservation</source>
          <volume>22</volume>
          (
          <issue>2</issue>
          ) (
          <year>2010</year>
          )
          <volume>4</volume>
          {
          <fpage>12</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>