<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semi-Automated Preservation and Archival of Scientific Data using Semantic Grid Services</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jane Hunter</string-name>
          <email>jane@dstc.edu.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sharmin Choudhury</string-name>
          <email>sharminc@dstc.edu.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DSTC, The University of Queensland</institution>
          ,
          <addr-line>Brisbane</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2003</year>
      </pub-date>
      <abstract>
        <p>Addressing the long term preservation issues associated with scientific data is a complex challenge compounded by: the scale and multidisciplinary nature of the problem; the wide range of formats and data types involved; and the often proprietary and transitory nature of the hardware and software used to generate the data. In this paper we present the PANIC system - an integrated, extensible architecture based on preservation metadata, automatic notification services, software and format registries and semantic grid services - that we believe, offers a sustainable, dynamic approach to the long term preservation of large collections of heterogeneous scientific data.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Preservation</kwd>
        <kwd>Semantic Grid Services</kwd>
        <kwd>Scientific Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Addressing the preservation and long -term access
issues associated with digital scientific data is one of
the key challenges facing research, government and
scientific organizations today. High performance grid
computing, space and earth observation sciences and
large scientific experiments produce enormous
quantities of data that require effective and efficient
management. Digital data files and information objects
require constant and expensive maintenance because
they depend on hardware, software, data, models and
standards which are upgraded or replaced every few
years. Accelerating rates of data collection and content
creation and the increasing complexity of digital
resources means that many organizations can no longer
keep pace with the preservation needs of all of the data
entrusted to them.</p>
      <p>In the field of science the problem is compounded
due to heterogeneous nature of the large quantities of
data being generated from experiments and
observation. For example, astronomers are producing
Flexible Image Transport System (FITS) files and
VOTable compliant XML files [1]; geoscientists are
producing Arc/Info SHP files and GeoTIFF images;
bioscientists are generating huge genomic databases
and associated EST (expressed sequence tags), GSS
(genome survey sequence) and HTGS (high throughput
genomic sequence) files; medical scientists are
generating vast sets of DICOM images; weather
researchers are generating HDF5 datasets.</p>
      <p>The task of ensuring long term access to scientific data
collections is so overwhelming that scientists spend
much of their time managing the data through special
purpose handcrafted solutions, rather than using their
time effectively for scientific investigation and
discovery. This has been partly due to the lack of
available tools and services but largely due to the
heterogeneous nature of scientific data, not only
between different science domains but also within
science domains. Organizations such as CODATA [2]
have been active in promoting improved management
of digital scientific data and the importance of
preserving contextual information with the archived
scientific datasets. In particular CODATA recommend
the use of the OAIS model [3], a high-level conceptual
model developed by a consortium of space agencies, to
facilitate scientific data management. However to date
there has been limited availability of practical,
available preservation tools and services . This is
beginning to change primarily as a result of activities
in the digital library domain. For example, Cornell’s
Virtual Remote Control (VRC) project [4] and OCLC’s
INFORM [5] project are developing risk measurement
and notification services. The Global Digital Format
Registry (GDFR) [6] initiative, the UK National
Archive’s PRONOM project [7] and VersionTracker
[8] are developing format and software registries that
can be used to determine required preservation actions.
The UK Digital Curation Centre is developing a
Representation Information Repository using ebXML
[9]. Projects such as the Typed Object Model (TOM)
[10] and IBM’s UVC Emulation project [11] are
generating migration and emulation services. Many
scientific communities are developing their own sets of
migration services (e.g., converter programs ).
Currently each of these components is being developed
independently but altogether they can be leveraged to
help build a complete preservation solution.</p>
      <p>Moreover, it is generally recognized that there is no
single best solution to digital preservation. Differences
in the needs and practices of various scientific
disciplines, make it difficult if not impossible to define
a ‘one size fits all’ approach to selecting, appraising
and retaining scientific data. [12] The most
appropriate strategy depends on the particular
requirements of the custodial organization, the
producers and users of its collection and the nature of
the objects in the collection. Hence within the PANIC
project [13] we combine the efforts of the different
domain-specific preservation initiatives by integrating
the range of tools and services being developed into a
single encompassing Grid framework. More
specifically PANIC uses a flexible, dynamic,
semiautomated approach which provides access to a range
of metadata tools and risk assessment, notification,
emulation and migration services through a Semantic
Web/Grid services architecture.</p>
      <p>The remainder of the paper is structured as follows.
The next section describes the system objectives
through a motivational example. Section 3 describes
the overall system architecture. Section 4 concludes
with an evaluation of the results and a discussion of
problem issues and future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Motivational Example</title>
      <p>Russel Coight is an astronomer at the Australian
Telescope National Facility (ATNF). He has a massive
collection of legacy FITS (Flexible Image Transport
System) files. The International Virtual Observatory
maintains online registries of the latest recommended
formats, format versions and software tools and
services, (for authoring, rendering, viewing, and
converting files) for the astronomy community. The
PANIC system periodically compares the metadata
associated with files and information objects ni the
ATNF collection with the IVO’s registries. PANIC
determines that IVO now recommends that FITS
format files should be replaced with VOTable 1.1
format because the latest version of the Xanadu editing
and analysis software will no longer support FITS f iles.
The system sends an email to Russel Coight notifying
him that certain datasets that he owns are in danger of
becoming obsolete. The email message includes a list
of the endangered FITS files. The message also
recommends that the FITS format be replaced by
VOTable 1.1 which has now become the defacto
astronomical data standard, recommended by the IVO.
Russel Coight must now find a FITS-to- VOTable
conversion service that meets all his service quality
parameters. He uses the PANIC system to specify the
parameters he requires in the conversion service. For
example, he specifies that he requires a free service
that converts from FITS to VOTable 1.1. He prefers a
distributed converter that will process multiple files
concurrently using parallel processors on the Grid. He
also specifies that the distributed converter must be
highly reliable, high-speed and not result in any loss in
data quality. His request is handled by a Discovery
Agent which searches a Grid Service registry for the
appropriate service description. The Discovery Agent
cannot find any exact matches to Russel’s request but it
can find one near match and two conversion services
which can be chained to approximate the required
service:
1. A service which converts from FITS to VOTable
1.1 but which is lossy - developed by NASA’a
HEASARC
2.</p>
      <p>One service which converts from FITS to
VOTable 1.0 and another which converts from
VOTable 1.0 to VOTable 1.1– both are lossless,
reliable and high speed and can be distributed
across GRID processors.</p>
      <p>
        The Discovery Agent pres ents these alternatives to
Russel, ranked according to how well they match his
request. Russel is able to choose his preferred option
and the selected Provider Agent then executes this
service, (ideally processing multiple files in parallel
across Grid procesors e.g., [
        <xref ref-type="bibr" rid="ref6">14</xref>
        ]) and returns the
VOTable 1.1 files. After migration is complete, the
associated provenance and events metadata (which
records a history of preservation actions associated
with each digital object in the collection) is also
automatically updated.
      </p>
      <p>The next two sections of this paper describe the
architecture, components and implementation details of
the PANIC system that we have developed in order to
turn the s cenario outlined above into reality.</p>
      <sec id="sec-2-1">
        <title>Notification component</title>
        <p>Notification
Service
Registry(s)
Requester
Agent</p>
      </sec>
      <sec id="sec-2-2">
        <title>Discovery component</title>
      </sec>
      <sec id="sec-2-3">
        <title>Provider component</title>
        <p>Discovery</p>
        <p>Agent</p>
        <p>Preservation</p>
        <p>Service</p>
        <p>Registry
Retrieve and Invoke
Appropriate Service(s)</p>
        <p>Preservation</p>
        <p>Service
Provider
Agent
Preservation
Web Services</p>
        <p>TOM
XENA
UVC</p>
        <p>
          There is strong agreement that preservation
metadata is crucial to the long term preservation of
digital objects. CODATA [12] recommend that
scientific organizations adopt the OAIS model [3].
METS [
          <xref ref-type="bibr" rid="ref7">15</xref>
          ], a Digital Library Federation initiative,
builds on OAIS and provides an XML document
format for encoding metadata necessary for both
management and exchange of digital objects within
and between repositories. Consequently the first
phase of PANIC involved developing:
• a preservation metadata schema based on METS
but with extensions for audiovisual and
discipline-specific needs to support a wide
variety of atomic and composite digital objects;
• a preservation metadata capture tool (PREMINT)
based on this schema.
        </p>
        <p>
          Although METS is the more widely used
preservation metadata schema, MPEG-21 [
          <xref ref-type="bibr" rid="ref8">16</xref>
          ] has
also been applied successfully as a preservation
metadata format [
          <xref ref-type="bibr" rid="ref9">17</xref>
          ]. Hence we decided, for
purposes, to develop an MPEG-21 compliant schema
to support our preservation metadata requirements.
        </p>
        <p>
          Based on these metadata schemas, we developed
both a stand-alone Java application and a JSP (Java
Server Pages) version of the PREservation Metadata
INput Tool (PREMINT) metadata input tools. The
application consists of a set of metadata input forms,
constrained by the underlying XML Schema.
PREMINT collects metadata by dynamically
presenting the user with a series of forms that collect:
Descriptive Metadata, Technical Metadata,
Instrumentation Metadata. In order to streamline the
preservation metadata capture process we plan to
integrate services such as JHOVE, the
JSTOR/Harvard Object Validation Environment [
          <xref ref-type="bibr" rid="ref10">18</xref>
          ],
to automatically extract the format-specific technical
metadata and services that automatically record
scientific instrument settings (e.g., OM E Open
Microscopy Environment) to capture precise
provenance data. Figure 2 shows a screenshot of the
PREMINT tool, illustrating the technical metadata
input form for a video file. Users have the option of
saving the metadata output to either
METS+Extensions or MPEG-21.
3.2.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Obsolescence</title>
    </sec>
    <sec id="sec-4">
      <title>Notification</title>
    </sec>
    <sec id="sec-5">
      <title>Detection and</title>
      <p>The Obsolescence Detection module periodically
compares the preservation (formatting) metadata for
each object in the collection with information stored
in the following three registries:
•</p>
      <p>Software Version Registry – this contains
information about the latest versions of
authoring, rendering, viewing, editing or analysis
software required to access and use objects in the
collections.</p>
      <p>For each softwa re tool, the registry sto res: Title,
Description, Creator, CurrentVersion,
ReleaseDate, DeveloperPage, License,
Requirements, DownloadSite, DownloadSize,
Rating, EaseOfUse, FormatsSupported,
Features, Stability, Price.</p>
      <p>
        Format Registry - this contains detailed
information about digital formats including:
Identifier, Description, Version, Author, Owner,
RelationshipsOtherFormats,
ApplicationsUsingThisFormat;
FormatSpecification; ProvenanceEvents etc [
        <xref ref-type="bibr" rid="ref11">19</xref>
        ]
Recommended Format Registry – this tracks the
latest recommended preservation formats and the
authority making the recommendation e.g ., “the
International Virtual Observatory (IVO)
recommends VOTable as the preferred
preservation format for astronomical data”.
      </p>
      <p>For the purposes of demonstrating the PANIC
prototype, we have developed three MySQL
databases containing sample data, to represent the
registries. However there are existing initiatives
focusing on developing and maintaining the
corresponding real-world registries. For example,
Global Digital Format Registry (GDFR) and
PRONOM are both developing digital format
registries for the purposes of long term preservation.
These mainly focus on digital library applications and
would have to be extended to include information
relevant t o scientific data formats . VersionTracker [8]
also maintains a website with a human searchable
registry of software versions that enables users to
determine whether they should update, upgrade or
patch their existing applications. Again, this mainly
focuses on commercial software. We envisage that
each specific scientific community (e.g.,
astronomers) would have an authoritative body (e.g.,
IVO) responsible for maintaining registries of
recommended formats and available software
versions.</p>
      <p>When there is an incompatibility between a
digital object’s current preservation/formatting
metadata and the latest version recommended in the
registry, then a message is sent to the owner of the
data or some nominated person(s) or software agent,
notifying them of a potential risk. Figure 3 is an
example of a notification window.
3.3.</p>
    </sec>
    <sec id="sec-6">
      <title>Preservation Service</title>
    </sec>
    <sec id="sec-7">
      <title>Discovery and Invocation</title>
    </sec>
    <sec id="sec-8">
      <title>Description,</title>
      <p>Our aim is to build a system which dynamically
incorporates the expanding range of preservation
services available and also provides decision-support
tools or recommender services which can assist the
scientific data manager to select the best single
service or combination of services for a particular
digital object or a particular set of circumstances.</p>
      <p>The modular, distributed nature of the Semantic
Grid/Web services architecture makes it perfectly
suited to the dynamic, large-scale, heterogeneous
nature of the digital preservation problem. A key
objective of the PANIC project is to test this
hypothesis by developing and evaluating a
semiautomated preservation system based on the
Semantic Web services architecture which provides
access to a suite of independent preservation service
components which can be discovered, linked, and
invoked in arbitrary combinations across the Grid to
fulfil the specific preservation tasks and
requirements of different scientific organizations.
3.3.1</p>
      <sec id="sec-8-1">
        <title>Semantic Web/GridServices</title>
        <p>
          Web services are enabling networked computer
programs to process and consume information. Based
on open standards such as XML, SOAP and WSDL,
Web services provide a standardized way of enabling
Web-based application-to-application
interoperability. More recently the Semantic Web
services initiative has developed OWL-S/DAML-S,
an OWL ontology which enables Web services to be
described semantically and their descriptions to be
processed and understood by software agents. A
number of projects are using OWL-S to describe their
domain-specific services and enable software agents
to automatically discover, compose, invoke and
monitor the most appropriate Web services [
          <xref ref-type="bibr" rid="ref12 ref14">20, 21</xref>
          ].
As far as we are aware, no one is currently applying
or extending OWL-S/DAML-S to generate semantic
descriptions of digital preservation services so that
they can be discovered, invoked and composed by
software agents in order to automate the preservation
tasks of large archival organizations.
3.3.2
        </p>
        <p>OWL-S</p>
      </sec>
      <sec id="sec-8-2">
        <title>Services</title>
      </sec>
      <sec id="sec-8-3">
        <title>Ontology for</title>
      </sec>
      <sec id="sec-8-4">
        <title>Preservation</title>
        <p>The purpose of OWL-S is to provide
computerinterpretable descriptions of services so that they can
be located, selected, employed, composed and
monitored automatically over the Internet. Multiple
web services can be matched and chained
interoperating to perform complex tasks and
transactions for users dynamically and on-demand.</p>
        <p>The OWL-S ontology has a top-level Service
ontology with three main subontologies :
• ServiceProfile – provides a description of what
the service does, enabling advertising and
discovery;
• ServiceModel – provides a detailed description
of a service’s opera tion or how it works;
• ServiceGrounding – provides details of how to
interoperate with or access a service using
messages.</p>
        <p>
          The advantage of OWL-S is that it provides
generic upper-level classes that can be refined to
describe any Web service. We have extended the
OWL-S classes to create more preservation-specific
subclasses. Figure 4 illustrates how we have extended
the generic Service class by defining a
PreservationService subclass. PreservationService
has two subclasses – emulation and migration. These
new types of service are defined in the
PreservationService ontology, which extends the
Service ontology provided by the OWL-Service
Coalition [
          <xref ref-type="bibr" rid="ref15">22</xref>
          ]. The normalization service is defined
as a further subclass of migration.
        </p>
        <p>Within the ServiceProfile ontology, the Profile
class provides three types of information:
• Service name, description and contact (person or
organization);
•
•</p>
        <p>Functional description in terms of inputs,
outputs, pre-conditions and effects;
An extensible set of properties used to describe
features of the service e.g., service category,
quality rating, etc.</p>
        <p>We have also extended the ServiceProfile
ontology to create a PreservationServiceProfile.</p>
        <p>
          Semantic Matchmaker [
          <xref ref-type="bibr" rid="ref16">23</xref>
          ] is used as the
Discovery Agent in PANIC. When the collections
manager (see Figure 1) specifies the parameters
required in the preservation web service, a query is
created and submitted to Semantic Matchmaker.
ServiceProfiles are used to match service requesters
to service providers. Requests from the service
requesters, are converted to ServiceProfile documents
Service
PreservationService
subClassOf
Migration
        </p>
        <p>ExecutionStatus
SystemRequirment</p>
        <p>Creator
ReleaseDate
ServiceQuality</p>
        <p>Speed
Reliability
Emulation</p>
        <p>Download
e.g. Windows
XP
e.g. John
Doe
e.g. 8-12-2003
e.g. Low
e.g High
subClassOf
Normalisation</p>
        <p>OriginalObjectFormat
OriginalObjectVersion e.g. 5.12
TargetObjectFormat
TargetObjectVersion
e.g. TIFF
e.g. JPEG
2000
e.g. 2.02
Lossiness
e.g. lossless</p>
        <p>EmulatedObject
EmulationType
SystemSetting
e.g. MAC OS
e.g. OS
e.g. 256 bit
palette
and compared against the stored ServiceProfiles for
available services. A ranked list of matching services
is retrieved and displayed. Matching services can be
either atomic o r chained, composite services.</p>
        <p>
          For example, a user might specify that he requires
a service that converts from FITS to VOTable 1.1. He
also prefers a distributed converter that must be
highly reliable, high-speed and not result in any loss
in data quality. His request is forwarded by the
Requester Agent to a Discovery Agent which
searches a Web Service registry for a service
description matching this specification. In the past
UDDI registries [
          <xref ref-type="bibr" rid="ref17">24</xref>
          ] were used to advertise available
Web services but dynamic discovery was difficult
due to the lack of semantics. Using OWL-S and
Semantic Matchmaker, enables more precise and
dynamic discovery of appropriate services. In this
case, the Semantic Matchmaker determines that there
are two matches to the service request: a simple
process (which converts directly from FITS to
VOTable 1.1) and a composite, chained process that
first converts FITS to VOTable 1.0 and then converts
VOTable 1.0 to VOTable 1.1 . The matching service
descriptions (the WSDL, ServiceGrounding,
ServiceProfile and ServiceProcess documents) are
sent back to the Requester Agent . These are then used
to present the search results to the user as shown in
Figure 5.
        </p>
        <p>Given the results of the search and the
recommendations of the Discovery Agent, the
collections manager can (through the system
configuration interface) choose to allow the system to
automatically invoke the best matching service or
interactively select a particular preservation action
and invoke it manually. The system configuration
interface enables the collections manager to set
certain runtime parameters prior to service execution
e.g., whether to process multiple files in parallel,
where to save the output files, whether to update
preservation metadata, where to email the logfile etc.</p>
        <p>
          After service selection, the Requester Agent sends
the inputs (e.g., FITS files) to the Provider Agent
which schedules and executes the conversion services
(ideally across multiple parallel procesors using a
Grid DataFarm such as [
          <xref ref-type="bibr" rid="ref6">14</xref>
          ]) and returns the outputs
(e.g., VOTable 1.1 files) to the Requester Agent. The
Requester Agent saves the output files locally to the
specified location and updates the preservation action
metadata – recording what files were converted,
when, authorized by whom, and the service that was
used. Finally an email is sent to the user, notifying
him that the migration of FITS files has been
completed.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Conclusions</title>
      <p>In this paper we have briefly described how PANIC,
a prototype preservation system which we have
developed based on : preservation metadata; software
and fo rmat registries ; and Semantic Grid Services ,
can streamline the long term preservation of scientific
data.</p>
      <p>The distributed nature of the pro posed Web/Grid
services architecture offers many advantages. It
leverages existing work on preservation metadata and
preservation software tools (e.g., emulation and
migration services) by integrating them and making
them available through a single interfa ce. It enables
institutions to coordinate and share their digital
preservation activities whilst retaining the flexibility
to meet local requirements. The proposed system is
scalable and extensible. It has the potential to provide
preservation services for a very wide variety of data
and media formats. Because the system is based on
standards including: METS, MPEG-21, XML,
SOAP, WSDL, UDDI, OWL, OWL-S,
interoperability between services and information is
optimized. The design offers maximum flexibility
as an organization’s preservation needs change, the
system adapts accordingly. As new preservation
services, tools, standards and recommendations
evolve, they can automatically be incorporated into
the system by adding their semantic descriptions to
the relevant registries. As well as providing unified
access to the wide range of preservation services
available, the system also provides decision-support
and recommender services to assist the scientific data
collections manager to select the best single servic e
or combination of services for a particular set of
objects . The user interface allows easy customization
of the system and human intervention where required
– offering the best combination of human and
software agents.</p>
      <p>To conclude, we believe that the semantic web
services approach, as employed within the PANIC
system, provides the optimum architecture for viable
long term preservation of large scale collections of
scientific data. By enabling the automatic detection of
potentially obsolescent data objects and the dynamic
discovery and execution of the most appropriate
preservation service – the system can potentially save
research, government and scientific organizations
vast amounts of time and effort, as well as prevent
the loss of valuable data.</p>
    </sec>
    <sec id="sec-10">
      <title>Ackno wledgements</title>
      <p>The work reported in this paper has been funded in
part by the Co -operative Research Centre for
Enterprise Distributed Systems Technology (DSTC)
through the Australian Federal Government's CRC
Programme (Department of Education, Science, and
Training).
Global Digital Format Registry (GDFR)
http://hul.harvard.edu/formatregistry/
PRONOM –The file format registry
http://www.nationalarchives.gov.uk/pronom/
VersionTracker http://www.versiontracker.com/
D. Giaretta, “Draft DCC Approach to Digital
Curation”, Jan 2005.
http://dev.dcc.rl.ac.uk/twiki/bin/view/Main/DCCApp
roachToCuration
[10] The Typed Object Model (TOM)</p>
      <p>http://tom.library.upenn.edu/
[11] J.R. van der Hoeven, “Permanent Access Technology
for the virtual heritage”, May 2004
http://jeffrey.famvdhoeven.nl/Researchtask%20IBM
%20TU%20Delft%20%20J.R.%20van%20der%20Hoeven.pdf
[13] Preservation webservice Architecture for Newmedia
&amp; Interactive Collections (PANIC),
http://www.metadata.net/panic/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>McGlynn</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Szalay</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Wicenec, “
          <article-title>VOTable: A Proposed XML Format for Astronomical Tables”</article-title>
          ,
          <source>Apr</source>
          <year>2002</year>
          , http://www.us-vo.org/VOTable/VOTable1-0.pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Lavoie</surname>
          </string-name>
          , “
          <article-title>Meeting the challenges of digital preservation: The OAIS reference model”</article-title>
          ,
          <source>Feb</source>
          <year>2002</year>
          , http://www.oclc.org/research/publications/archive/20 00/lavoie/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Buckley</surname>
          </string-name>
          , “Virtual Remote Control:
          <article-title>Building a Preservation Risk Management Toolbox for Web Resources”</article-title>
          ,
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <year>April 2004</year>
          http://www.dlib.org/dlib/april04/mcgovern/04mcgov ern.html
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>A.</given-names>
            <surname>Stanescu</surname>
          </string-name>
          , “
          <article-title>Assessing the Durability of Formats in a Digital Preservation Environment: The INFORM</article-title>
          <string-name>
            <surname>Methodology” D-Lib Magazine</surname>
          </string-name>
          November 2004 Volume
          <volume>10</volume>
          Number 11
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>http://www.dlib.org/dlib/november04/stanescu/11sta nescu.html</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Yamamoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Tatebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sekiguchi</surname>
          </string-name>
          ,
          <article-title>"Parallel and Distributed Astronomical Data Analysis on Grid Datafarm"</article-title>
          ,
          <source>Proceedings of 5th IEEE/ACM International Workshop on Grid Computing (Grid</source>
          <year>2004</year>
          ), pp.
          <fpage>461</fpage>
          -
          <lpage>466</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Metadata</given-names>
            <surname>Encoding</surname>
          </string-name>
          and
          <article-title>Transmission Standard (METS) http</article-title>
          ://www.loc.gov/standards/mets/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[16] ISO/IEC TR 21000-1</source>
          :
          <fpage>2001</fpage>
          <string-name>
            <surname>(E) (</surname>
            <given-names>MPEG</given-names>
          </string-name>
          <source>-21) Part</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Vision</surname>
            , Technologies and Strategy,
            <given-names>MPEG</given-names>
          </string-name>
          , Document: ISO/IEC JTC1/SC29/WG11 N3939 http://www.cselt.it/mpeg/public/mpeg-21_pdtr.zip
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bekaert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hochstenbach and H. Van de Sompel</surname>
          </string-name>
          , “
          <article-title>Using MPEG-21 DIDL to Represent Complex Digital Objects in the Los Alamos National Laboratory Digital Library”</article-title>
          ,
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <year>November 2003</year>
          http://www.dlib.org/dlib/november03/bekaert/11beka ert.html
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>[18] JSTOR/Harvard Object Validation Environment (JHOVE) http://hul.harvard.edu/jhove/jhove.html</mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Global</given-names>
            <surname>Digital Format</surname>
          </string-name>
          <article-title>Registry (GDFR), Data Model v.3 Dec 2003 http://hul</article-title>
          .harvard.edu/gdfr/DataModel_v3.doc
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Parsia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sirin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hendler</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Nau</surname>
          </string-name>
          , “
          <string-name>
            <surname>Automating DAML-S Web Services Composition Using</surname>
          </string-name>
          <article-title>SHOP2”</article-title>
          , 2nd International Semantic Web Conference,
          <string-name>
            <surname>ISWC</surname>
          </string-name>
          <year>2003</year>
          ,
          <string-name>
            <given-names>Sanibel</given-names>
            <surname>Island</surname>
          </string-name>
          , Florida, USA,
          <year>October 2003</year>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>http://www.mindswap.org/papers/ISWC03- SHOP2.pdf</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tsai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yan</surname>
          </string-name>
          , S. Chou, “Ontology -Mediated Integration of Intranet Web Services”, Computer No 10, Volume
          <volume>36</volume>
          ,
          <year>October 2003</year>
          http://computer.org
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [22]
          <string-name>
            <surname>The</surname>
            <given-names>OWL Services</given-names>
          </string-name>
          <string-name>
            <surname>Coalition</surname>
          </string-name>
          , “
          <article-title>OWL-S: Semantic Markup for Web services</article-title>
          ”
          <year>July 2004</year>
          http://www.daml.org/services/owl-s/1.1B/owl-s/owls.html
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Massimo</surname>
            <given-names>Paolucci</given-names>
          </string-name>
          , Katia Sycara, Takuya Nishimura, Naveen Srinivasan, “
          <article-title>Using DAML-S for P2P Discovery”</article-title>
          , International Conference on Web Services,
          <string-name>
            <surname>ISWS</surname>
          </string-name>
          <year>2003</year>
          ,
          <string-name>
            <given-names>Las</given-names>
            <surname>Vegas</surname>
          </string-name>
          , Nevada, USA, June 2003 http://www2.cs.cmu.edu/~softagents/papers/p2p_icws.pdf
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [24] OASIS, “UDDI Spec Technical Committee Draft”,
          <year>Nov 2004</year>
          http://uddi.org/pubs/uddi_v3.htm
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>