<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards A Semantic &amp; Domain-agnostic Scienti c Data Management System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuan-Fang Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gavin Kennedy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Faith Davies</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jane Hunter</string-name>
          <email>j.hunterg@uq.edu.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of ITEE The University Of Queensland Brisbane</institution>
          ,
          <addr-line>Queensland 4072</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data management has become a critical challenge faced by a wide array of scienti c disciplines in which the provision of sound data management is pivotal to the achievements and impact of research projects. Massive and rapidly expanding amounts of experimental data combined with evolving domain models contribute to making data management an increasingly challenging task that warrants a rethinking of its design. In this paper we present PODD, an ontology-centric data management system architecture for scienti c experimental data that is extensible and domain independent. In this architecture, the behaviors of domain concepts and objects are speci ed entirely by ontological entities, around which all data management tasks are carried out. The open and semantic nature of ontology languages also makes PODD amenable to greater data reuse and interoperability. To evaluate this architecture, we have developed a data management system and applied it to the challenge of managing phenomics data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Data management is the practice of managing (digital) data and resources,
encompassing a wide range of activities including acquisition, storage, retrieval,
discovery, access control, publication and archival. For many data-intensive
scienti c disciplines such as life sciences and bioinformatics, sound data management
informs and enables research and has become an indispensable component [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The need for e ective data management is, in a large part, due to the fact
that massive amounts of digital data are being generated by modern instruments.
Furthermore, the fast evolution of technologies/processes and discovery of new
scienti c knowledge require exibility in handling dynamic data and models in
data management systems. Among others, there are three core challenges for
e ective data management in scienti c research.</p>
      <p>{ The ability to provide a data management service that can manage large
quantities of heterogeneous data in multiple formats (text, image, and video)
and not be constrained to a nite set of experimental, imaging and
measurement platforms or data formats.
{ The ability to support metadata-related services to provide context and
structure for data within the data management service to facilitate e
ective search, query and dissemination.
{ The ability to accommodate evolving and emerging knowledge, technologies
and processes.</p>
      <p>
        Database systems have traditionally been used successfully to manage
research data [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] in which database schemas are used as domain models to capture
attributes and relationships of domain concepts. One implication of the above
approach is that domain models need to stay relatively stable as database
extension and migration is often an error-prone and laborious task. Consequently,
this approach is not suitable for domains where data and model evolution is the
norm rather than the exception.
      </p>
      <p>
        Ontology language OWL possesses expressive, rigorously-de ned semantics
and non-ambiguous syntaxes. it has been designed to be open and extensible
and to support knowledge and data exchange on the Web [
        <xref ref-type="bibr" rid="ref1 ref10 ref8">1, 8, 10</xref>
        ]. These
intrinsic characteristics make them an ideal conceptual platform on which a exible
scienti c data management system can be built.
      </p>
      <p>In this paper, we present our work in designing PODD (Phenomics Ontology
Driven Data Management), a semantic, domain-agnostic architecture for systems
managing data generated from scienti c experiments that employes ontologies as
domain models. The ontology-based domain model is at the core of PODD as it
de nes the behavior of entities in scienti c experiments. Logical structure of data
is therefore maintained and enforced via ontological de nitions and reasoning,
and not via database schemas and associated constraints.</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] we described an early version of the PODD data repository to meet
the above challenges facing the Australian phenomics research community. We
would like to emphasize that although the PODD system presented in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is
geared towards phenomics research, the ontology-centric architecture we propose
in this paper is actually domain-independent and can be applied in any scienti c
discipline where research activities and output can be conceptually organized in
a structured manner.
      </p>
      <p>The rest of the paper is organized as follows. In Section 2 we present related
work and give a brief overview of the motivation and goals of the PODD project.
Section 3 presents the ontology-based architecture for data management systems.
In Section 4, we discuss the PODD ontologies in more detail and show how the
ontology-based modeling approach is used in the life cycle of repository concepts
and objects. In Section 5, we describe the PODD data management system we
developed based on the ontology-driven architecture. Finally, Section 6 concludes
the paper and identi es future directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Overview</title>
      <p>In this section, we survey a number of related systems and architectures.
Following the survey, we present the motivation behind the ontology-centric
architecture and the goals we wish to achieve with the PODD data management
system.
2.1</p>
      <p>Related Work
A number of ontology repositories and search engines have been developed.
Repositories such as NCBO Bioportal1 and Cupboard2 publish ontologies and
usually support functionalities including full-text &amp; faceted search, hierarchical
browsing, visualization and cross references. Ontology search engines such as
Swoogle3 and Watson4 index and store large numbers of ontologies and make
them searchable.</p>
      <p>
        There are also prior works in developing content repository systems.
Fedora Commons5 is a widely used open-source, general-purpose digital resource
management system based on the principles of modularity, interoperability and
extensibility. In Fedora Commons, abstract concepts are de ned as models, on
which inter-relationships and behaviors can be further de ned. Data in Fedora
Commons repositories are organized into objects, which have datastreams that
stores either metadata or data. PhenomicDB [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a multi-organism
phenotypegenotype database for a number of model organisms. It contains data from a
number of primary databases including FlyBase, Phenobank, OMIM and NCBI
Gene. More recently, an ontology-based approach has been taken in VIVO [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
to model, organize and integrate research activities and researcher pro le in an
institutional setting.
      </p>
      <p>The Ontology for Biomedical Investigations (OBI)6 is an ongoing e ort aimed
at developing an integrative ontology for biological and clinical investigations. It
takes a top-down approach by reusing high-level, abstract concepts from other
ontologies. It includes 2,600+ OWL classes and 10,000+ logical axioms (in the
import closure of the OBI ontology). OBI is very comprehensive and is suitable as
an annotation vocabulary for structured data. However, its size and complexity
(SHOIN (D)) makes reasoning and querying of OBI-based ontologies and RDF
graphs computationally expensive and time consuming7, making it impractical
as a domain model for a data management system where such reasoning may
need to be performed repeatedly.</p>
      <p>
        Functional Genomics Experiment Model (FuGe) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is an extensible
modeling framework for high-throughput functional genomics experiments, aiming at
increasing the consistency and e ciency of experimental data modeling for the
1 http://bioportal.bioontology.org/
2 http://kmi-web06.open.ac.uk:8081/cupboard
3 http://swoogle.umbc.edu/
4 http://watson.kmi.open.ac.uk/
5 http://www.fedora-commons.org/
6 http://purl.obolibrary.org/obo/obi
7 On a MacBook Pro with 2GB memory and an Intel Core 2 Duo 2.4 GHz processor,
classifying the OBI ontology (version \2009-11-06") takes more than 6 minutes using
Pellet in Protege. Such performance is clearly inadequate for a data management
system.
molecular biology research community. Centered around the concept of
experiments, it encompasses domain concepts such as protocols, samples and data.
FuGe is developed using UML from which XML Schemas and database de
nitions are derived. The FuGe model covers not only biology-speci c information
such as molecules, data and investigation; it also de nes commonly used
concepts such as audit, reference and measurement. Extensions in FuGe are de ned
using inheritance of UML classes.
      </p>
      <p>We feel that the extensibility we require is not met by FuGe as any
addition of new concepts would require amendment of database schemas and code.
Moreover, the concrete objects reside in relational databases, making subsequent
integration and dissemination more di cult.
2.2</p>
      <p>Motivation &amp; Goals
Phenomics is a fast-growing, data-intensive discipline with new technologies and
processes rapidly emerging and evolving. As a result, its domain model and
data management systems must also be able to evolve to handle the complexity,
dynamics and scale of the data.</p>
      <p>In phenomics, data is usually captured and measured by both high- and
low-throughput phenotyping devices. The scale of measurement can be from the
micro or cellular level, through the level of a single organism, and up to the
macro or eld level. Imaging, measurement and analysis of organisms on such a
large scale will produce an enormous amount of data.</p>
      <p>Phenomics research makes use of a large variety of imaging and measurement
platforms. For example, in mouse histopathology and organ pathology research,
the Zeiss \Mirax Scan" scanner is used to scan microscope slides. In clinical
pathology, a Flow Cytometer is used to capture laser di raction images of blood
samples. In plant research, the Lemnatec Scanalyzer is used to capture RGB
images of plants in growth cabinets. The Fluorogroscan system is used in quenching
analysis: the partitioning of light energy used in photosynthesis on model plants
such as Arabidopsis. Other devices, such as the Infrared Thermography Camera
are used to capture leaf temperature and the SPAD Meter is used to measure
the chlorophyll content of plant leaves. New devices and instruments will also
be employed as they become available. Moreover, existing instruments may be
upgraded so that they can capture more information. The PODD domain model
needs to be exible to accommodate these continual changes in the formats,
resolution and source of the data.</p>
      <p>Because an organism's phenotype is often the product of the organism's
genetic makeup, its development stage, disease conditions and its environment,
any measurement made against an organism needs to be recorded in the context
of these other metadata. Consequently the opportunity exists to create a
repository to record the data, the contextual data (metadata) and data classi ers in
the form of ontological or structured vocabulary terms. The structured nature
of this repository will support both manual and autonomous data discovery as
well as provide the infrastructure for data based collaborations with domestic
and international research institutions. Currently there are no such integrated
systems available. The goals of PODD are to capture, manage, annotate and
distribute the data generated by mouse and plant phenomics research activities.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Architecture of the Ontology-Centric Data</title>
    </sec>
    <sec id="sec-4">
      <title>Management System</title>
      <p>The most distinguishing characteristic of PODD is the central role that
ontologies play. In this architecture, raw data is not stored in a at structure but
is attached to domain objects organized in a logical, hierarchical system,
dened according to the domain model that represents the structure of research
activities.</p>
      <p>Current content management systems typically have a relatively static
domain model and hardwire it as relational schemas and foreign key constraints in
a custom relational database independent from the underlying repository
system. Consequently, the information pertinent to each concrete object is stored
in this custom database as well. As stated in the previous section, this approach
is unsuitable for dynamic environments where conceptual changes are common.</p>
      <p>To e ectively support a dynamic conceptual framework, the domain model
in the proposed architecture is de ned using OWL ontologies, in which: OWL
classes represent domain concepts; OWL properties de ne concept attributes
and their relationships; OWL restrictions specify constraints on concepts and
nally; OWL individuals de ne concrete domain objects where attributes and
relationships are de ned using OWL assertions. Raw data les are attached to
concrete domain objects.</p>
      <p>Such a conceptual architecture alleviates the problem of imposing hard
relational constraints in a database which is di cult to extend/change.</p>
      <p>Another drawback of existing systems is that there can be only one domain
model. When a concept needs to be updated, all the existing objects de ned by
that concept need to be updated accordingly, which may be undesirable,
inappropriate and time-consuming. This is, unfortunately, unavoidable as long as the
domain model is de ned using database schemas. In our proposed architecture,
as concept and object de nitions are stored in the repository, such changes can
be versioned so that existing instance objects can remain legitimate when
integrity validation is performed as they can still refer to the previous conceptual
de nitions.</p>
      <p>The high-level design of ontology-centric architecture takes a modular and
layered approach, as can be seen in Figure 1. At the foundation is the data
access layer, consisting of an underlying repository system, an RDF triple store,
an in-house database that stores essential information and a full-text search
engine. This layer is responsible for low-level tasks when the creation, modi cation
and deletion of concepts and objects occur. The business logic layer in the
middle is responsible for managing concepts and objects, such as versioning,
object conversion and integrity validation. The security layer controls access
(authentication and authorization) to concepts and objects and guards all
operations on them. At the top of the stack is the interface layer, where the data</p>
      <p>Object
Management
Repository</p>
      <p>Interface Layer
Metadata Publishing
Services Services</p>
      <p>Security Layer
Business Logic Layer</p>
      <p>Concept
management</p>
      <p>Data Access Layer
RDF Triple Database</p>
      <p>Store
users, roles
Reasoning
Service</p>
      <p>Search
Index
management system can be accessed using a number of interfaces such as a Web
browser or API calls.</p>
      <p>In developing the ontology-centric architecture, the following design decisions
have been made to balance expressivity, exibility and conceptual clarity. These
decisions have also been based on a survey of user requirements from scientists
within a range of research organizations including the Australian Plant
Phenomics Facility (APPF) as well as the Institute of Molecular Biology (IMB),
Queensland Brain Institute (QBI) and Australian Institute of Bioengineering &amp;
Nanotechnology (AIBN), working on collaborative research projects that involve
large scale data and distributed teams:
{ There is a top-level domain concept, called Project , under which other
concepts (such as Investigation and Material) reside in a hierarchical manner.
{ Access control (authorization) is de ned on the Project level but not on an
individual object level, i.e., a given user will have the same access rights for
all objects within a given project.
{ Within a Project hierarchy, objects are in a parent-child relationship in a
tree structure such that each child can only have one parent. This ensures
that access rights are properly propagated from parent to child and there is
no chance of confusion.
{ Additionally, inter-object, many-to-many reference relationships can be
dened to enhance exibility of the architecture as it allows arbitrary links
between objects to be established.
{ Objects cannot be shared across Projects. Instead, objects must be copied
from one project and pasted into another one. Such a rule simpli es object
management with the elimination of possible side-e ects caused by sharing
object between projects.
{ There should be no interference between di erent versions of a given concept
and between objects that are instances of di erent concept versions.</p>
    </sec>
    <sec id="sec-5">
      <title>Ontology-based Domain Modeling</title>
      <p>As we emphasized previously, the domain model should be exible enough to
accommodate the rapid changes and dynamic nature of scienti c research. In
this section, we present the base ontology and the roles it plays in the
ontologycentric architecture. It should be noted that the architecture proposed here is
domain-independent and it can be applied to any scienti c discipline that shares
a similar high-level domain model.</p>
      <p>Note that concepts in the ontologies presented here are models of entities
(activities and objects) in scienti c investigations: they de ne the logical
structure of investigations - objects in an investigation and logic relationship between
these objects. In a sense, the ontologies serve as a data model of the investigations
in the scienti c domain for which the data management system is developed. In
other words, the ontologies are used as a model for the data management system
implementation.
4.1</p>
      <p>
        The Base Domain Ontology
Inspired by FuGe [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and OBI8, we created the base domain ontology in OWL to
de ne essential domain concepts, their attributes and inter-relationships in an
object-oriented fashion. As stated in the previous section, domain concepts will
be modeled as OWL classes; relationships between concepts and object attributes
will be modeled as OWL object and datatype properties. Concrete objects will
be modeled as OWL individuals.
      </p>
      <p>Investigation</p>
      <p>Event</p>
      <p>Process</p>
      <p>Protocol</p>
      <p>For an overview, inter-relationships of some of the domain concepts in this
ontology are shown in Figure 2. It is worthing pointing out that concepts in
this gure are shown in the logical hierarchy but not the inheritance hierarchy:
it describes the structure of a scienti c investigation and how di erent
activities/objects in it are related to each other. For brevity reasons, OWL object
properties and cross references between classes are not shown. We also de ned
the following design principles for the domain ontology.
8 http://purl.obolibrary.org/obo/obi
{ All essential domain concepts are modeled as sub-classes of an abstract
toplevel OWL class PODDConcept that captures common attributes and
relationships.
{ All relationships between domain concepts are captured by domain
properties, which can be further divided into two property hierarchies, one for
parent-child relationships and the other for reference relation-ships. Each of
the two hierarchies have an abstract top-level property, called contains and
refersTo, respectively.
{ All parent-child relationships are modeled in a property hierarchy as
subproperties of the abstract property contains, and all reference relationships
are modeled in another property hierarchy as sub-properties of the abstract
property refersTo.
{ For each domain concept C, one property is de ned in each of the above
hierarchies with its range de ned to be C. The domains of such properties
are not speci ed so that they can be used by any applicable domain concept
to establish a relationship between them.
{ Class attributes are modeled using OWL restrictions.
{ Essential domain concepts can be sub-classed to provide more specialized
and re ned information.
{ To ensure that each object can have at most one parent object, the inverse
property of contains, isContainedBy, is de ned so that a max cardinality
restriction can be added to the top-level concept PODDConcept to enforce
it.</p>
      <p>
        The de nitions of some top-level constructs are summarized in Figure 3, in
OWL DL syntax [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <sec id="sec-5-1">
        <title>PODDConcept v &gt;</title>
        <p>isContainedBy v ( contains)
&gt; v refersTo:PODDConcept
&gt; v 8 contains:PODDConcept</p>
      </sec>
      <sec id="sec-5-2">
        <title>PODDConcept v</title>
        <sec id="sec-5-2-1">
          <title>1 isContainedBy</title>
          <p>In our model, we use OWL properties to model object attributes. When the
possible values of a particular attribute can be enumerated, such as project status
(active, inactive and completed ), an enumerated OWL class is used to represent
all the values. When an attribute represents a grouping of some values, such
as accessions, where an accession has a source and a number, an OWL class is
also de ned to represent the grouping. In this case, auxiliary OWL properties
are de ned to project out speci c values in the grouping. In all other cases,
attributes are modeled using datatypes.</p>
          <p>Figure 4 shows the partial de nition of the OWL class Project .
Restriction (1), for example, states that any Project instance must have exactly one
ProjectPlan (through the predicate hasProjectPlan, the range of which is ProjectPlan).
The other 3 restrictions are similarly de ned.</p>
          <p>Project v = 1 hasProjectPlan u
v
v = 1 hasStartDate u</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>1 hasInvestigation u</title>
        <p>v</p>
        <sec id="sec-5-3-1">
          <title>1 hasPublicationDate</title>
          <p>The base ontology de nes essential concepts independent of the domain.
Domainspeci c knowledge can be incorporated by extending the base ontology for
disciplinespeci c systems.</p>
          <p>As stated in Section 1, the ontology-based domain model is at the center of
the whole life cycle of objects. In this subsection, we brie y describe the roles
that the domain ontologies perform at various stages of the object life cycle.
Ingestion When an object is created, its de nition is expressed in ontological
terms. Such de nitions will be used to (a) guide the rendering of object
creation interfaces and (b) validate the attributes and inter-object
relationships the user has entered before the object is ingested. When an object is
ingested, its de nitions are stored as RDF assertions.</p>
          <p>Retrieval &amp; update When an object is retrieved from the repository, its
attributes and inter-object relations are retrieved from its RDF assertions,
which are used to drive the on-screen rendering. When any value is updated,
it is validated and updated in this object's RDF assertions.</p>
          <p>Query &amp; search An object's assertions will be stored in an RDF triple store,
which can be queried using SPARQL. Similarly, ontological de nitions are
indexed to provide functionalities such as full-text search and faceted
browsing.</p>
          <p>Publication &amp; export When an object is published or exported, its metadata,
in RDF, will be retrieved and exported.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>The PODD Data Management System</title>
      <p>Based on the ontology-centric architecture presented in Section 3 and the base
ontology presented in Section 4 we implemented the PODD data management
system - with the aim being to meet the data management challenges faced by
the Australian phenomics research community.</p>
      <p>To describe domain knowledge in phenomics, we extend the base ontology by
de ning additional concepts including Genotype, Gene, Phenotype and Sequence
as subclasses of PODDConcept. Additional OWL object and datatype properties
are also de ned to model the attributes and relationships of these concepts, as
shown in Figure 5. Note that Phenotype is a subclass of Observation.</p>
      <p>Data</p>
      <p>Data
Event</p>
      <p>Gene</p>
      <p>Allele
Marker
Sequence
Observation/
Phenotype</p>
      <p>Figure 6 shows some new de nitions in the domain ontology. Note that the
last two de nitions integrate the new de nitions with those in the base ontology.</p>
      <p>Also note that the concepts de ned in the PODD ontologies do not necessarily
represent real-world entities/reality. For example, the OWL class Gene does not
intend to be a class that describes genes in general. Rather, it is used to describe
genes that are observed/involved in scienti c investigations.</p>
      <sec id="sec-6-1">
        <title>Genotype v PODDConcept</title>
      </sec>
      <sec id="sec-6-2">
        <title>Gene v PODDConcept</title>
        <p>v
v 8 hasGene:Gene</p>
        <sec id="sec-6-2-1">
          <title>1 hasEcotype</title>
          <p>v</p>
        </sec>
        <sec id="sec-6-2-2">
          <title>1 hasSubspecies</title>
        </sec>
      </sec>
      <sec id="sec-6-3">
        <title>Project v 8 hasGenotype:Genotype Material v 8 hasPhenotype:Phenotype v 8 refersToGenotype:Genotype</title>
        <p>v
v
v 8 hasSequence:Sequence</p>
        <sec id="sec-6-3-1">
          <title>1 hasAlias</title>
        </sec>
        <sec id="sec-6-3-2">
          <title>1 hasChromosome</title>
          <p>In developing the PODD system, we chose to employ a number of mature
technologies. (1) We use Fedora Commons for the storage and retrieval of
domain objects. Together with raw data les, the OWL (for concepts) and RDF
(for objects) de nitions of each concept and object are stored in a versioned
datastream PODD, which is used by the PODD system in various tasks such
as object creation, rendering, validation, update and visualization. (2) We
incorporate the Sesame triple store9 to support complex query answering using
SPARQL. Sesame contexts are used to give scope to the RDF triples for each
domain object. As described in Section 3, access control needs to be enforced on
a per project level. Similarly, it also needs to be enforced on query answering
in the triple store. By identifying triples of individual objects, we are able to
9 http://www.openrdf.org/
control contexts a user can access through query expansion. (3) Lastly, we use
the Lucene and Solr open-source search engine platform10 to provide full-text
search and faceted browsing capabilities. Similar to the structure of the Sesame
triple store, there is a one-to-one correspondence between domain objects in the
repository and the Solr documents, the logical indexing units.</p>
          <p>Although the architecture and the system are based on ontologies, the
interface is designed to hide ontology-related complexity from the user and present
information in an easy to use manner for all repository functions. For example,
Figure 7 shows the browser view of a plant phenomics project that investigates
salt tolerance of wheat. In this view, the objects are shown in a tree-like
structure by following property assertions of subproperties of contains de ned in the
base and domain ontologies.</p>
          <p>We have started to deploy the PODD system in Australian phenomics
research centers including APPF and APN and begun engaging users in the
evaluation of the performance, exibility, usability and scalability of the system. User
feedback to date has shown that the system is intuitive and e cient.
6</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>In summary, our contribution to scienti c data management is three-fold: rstly,
the proposal of the ontology-centric architecture for developing data
management systems; secondly, the development of a base ontology that de nes essential
10 http://lucene.apache.org/
domain knowledge; and thirdly, the development of the PODD data management
system (based on both existing and new technologies) that validates the
feasibility of the proposed approach.</p>
      <p>We have identi ed a number of future work directions that we would like to
pursue. Firstly, we will investigate integration with existing domain ontologies
such as the Gene Ontology and the Plant Ontology. One possibility would be
to use terms de ned in these ontologies to annotate metadata objects. Secondly,
we would like to investigate the generalization of the ontology-centric approach
so that it can be applied to other areas such as work ow management systems.
Thirdly, we will continue the development of the PODD system to provide
additional functionalities such as data visualization, automated data integration and
Linked Data-style data discovery and publication.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          , G. Kobilarov,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cyganiak</surname>
          </string-name>
          , and
          <string-name>
            <surname>Z. Ives.</surname>
          </string-name>
          <article-title>DBpedia: A Nucleus for a Web of Open Data</article-title>
          .
          <source>In Proceedings of 6th International Semantic Web Conference, 2nd Asian Semantic Web Conference (ISWC+ASWC</source>
          <year>2007</year>
          ), pages
          <fpage>722</fpage>
          {
          <fpage>735</fpage>
          ,
          <string-name>
            <surname>November</surname>
          </string-name>
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J.</given-names>
            <surname>Gray</surname>
          </string-name>
          , D. T. Liu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nieto-Santisteban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Szalay</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. J. DeWitt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Heber</surname>
          </string-name>
          .
          <article-title>Scienti c data management in the coming decade</article-title>
          .
          <source>SIGMOD Rec</source>
          .,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <volume>34</volume>
          {
          <fpage>41</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Groth</surname>
          </string-name>
          , Philip, Pavlova, Nadia, Kalev, Ivan, Tonov, Spas, Georgiev, Georgi, Pohlenz, Hans-Dieter,
          <article-title>Weiss, and</article-title>
          <string-name>
            <surname>Bertram. Phenomicdb:</surname>
          </string-name>
          <article-title>a new cross-species genotype/phenotype resource</article-title>
          .
          <source>Nucleic Acids Research</source>
          ,
          <volume>35</volume>
          (Supplement 1):D696{
          <fpage>D699</fpage>
          ,
          <year>January 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>I.</given-names>
            <surname>Horrocks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Patel-Schneider</surname>
          </string-name>
          , and
          <string-name>
            <surname>F. van Harmelen. From SHIQ</surname>
          </string-name>
          and
          <article-title>RDF to OWL: The Making of a Web Ontology Language</article-title>
          .
          <source>Journal of Web Semantics</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):7{
          <fpage>26</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aebersold</surname>
          </string-name>
          , et al.
          <article-title>The Functional Genomics Experiment model (FuGE): an Extensible Framework for Standards in Functional Genomics</article-title>
          .
          <source>Nature Biotechnology</source>
          ,
          <volume>25</volume>
          (
          <issue>10</issue>
          ):
          <volume>1127</volume>
          {
          <fpage>1133</fpage>
          ,
          <year>October 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. D. B. Kra t, N. A.
          <string-name>
            <surname>Cappadona</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Caruso</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Corson-Rikert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Devare</surname>
            ,
            <given-names>B. J.</given-names>
          </string-name>
          <string-name>
            <surname>Lowe</surname>
          </string-name>
          , and VIVO Collaboration.
          <article-title>VIVO: Enabling National Networking of Scientists</article-title>
          .
          <source>In Proceedings of the WebSci10: Extending the Frontiers of Society On-Line, Apr</source>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Y.-F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Kennedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Davies</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Hunter. PODD:</surname>
          </string-name>
          <article-title>An Ontology-driven Data Repository for Collaborative Phenomics Research</article-title>
          .
          <source>In Proceedings of 12th International Conference on Asian Digital Libraries (ICADL</source>
          <year>2010</year>
          ), pages
          <fpage>179</fpage>
          {
          <fpage>188</fpage>
          . Springer-Verlag,
          <year>June 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Ruttenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rees</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samwald</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Marshall</surname>
          </string-name>
          .
          <source>Life Sciences on the Semantic Web: the Neurocommons and Beyond. Brie ngs in Bioinformatics</source>
          ,
          <volume>10</volume>
          (
          <issue>2</issue>
          ):
          <volume>193</volume>
          {
          <fpage>204</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Klicker</surname>
          </string-name>
          , E. Stephan, H. Wiley, and
          <string-name>
            <given-names>K.</given-names>
            <surname>Waters</surname>
          </string-name>
          .
          <article-title>Enabling high-throughput data management for systems biology: The bioinformatics resource manager</article-title>
          .
          <source>Bioinformatics</source>
          ,
          <volume>23</volume>
          (
          <issue>7</issue>
          ):
          <volume>906</volume>
          {
          <fpage>909</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>B.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ashburner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rosse</surname>
          </string-name>
          , et al.
          <article-title>The OBO Foundry: Coordinated Evolution of Ontologies to Support Biomedical Data Integration</article-title>
          .
          <source>Nature Biotechnology</source>
          ,
          <volume>25</volume>
          (
          <issue>11</issue>
          ):
          <volume>1251</volume>
          {
          <fpage>1255</fpage>
          ,
          <string-name>
            <surname>November</surname>
          </string-name>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>