<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DDaattaabbaassee EEnnggiinneeeerriinngg ffrroomm tthhee CCaatteeggoorryy TThheeoorryy VViieewwppooiinntt</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Toth Toth</string-name>
          <email>tothd1@fel.cvut.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>David Toth David Toth Dept. of Computer Science, FEE CTU Prague, Department oKfaCrloomvopunt ́aemr .S1c3ie,n1c2e1a3n5d Engineering Faculty of ElectricaPl rEanhgai,nCeezreicnhg,RCepzeucbhliTcechnical University Karlovo</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper gives an overview of XML formal models, summarizes database engineering practices, problems and their evolution. We focus on categorical aspects of XML formal models. Many formal models such as XML Data Model, XQuery Data Model or Algebra for XML can be described in terms of category theory. This kind of description allows to consider generic properties of these formalisms, e.g. expressive power, optimization, reduction or translation between them, among others. These properties are rather crucial to comparison of different XML formal models and to consequent decision which formal system should be used to solve a concrete problem. This work aim is to be the basis for further research in the area of XML formal models where category theory is applied.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In this paper we will focus on some peculiarities from today’s database world.
Now, in spring 2008, we have many database technologies, many technical
frameworks, many solutions for different and similar problems. What we do not have
is a global point of view of databases (DB); theoretical approach stating
theorems about database models and languages. This paper summarizes database
technologies from higher perspective and introduces some of the terms from
mathematical category theory (CT). These two aspects, databases and category
theory, are put together in order to give new look at the database technologies, to
give new way of data model and languages description; and to find new language
in which we could ask and answer more generic questions, e.g. about expressive
power of (query) languages of particular data models.</p>
      <p>
        This paper deals with databases. More precisely we should say it treats
problems which appear when we would like to know which database technology should
be used in software project. There are generaly more requirements leading one to
use DB, e.g. to make the data persistent, to assure concurrency, etc. More about
database technology in general can be found in Date’s Introduction to Database
Systems [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. There are many factors influencing the decision. In fact in praxis
it is more subjective (personal or team) decision.
      </p>
      <p>From the software engineering point of view there are at least these kinds
of factors; (1) Human factors, as e.g.: knowledge about particular DB product,
concrete DB technology experience; individual or team, subjectivelly favourite
/ preferred DB technology; (2) Technical aspects: vendor influence, e.g. offered
support, problem solving time, programming languages support, performance,
accessibility, clustering, and many others. (3) Problem definition: teoretical
aspects of problem, data itself, its nature (data character), i.e. (a) structuralization
(no inner structure e.g. streams, files respectively, weak structure e.g.
newspaper articles, strong structure any well structured forms, e.g. tax return form),
(b) data contain metadata (typical for XML documents), data separated from
metadata respectively (typical for tables—relations). (4) Possibly other aspects.</p>
      <p>
        Some of these factors are summarized in SWEBOK [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ]. In software
engineering paper we would like to address especially the first two categories. In
SIGSOFT [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] and especially in SEN [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] can be found more on these topics.
But this text is intended to be considered as more database-oriented. Therefore
we will focus more on the problem’s aspects as the third factor mentioned above.
Nevertheless all topics covered here are closely related to software engineering
and even to database engineering which we deal with later. Next we will take a
closer look at particular database technologies emphasizing the problem’s aspect.
1.1
      </p>
      <sec id="sec-1-1">
        <title>Relational Database Technology</title>
        <p>
          Historically the first database approach which solved the inconsistencies,
redundancy, concurrency and other problems was the relational model. C. J. Date in
his Introduction [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] deals with the relational approach to represent data (i.e.
relational data modeling and storing among other aspects). Other very deep
insight into relational data model can be found in E. F. Codd’s Relational model
for database management [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. One can imagine the main idea as data grasped
via a relation in mathematical notion, i.e. all data can be viewed as relations, in
other words sets with internal structure of its elements. Relations are
interconnected together using values of particular set of elements which is usually called
foregin key usage.
        </p>
        <p>Can any data be represented using relational approach, i.e. can any data
be stored as relations, tables respectively? We must consider the fact that the
data could possibly change its structure, and even that we do not know the
structure before we have the data physically. Can the changing data structure
be modeled using relational approach? What other questions play a significant
role when we consider expressive power of e.g. relational algebra, etc.? These
and related questions will be considered in future works which will contain CT
oriented features. In this paper data model description and related topics will
be treated. Further we focus on object technologies.
1.2</p>
      </sec>
      <sec id="sec-1-2">
        <title>Object Database Technology</title>
        <p>
          Object and object-oriented databases arose out of the impedance mismatch
between relational and object data models. The essence of this problem lies in a
different kind of data representation, i.e. once as relations or n-aries and once
as objects. The problem inhere in data translation. In other words data must
be mapped between classes of objects and relations of n-aries. More about the
object relational mapping can be found in Fussel’s Foundations of Object
Relational Mapping (ORM) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Another paper about ORM can be found in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ];
implementation issues are covered in persistence framework Hibernate [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
        </p>
        <p>
          To object and object-oriented databases and to object database
management systems is dedicated the web site [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. At this place many object and
object-oriented technologies, especially database technologies, of course, and
open source as well, can be found.
The XML Databases were born actually very shortly after the XML, the new
language for semistructred data description, a W3C’s standard respectively [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ],
emerged in 1998. More about XML evolution can be found at [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ]. It did not
took a long time and new term Native XML Database, often just NXD, came
abroad [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ]. We will use the term, as e.g. R. P. Bourret in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] does. Why NXD
appeared and what to expect from them is described in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] or [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. R. P. Bourret
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] also maintain fresh list of NXD products and even more wide XML products
in general.
        </p>
        <p>
          The main motivation for NXD usage resides in the impedance problem again,
as in case of ODB as well. The typical situation where NXD are used is web
portals and web applications communicating through web services. The web
services standards are based on XML and related standards. Therefore it is
evident that the need for XML document transformation should be avoided to
speed up the performance of applications of this type. We have treated this yet
earlier in [
          <xref ref-type="bibr" rid="ref42">42</xref>
          ].
        </p>
        <p>
          The principle of NXD consists in XML data model as intrinsic data model
of the database engine. R.P. Bourret is more specific about what XML-Enabled
and what XML-native suppose to mean, e.g. in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>The paper is structured as follows. Section 2 deals with relationship of software
engineering and database engineering. Section 2.2 refers problems relevant to
appropriate database technology selection in the introduction above bearing in
mind. Section 3 summarizes issues related to essences of particular data
models. In section 4 there are mostly XML related standards and technical part of
XML databases dealt with and section 5 treats formalisms developped for the
purpose of XML Data model description. Section 6 introduce basic terms from
category theory (CT) and gives formal background for cited formalisms. Section
7 summarizes exhibited XML formalism. Last section 8 reveals our future plans.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Software Engineering and Database Engineering</title>
      <sec id="sec-2-1">
        <title>Terminological note on database engineering. Normally the term database</title>
        <p>engineering is used to describe an area of processes, methods and techniques,
formalisms and languages useful for database designing, and in general,
useful for database application development. There are several conferences around
database engineering topics. Clear definition of what is database engineering does
not exist. What we mean under database engineering is specialization of
software engineering practices for purposes of database application, i.e. application
strongly related to data which makes persistent and which further operates with.
Practically we mean specialization of all techniques where arbitrary database
artifact, most typically it is database schema, is created, changed (most common
case), or removed.</p>
      </sec>
      <sec id="sec-2-2">
        <title>From waterfall to iterative development. Software engineering in past was</title>
        <p>
          understood as sequential processes equivalent to phases which must be performed
in a serial way. The phases typically are: feasibility study, business analysis,
requirements analysis, architecture analysis and design, logical design, GUI design,
DB design, physical design, coding, testing, refactoring, installing, deploying,
measuring, among others. The same can be said, as an analogy of course, about
database engineering, i.e. database design, database tuning and administration
among others. But in today’s world when agile methodologies in sofware
engineering become successfull and more and more widespread, it also seems to be
inevitable to use agile or generally speaking iterative approaches in database
community. It is a paintful step for every single database expert long time
experienced sequential approach when starting use agile principles.
MDA—Model Driven Architecture. The MDA approach [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] seems to be
contradictory in the context of agile methodologies. But not necessarilly.
Iterative approaches allow to build software systems more focused on one particular
problem, emphasizing one aim in time (during iteration). The basic imagination
could be as very little waterfalls chaining every iteration stressing analysis or
design or programming depending on current phase. We can figure out here the
semantics of the word phase depends strictly on chosen methodology.
2.1
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Database Engineering and Evolutionary Approach</title>
        <p>
          From the point of view of the database engineering there is need to elaborate
database design. Typically conceptual model is considered as a part of database
modeling and as a part of a database design phase. In fact we would like to stress
here that there is no need to create domain model as UML [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ] class diagram
during business analysis and also E-R diagram as a part of database modeling
independently. It is possible to create or even generate E-R diagram or UML class
diagram in Data Modeling Profile from domain model. It is typically expressed
as UML class diagram which is done during business analysis. Actually some
CASE tools offer this functionality in these days, e.g. the Enterprise Architect
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>Database modeling, a part of database design, can be viewed as a
transformation from domain model. And this does not mean that all the modeling must
be finished before normalization or tuning starts. The core of the evolutionary
approach lies in doing the whole step by step in very small parts which have to
be integrated. Continuous refinement is necessary. One of the biggest argument
against iterative database development is the need for neverending reworking
and refining of non-stabilized artifacts—which is possibly a great number.</p>
        <p>As a resume here we would like to pinpoint the possibility to look at database
evolution concurrently with regular software evolution. And therefore to see
database engineering as a specialization of software engineering. The
principles of MDA—model transformations are essentially the same in software and
database engineering. This kind of abstraction should help us thinking in
software engineering and database engineering in very similar way. Furthermore
CT can help us when dealing with models, their properties and qualities, and
transformations.
Which particular DB technology should we choose to use? What should lead us
— help us? The discussion below involves these questions.</p>
        <p>Relational Databases (RDB). From the historical perspective there is a
big argument which says to use RDBs. It is deep insight into relational
technology, strong mathematical background in form of data relational model and
relational algebra. Many people made refinements of this technology for a long
time. Shortly, RDBs are greatly elaborated in comparison to other (and younger)
technologies.</p>
        <p>
          Object Databases (ODB). In [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ] we could find at least these important
reasons why to select ODB instead of RDB or XDB: embedded DBMS
application, complex data relationships, deep object structures, changing data
structures, development team is using agile techniques, massive use of object oriented
programming language, there are many objects including collections, data is
accessed by navigation rather than query.
        </p>
        <p>
          One of the most popular ODBMS in open source community is db4objects
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Another example could be the NeoDatis ODB [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] or GemStone/S [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
XML Databases (XDB). With XDB, and NXD respectively, fine-grained
reuse of content is possible; NXD allows sophisticated hypertext applications
with mixture of stuctural and fulltext query. The most typically cited NXD
benefits are flexibility and reuse.
        </p>
        <p>
          We have treated of this issue in greater detail in [
          <xref ref-type="bibr" rid="ref41">41</xref>
          ]. Three distinct metrics,
ρ, τ , and ξ, were proposed for different kinds of database technologies.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Essentials of Data Models</title>
      <p>Does exist essential difference between different data models? In words of CT
we could say: belong all categories of all data models into the same category
(of categories)? We will focus a bit more on this in section 6 — The Category
Theory Standpoint.</p>
      <p>Now imagine not to use CT. The question if there exists any problem which
cannot be solved using arbitrary technology would have to be proven hardly. We
would have to prove that every single case of data expressed in one data model
could also be expressed in every other data model.</p>
      <p>Theoretically any data can be expressed in arbitrary format, i.e. (1) tables,
nested tables respectively, (2) the web of objects or (3) hierarchy of elements if
we found mappings between all data instances.</p>
      <p>Mapping from XDB to RDB can be viewed so that any XML document can
be stored (represented) in RDB in generic tables (ELEMENTS, ATTRIBUTES,
DOCUMENTS, etc.). That objects can be stored as record in tables which can
be seen e.g. in Object Relational Mapping (ORM) Pattern. The other way can be
imaginated as direct overwriting of RDB data using wrapping method for column
content and nesting in case of foreign keys (FK). FK can also be represented as
ID and IDREF attributes in XML documents.</p>
      <p>
        Mapping from XML documents to objects can be grasped in a way that XML
data model will be grasped as a tree, object model would be accessed as a graph.
A tree is also a kind of a graph. This idea is demonstrated e.g. in previous work
[
        <xref ref-type="bibr" rid="ref42">42</xref>
        ], and it is implemented in java programming language in JAXB—Java API
for XML Binding [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ]. These mappings are typically based on DTDs or XML
Schema or even RelaxNG. R. P. Bourret wrote general paper on XML document
mapping between relational and object models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The same could be obtained if we find all the mappings between one and
another DB structure — technology (RDB, ODB, XDB). But a more convenient
way would be to find out the way of general description and prove that these
mappings have to exist or that it is impossible these mappings would exist. And
not only convenient, we should consider all data models; even those which do not
exist yet. It seems that different technologies fit for different kind of problems
but they are essentially the same after all. Are they? Can we prove this using
category theory? We would like to focus our future research on these questions.</p>
      <p>And there are other interesting questions leading us to finding one framework
only, CT, e.g. is it possible to store and effectively retrieve data with unknown
and/or changing data structure in RDB, ODB and XDB?
4</p>
    </sec>
    <sec id="sec-4">
      <title>XML Databases</title>
      <p>
        XML standard [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] first released in 1998 and last updated in 2006 has initiated
the great interest in XML Databases and NXDs.
      </p>
      <p>
        R. P. Bourret summarizes and yet reconciles not only basic problems and
principles of native XML databases in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In his article R. P. Bourret says: “...
the problem is practical, not teoretical ...” about the problem of arbitrary data
expressed in any data model. He also states “... in RDB there is an impractical
number of joins ...” in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Another resource concluding the benefits of NXD usage [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] tries to list not
only the advantages but also the possible problems.
      </p>
      <p>
        We will very shortly summarize here XML database technologies and in the
next section we will cover formal models for technological standards and data
models treated here. Three typical NXD representants are as follows: (1) One
of the most common NXD’s is eXist [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. (2) Another very popular NXD from
Apache is called Xindice [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ]. Oracle Berkeley XML DB is described at [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ].
      </p>
      <p>
        XML formal models and languages from the point of view of XML standard
are at least as follows. We could say the following list is an extension of XML
Data Models according to R. P. Bourret [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The majority of all treated models
are tree-based formal models and algebras.
      </p>
      <p>
        – DOM — Document Object Model [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ].
– SAX — Simple API for XML [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
– Infoset [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
– XPath [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
– XQuery [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
– XML-λ: functional approach to XML description [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
– Most likely there are many other standard-based formalisms.
      </p>
      <p>Next section reveals the formal background of stated standards and needed
relationships.
5</p>
    </sec>
    <sec id="sec-5">
      <title>XML Databases Formal Models</title>
      <p>
        A Formal Data Model and Algebra for XML [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is the name of the article
suggesting a tree-based model as a formal data model for XML and as an algebra
for XML, an algebra based on such trees, i.e. essentially same structure as DOM
and the related.
      </p>
      <p>
        XML Data Model as it is defined in XPath or XQuery is basically grasped
as a forrest of trees of nodes representing elements and attributes and texts
and other XML features mentioned in previous section. Many of the XML Data
Model facets are explained in XML infoset [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>As an XML data model could be grasped DOM. Basically the formalism is
build on same terms as in case of XPath or XQuery. So from the CT point of
view it would be grasped as one formalism.</p>
      <p>
        Sengupta and Mohan summarized in [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] the formalisms used to describe
data in XML format. They found these formalisms:
– Tree-based formalisms (XAlgebra, DOM and others).
– SAL — Semi-structured Algebra [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
– The ENF — Element Normal Form concept: it is proved that attributes can
be avoided in cases of general description because every XML document with
attributes can be (without any information loss) transformed onto the XML
document variant without attributes and vice versa.
– HNR — Heterogeneous Nested Relations — also arise from NF2 (Non-first
normal form).
– HNRC — HNR Calculus — analogously to relational calculus.
– HNRA — HNR Algebra — analogously to relational algebra.
– DSQL — Document SQL — as an analogy to SQL.
      </p>
      <p>For all of these we would like to find the proper meta-formal way of
description in terms of CT; and finally find out the properties valid among these
categories.</p>
      <p>
        XML Algebra based on monads is another interesting formalism [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. But
this approach, this XML Algebra, lacks references and dereferences. The algebra
specified in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] count on it and offer a way of how to solve this problem.
      </p>
      <p>
        P. Wadler proposed several formal models. Especially formal semantics for
XSL [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ], and semantics for XPath [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ].
      </p>
      <p>Future challenges would be to describe formalisms used for metamodels—
conceptual models and visualisations e.g. via UML.</p>
      <p>Having in mind the extent of all this we will focus on just few factors from
the previous list in CT. We introduce CT in the next section.
6</p>
    </sec>
    <sec id="sec-6">
      <title>The Category Theory Standpoint</title>
      <p>This section deals with an introduction to CT and categorical description of
XML formal models defined above. Let us take a look at the word category
itself.
6.1</p>
      <sec id="sec-6-1">
        <title>Three semantics of the word Category</title>
        <p>
          Categories originally arose in mathematics out of the need of a formalism to
describe the transformation from one type of mathematical structure to another.
Category represents a kind of mathematics. Barr and Wells [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] state then
category as a mathematical workspace.
        </p>
        <p>A category is also a mathematical structure. It is then a generalization of
both ordered sets and monoids. Barr and Wells call it in this case category as a
mathematical structure.</p>
        <p>Category as a theory is the third recognized point of view. Category can
be seen as a structure that formalizes a mathematican’s description of a type
of structure. Traditional way to do this in mathematics, in mathematical logic
respectively, is to use formal languages with rules, terms, axioms and equations.</p>
        <p>
          We now define the term category more precisely. We will use the notation
and mathematical formalism used in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] for the rest of this section.
6.2
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>Definition of a Category</title>
        <p>Definition 1. A category C consists of objects (denoted by A, B, C, ...) and
morphisms between them (denoted by f : A → B, g : B → C, ...). These data
are subject to obvious axioms expressing composition, its associativity, and
existence of identity morphisms (units w.r.t. composition).</p>
        <p>
          A paradigm category is the category Set of all sets and mappings. See [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] for
more details.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>Definition of CCC, Connections to λ-Calculus</title>
        <p>
          We define now the concept of a cartesian closed category (CCC). It is proved
in [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] that CCC’s are essentially the same thing as simply typed λ-calculus.
Altough the following definition is rather a technical, one may bear in mind that
the category Set forms a paradigm example of a CCC.
        </p>
        <p>Definition 2. A category C is called a cartesian closed category (CCC) if it
satisfies the following:
(1) There is a terminal object 1.
(2) Each pair of objects A and B of C has a product A × B with projections
p1 : A × B → A and p2 : A × B → B.</p>
        <p>(3) For every pair of objects A and B, there is an object [A → B] and an arrow
eval : [A → B] × A → B with the property that for any arrow f : C × A → B,
there is a unique arrow λf : C → [A → B] such that the composite
λf × A eval</p>
        <p>C × A −−−−−→ [A → B] × A −→ B
is f .</p>
        <p>Note that we call an object 1 of a category C terminal iff there is exactly one
arrow A → 1 for each object A of C.</p>
        <p>
          λ-calculus is one of the formal description of what is usually called an XML
data model. In [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], there is XML-λ approach to view XML. The typical point
of view of XML is a tree. But this approach emphasizes the notion of a function.
And there is a hypothesis that as functions or as trees we describe the same,
and that both kinds of description are of the same power. We will try to prove
this in future work. This proof will rely on what is stated above.
        </p>
        <p>
          Related works involving Object Databases description using CT are [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and
[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. Altough it is not about the XML data model the principles of formal
description are very similar.
        </p>
        <p>
          Because of the λ-calculus is one of the formal description or precise point of
view of XML data model and because of what Lambek and Scott proved in their
work [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], the XML data model can be described as CCC. We would like to find
out if also other categories as descriptions of other formal models are also CCC.
The main idea is to determine if all the models are also essentially the same in
the sense of Lambek and Scott; which should be done in next work.
        </p>
      </sec>
      <sec id="sec-6-4">
        <title>Proposed Descriptions based on Category Theory</title>
        <p>
          The very first description we considered was the XML-λ approach. This approach
is an instance of the λ-calculus theory which, grasped as a category, is CCC [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ].
        </p>
        <p>
          Let G be a graph. As a graph we mean special case of oriented graph with
loops on nodes. The category CGraph of such graphs is defined as follows:
Collection of objects consists of all possible graphs G; Collection of arrows consists
of all graph homomorphisms φG. Identity arrows are isomorphisms of objects.
It is needed to be verified, that this mathematical structure is a category, but it
is obvious; we let this to the kind reader. Furthermore this category is CCC, as
is proved e.g. in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. This model, category CGraph, is actually a useful model for
object databases [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] when other aspects than object visibility are ignored.
Objects in this category can be grasped as objects from object programming. But
as a model for XML databases it cannot be used because of the loops. When the
arrow is interpreted as a relation of nesting, element in XML document cannot
be nested into itself.
        </p>
        <p>Let CT ree be the category of trees (derived from the category above). Let
objects be trees and arrows tree homomorphisms. Again that it is a category is
needed to be verified as above. This category is not CCC. Because there would
needed to exist the terminal object with loop node. But such an object cannot
be interpreted as any XML document.</p>
        <p>Let CHF S be the category of hereditary finite sets. All these sets can be
undrestood as -trees. This approach seems to be very promising and is currently
under development.</p>
        <p>There are many other approaches which will be in detail covered in
subsequent works.
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Conclusions</title>
      <p>We have shown that the database technology selection is in praxis mostly
subjective problem. There are few practical reasons which would lead us to develop
theoretical framework for data modeling.</p>
      <p>We have stressed the natural evolution in software engineering from waterfall
to iterative database evolution approaches which still become more common.</p>
      <p>We have discussed XML formal models and their properties.</p>
      <p>The conclusions from CT applications are rather poor. But we tried to
summarize the database problems, existing solutions and we tried to offer another,
originial, approach.</p>
      <p>Furthermore this work open the doors for further more specific research, using
very strong mathematical background. Next section reveals our future plans.
8</p>
    </sec>
    <sec id="sec-8">
      <title>Future Works</title>
      <p>Subsequent work will be focused on the question of essentiality of the RDB, ODB
and XDB models, their computational equivalence, expressive power of relative
languages and similar aspects.</p>
      <p>In the near future we will try to categorify every formal model for XML data
which would be found.</p>
      <p>
        In far future there is a huge space for using CT formalism to describe itself,
i.e. use the notion of categories of categories. And according to Lambek and
Scott [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] it seems to be possible use only one notation, one language and one
formalism, we mean CT of course, for all (types of) data models. We would like
to try to find out such a way of description of data models.
      </p>
      <p>
        Next, in the future, not only XML databases and NXD will be described
using CT. But we would like to try to give formal basis for all data models. Good
example could be relational algebra and Crole’s way of categorical description
which should be further elaborated [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; using categorical semantics. And there
are many other similar examples as an inspiration for future research activities.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S. W.</given-names>
            <surname>Ambler</surname>
          </string-name>
          . Mapping Objects to Relational Databases: O/R Mapping In Detail.
          <year>2006</year>
          . http://www.agiledata.org/essays/mappingObjects.html.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Barr</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Wells</surname>
          </string-name>
          .
          <article-title>Category Theory for Computing Science</article-title>
          . International Series in Computer Science. Prentice-Hall,
          <year>1990</year>
          . Second edition,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.</given-names>
            <surname>Beech</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Malhotra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Rys</surname>
          </string-name>
          .
          <article-title>A formal data model</article-title>
          and
          <source>algebra for XML</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Beeri</surname>
          </string-name>
          and
          <string-name>
            <surname>Y. Tzaban. SAL:</surname>
          </string-name>
          <article-title>An algebra for semistructured data and XML</article-title>
          .
          <source>In WebDB (Informal Proceedings)</source>
          , pages
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Bourret</surname>
          </string-name>
          . Mapping DTDs to databases,
          <year>2001</year>
          . http://www.xml.com/lpt/a/2001/05/09/dtdtodbs.html.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Bourret</surname>
          </string-name>
          .
          <article-title>Going native: Making the case for XML databases</article-title>
          ,
          <year>2005</year>
          . http://www.xml.com/pub/a/2005/03/30/native.html.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Bourret</surname>
          </string-name>
          . XML and databases,
          <year>2005</year>
          . http://www.rpbourret.com/xml/XMLAndDatabases.htm.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Bourret</surname>
          </string-name>
          . XML database products,
          <year>2007</year>
          . http://www.rpbourret.com/xml/XMLDatabaseProds.htm.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Chamberlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Berglund</surname>
          </string-name>
          , and
          <string-name>
            <given-names>e. a. Scott</given-names>
            <surname>Boag. XML Path</surname>
          </string-name>
          <article-title>Language (XPath) 2</article-title>
          .0,
          <year>September 2005</year>
          . http://www.w3.org/TR/xpath20/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>E. F.</given-names>
            <surname>Codd</surname>
          </string-name>
          .
          <article-title>The relational model for database management: version 2</article-title>
          . AddisonWesley Longman Publishing Co., Inc., Boston, MA, USA,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>J.</given-names>
            <surname>Cowan</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Tobin</surname>
          </string-name>
          .
          <article-title>XML information set (second edition</article-title>
          ),
          <year>April 2004</year>
          . http://www.w3.org/TR/2004/REC-xml-infoset-
          <volume>20040204</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>R.</given-names>
            <surname>Crole</surname>
          </string-name>
          . Categories for Types. Cambridge Mathematical Textbooks. Cambridge University Press,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>C. J. Date</surname>
          </string-name>
          .
          <article-title>An Introduction to Database Systems</article-title>
          . Addison-Wesley Publishing Co., Inc.,
          <year>2003</year>
          . 8th ed.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <fpage>db4objects</fpage>
          - Open Source ODBMS. http://www.db4o.com.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sparx</surname>
          </string-name>
          <article-title>'s Systems Enterprise Architect UML CASE Tool</article-title>
          . http://www.sparxsystems.com.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>16. eXist - Open Source Native XML Database, Home Page. http://exist.sourceforge.net.</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Extensible Markup</surname>
          </string-name>
          <article-title>Language (XML) 1.0 (Fourth Edition</article-title>
          ),
          <year>2006</year>
          . http://www.w3.org/XML.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. M.
          <article-title>Fern´andez, A</article-title>
          . Malhotra,
          <string-name>
            <given-names>J.</given-names>
            <surname>Marsh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nagy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <source>XQuery 1.0 and XPath 2</source>
          .
          <article-title>0 Data Model</article-title>
          ,
          <year>September 2005</year>
          . http://www.w3.org/TR/xpath-datamodel/.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>M. Fernandez</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Simeon</surname>
            , and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Wadler</surname>
          </string-name>
          .
          <article-title>A semi-monad for semi-structured data</article-title>
          .
          <source>Lecture Notes in Computer Science</source>
          ,
          <year>1973</year>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Fussel</surname>
          </string-name>
          .
          <article-title>Foundations of Object Relational Mapping</article-title>
          . http://www.chimu.com.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>21. GemStone/S ODB. http://www.gemstone.com/products/smalltalk.</mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>J. Gu</surname>
          </string-name>
          <article-title>¨ttner. Object Databases and the Semantic Web</article-title>
          .
          <source>PhD thesis</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>E. R.</given-names>
            <surname>Harold</surname>
          </string-name>
          .
          <article-title>Managing XML data: Native XML databases</article-title>
          ,
          <year>2005</year>
          . http://www.ibm.com/developerworks/xml/library/x-mxd4.
          <fpage>html</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <article-title>Hibernate - Java and</article-title>
          .
          <article-title>NET persistence framework</article-title>
          . http://www.hibernate.org.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25. P.
          <article-title>Kolenˇc´ık. Categorical Framework for Object-Oriented Database Model</article-title>
          .
          <source>PhD thesis</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>J.</given-names>
            <surname>Lambek</surname>
          </string-name>
          and
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Scott</surname>
          </string-name>
          .
          <article-title>Introduction to Higher-Order Categorical Logic</article-title>
          . Cambridge University Press,
          <year>March 1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>P.</given-names>
            <surname>Loupal</surname>
          </string-name>
          .
          <article-title>Querying XML with lambda calculi</article-title>
          .
          <source>In Ph.D. Workshop, VLDB2006</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <given-names>D.</given-names>
            <surname>Megginson. SAX - Simple</surname>
          </string-name>
          <string-name>
            <surname>API</surname>
          </string-name>
          for XML,
          <year>2005</year>
          . http://www.saxproject.org/.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>29. NeoDatis ODB. http://wiki.neodatis.org.</mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Object Management Group (OMG). MDA - Model Driven Architecture</surname>
          </string-name>
          ,
          <year>2007</year>
          . http://www.omg.org/mda.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Object Management Group (OMG). UML - Unified Modeling Language</surname>
          </string-name>
          ,
          <year>2007</year>
          . http://www.uml.org.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>ODBMS - Object And Object Oriented Database Management Systems</surname>
          </string-name>
          . http://www.odbms.org.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <article-title>Oracle Berkeley XML DB, home page</article-title>
          . http://www.oracle.com/database/berkeley-db/xml/index.html.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>34. SEN - Sigsoft Software Engineering Notes. http://www.sigsoft.org/SEN/surfing.html.</mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <given-names>A.</given-names>
            <surname>Sengupta</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Mohan</surname>
          </string-name>
          .
          <article-title>Formal and conceptual models for xml structures - the past, present</article-title>
          , and future,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>36. SIGSOFT - ACM's Special Interest Group, dedicated to Software Engineering. http://www.sigsoft.org.</mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <given-names>S.</given-names>
            <surname>Staken</surname>
          </string-name>
          . Introduction to Native XML Databases.
          <year>2001</year>
          . http://www.xml.com/pub/a/2001/10/31/nativexmldb.html.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Sun</surname>
            <given-names>Microsystems</given-names>
          </string-name>
          , Inc.
          <article-title>Java architecture for XML binding (JAXB</article-title>
          ),
          <year>2003</year>
          . http://java.sun.com/webservices/jaxb/.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>SWEBOK - Software Engineering</surname>
          </string-name>
          Body Of Knowledge. http://www.swebok.org.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <article-title>The W3C Consortium</article-title>
          .
          <source>Document Object Model (DOM)</source>
          ,
          <year>2005</year>
          . http://www.w3.org/DOM/.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <given-names>D.</given-names>
            <surname>Toth</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Loupal</surname>
          </string-name>
          .
          <article-title>Metrics analysis for relevant database technology selection</article-title>
          .
          <source>In Objekty</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          42.
          <string-name>
            <given-names>D.</given-names>
            <surname>Toth</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Valenta</surname>
          </string-name>
          .
          <article-title>Using Object And Object-Oriented Technologies for XML-native Database Systems</article-title>
          . In J. Pokorny´,
          <string-name>
            <surname>V.</surname>
          </string-name>
          <article-title>Sn´aˇsel, and</article-title>
          K. Richta, editors,
          <source>DATESO, CEUR Workshop Proceedings. CEUR-WS.org</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          43.
          <string-name>
            <given-names>P.</given-names>
            <surname>Wadler</surname>
          </string-name>
          .
          <article-title>A formal model of pattern matching in XSL</article-title>
          .
          <source>Technical report</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <given-names>P.</given-names>
            <surname>Wadler</surname>
          </string-name>
          . Two semantics for xpath,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45.
          <string-name>
            <surname>Apache</surname>
            <given-names>Xindice</given-names>
          </string-name>
          , Home Page. http://xml.apache.org/xindice/.
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          46.
          <string-name>
            <given-names>XML</given-names>
            <surname>Main</surname>
          </string-name>
          <article-title>Page</article-title>
          . http://www.w3.org/XML.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>