<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>AutoModelGen: A Generic Data Level Implementation of ModelGen</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrew Smith</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter McBrien</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computing, Imperial College London</institution>
          ,
          <addr-line>Exhibition Road, London SW7 2AZ</addr-line>
        </aff>
      </contrib-group>
      <fpage>65</fpage>
      <lpage>68</lpage>
      <abstract>
        <p>The model management operator ModelGen translates a schema expressed in one modelling language into an equivalent schema expressed in another modelling language, and in addition produces a mapping between those two schemas. AutoModelGen is a generic data level implementation of ModelGen that meets these desiderata. Our approach is distinctive in that (i) it takes a generic approach that can be applied to any modelling language, and (ii) it does not rely on knowing the modelling language in which the source schema is expressed in.</p>
      </abstract>
      <kwd-group>
        <kwd>ModelGen</kwd>
        <kwd>Model Management</kwd>
        <kwd>Data Transformation</kwd>
        <kwd>Data Integration</kwd>
        <kwd>Meta Modelling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        ModelGen is a model management [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] operator that translates a schema in a
source data modelling language (DML), for example XML Schema, into
an equivalent schema in a target DML, for example SQL, and also generates
a mapping between the two schemas. To date, no implementation of ModelGen
generates both a target schema and a mapping between the source and target
schemas [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this demonstration we present an implementation of ModelGen
that automatically creates a data level mapping that describes how instances
of the source schema should be translated [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Further distinguishing features
of our approach are that (1) the translations are made on a Universal Meta
Model (UMM) that has previously been shown to able to represent schemas
from a large number of data modelling languages, and (2) the mappings created
are bidirectional i.e. we also create a mapping from the target to the source
schema.
      </p>
      <p>
        Fig. 1 gives an overview of our approach. In step (1) the source schema Ss is
transformed into an equivalent schema, Shdm¡s expressed in the UMM. In step
(2), a series of information preserving [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] transformations are applied to Shdm¡s
to transform it into Shdm¡t, that matches the structure of a schema in the target
DML. In step (3) the constructs in Shdm¡t are transformed into their equivalents
in the target DML to create St. We will ¯rst discuss the overall architecture of
our system and then discuss the details of the two algorithms used in step (2).
AutoModelGen is a tool that creates schemas and both-as-view (BAV)
transformations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] in the AutoMed data integration system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. AutoMed
allows for schemas to be stored in both the native modelling language of a
data source (eg XML or SQL/relational) and in AutoMed's UMM called the
hypergraph data model (HDM) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        The HDM uses three modellings constructs (nodes, edges and constraints)
to represent the constructs of a high level DML [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. HDM nodes and edges have
associated data values (called their extent). Constraints place restrictions on
the data values that may appear in the extent. Each variant of a high level DML
construct has a particular representation in the HDM. For example, the set of
HDM constraints generated by a nullable SQL column will be di®erent those
generated by a not null column. This is important when it comes to identifying
whether a group of HDM constructs matches a particular construct in the target
DML.
      </p>
      <p>
        A BAV information preserving [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] mapping is made up of a sequence of
transformations called a pathway, where each transformation either adds, deletes or
renames a single schema object (such as a single SQL column, SQL primary key
de¯nition, XML element, etc), thereby incrementally generating a new schema
from an old schema. The extent of the schema object being added or deleted is
de¯ned as a query on the extents of the existing schema objects.
      </p>
      <p>BAV transformations can be grouped into information preserving
composite transformations (CT), that act as templates of a fragment of a
pathway, describing common patterns of transformation steps. For example the CT
id node expand is useful when the target DML has explicit keys (such as a key
attribute in ER or SQL models) but the source model has implicit keys (such as
in XML Schema).
3</p>
    </sec>
    <sec id="sec-2">
      <title>Algorithms</title>
      <p>
        AutoMatch inspects a given HDM schema Shdm¡x and determines which of
the nodes, edges and constraints match a construct in the target DML.
AutoTransform searches for a schema in which all the HDM schema objects
match the structure of the target DML by repeatedly applying CTs to the schema
objects in Shdm¡x that AutoMatch identi¯es as not matching constructs in the
target DML. The set of possible schemas created in this way is called the world
space [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] of the problem. It can be represented as a graph whose nodes are
individual HDM schemas and whose edges are the CTs needed to get from one node
in the graph to the next. To limit the number of possible CTs that have to be
performed at each node of the world space graph, CTs must satisfy certain
preconditions before they can be executed. In our algorithm the preconditions rely
on the structure of the HDM schema, in particular the constraints, surrounding
the schema object that the CT is to be applied to.
      </p>
      <p>The world space graph for the example in the demonstration is shown in
Fig. 2. Each node is labelled with a schema name (S; S0; : : :) and a list of the
unidenti¯ed schema objects e1,e2,. . . in that schema, or the word Solution. All
the constructs in a Solution node match those of the target model. Above each
node is a list of CTs that meet the preconditions for the unidenti¯ed schema
objects in that schema. Those CTs that meet the preconditions most closely are
sorted to the top of the list and executed ¯rst. AutoTransform performs a
depth ¯rst search on the world space graph starting from the initial state, by
executing the CT at the head of the list, until a solution or a dead end is reached.
If a dead end is reached the algorithm back tracks to the last node in the world
space graph where an untried CT/edge combination exists, and executes the
next CT in the list for that node. If all the edges on all the nodes have been
tried without ¯nding Solution then AutoTransform has failed.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Execution of the Tool</title>
      <p>The current prototype of the tool is capable of translating between schemas
represented in the XML, ER and SQL DMLs, and of materialising the data
instances of a schema in one DML as instances of a second DML. In a typical
execution of the tool the following steps are performed:
1. A Source schema is imported into AutoMed, and then translated into the</p>
      <p>HDM.
2. The AutoMatch and AutoTransform algorithms are run on the newly
generated HDM schema, hdm, to generate a new schema hdm0, where the
hdm0 schema is one that matches the structure of the target DML.
3. The hdm0 schema is translated into a schema in target DML.
4. This target schema along with its data is then materialised.
Since the result of the tool's output is a set of schemas and BAV mappings
held in AutoMed, the standard AutoMedtoolkit may be used to view results
of the tool's execution. Fig. 3 shows a screen shot from the AutoMed GUI
featuring an ER schema and then in an anti-clockwise direction, the SQL and
XML equivalents of the schema generated automatically by AutoModelGen.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halevy</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pottinger</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A vision of management of complex models</article-title>
          .
          <source>SIGMOD Record</source>
          <volume>29</volume>
          (
          <issue>4</issue>
          ) (
          <year>2000</year>
          )
          <volume>55</volume>
          {
          <fpage>63</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melnik</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Model management 2.0: manipulating richer mappings</article-title>
          .
          <source>In: SIGMOD Conference</source>
          .
          <article-title>(</article-title>
          <year>2007</year>
          )
          <volume>1</volume>
          {
          <fpage>12</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBrien</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A generic data level implementation of modelgen</article-title>
          . In: BNCOD. (
          <year>2008</year>
          ) To appear
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hull</surname>
          </string-name>
          , R.:
          <article-title>Relative information capacity of simple relational database schemata</article-title>
          .
          <source>SIAM J. Comput</source>
          .
          <volume>15</volume>
          (
          <issue>3</issue>
          ) (
          <year>1986</year>
          )
          <volume>856</volume>
          {
          <fpage>886</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>McBrien</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poulovassilis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Data integration by bi-directional schema transformation rules</article-title>
          .
          <source>In: ICDE</source>
          . (
          <year>2003</year>
          )
          <volume>227</volume>
          {
          <fpage>238</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Boyd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kittivoravitkul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lazanitis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.J.</given-names>
            <surname>McBrien</surname>
          </string-name>
          and
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Rizopoulos: AutoMed: A BAV Data Integration System for Heterogeneous Data Sources</article-title>
          .
          <source>In: CAiSE04</source>
          . Volume
          <volume>3084</volume>
          of LNCS., Springer Verlag (
          <year>2004</year>
          )
          <volume>82</volume>
          {
          <fpage>97</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>McBrien</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poulovassilis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A general formal framework for schema transformation</article-title>
          .
          <source>In: Data and Knowledge Engineering</source>
          . Volume
          <volume>28</volume>
          . (
          <year>1998</year>
          )
          <volume>47</volume>
          {
          <fpage>71</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Boyd</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McBrien</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Comparing and transforming between data models via an intermediate hypergraph data model</article-title>
          .
          <source>J. Data Semantics IV</source>
          (
          <year>2005</year>
          )
          <volume>69</volume>
          {
          <fpage>109</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Weld</surname>
            ,
            <given-names>D.S.:</given-names>
          </string-name>
          <article-title>An introduction to least commitment planning</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>15</volume>
          (
          <issue>4</issue>
          ) (
          <year>1994</year>
          )
          <volume>27</volume>
          {
          <fpage>61</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>