<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Capturing Interactive Data Transformation Operations using Provenance Work ows</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tope Omitola</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andre Freitas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edward Curry</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sean O'Riain</string-name>
          <email>sean.oriaing@deri.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicholas Gibbins</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nigel Shadbolt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Enterprise Research Institute (DERI) National University of Ireland</institution>
          ,
          <addr-line>Galway</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Web and Internet Science (WAIS) Research Group School of Electronics and Computer Science University of Southampton</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The ready availability of data is leading to the increased opportunity of their re-use for new applications and for analyses. Most of these data are not necessarily in the format users want, are usually heterogeneous, and highly dynamic, and this necessitates data transformation e orts to re-purpose them. Interactive data transformation (IDT) tools are becoming easily available to lower these barriers to data transformation e orts. This paper describes a principled way to capture data lineage of interactive data transformation processes. We provide a formal model of IDT, its mapping to a provenance representation, and its implementation and validation on Google Re ne. Provision of the data transformation process sequences allows assessment of data quality and ensures portability between IDT and other data transformation platforms. The proposed model showed a high level of coverage against a set of requirements used for evaluating systems that provide provenance management solutions.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>Public Open Data</kwd>
        <kwd>Data Publication</kwd>
        <kwd>Data Consumption</kwd>
        <kwd>Semantic Web</kwd>
        <kwd>Work ow</kwd>
        <kwd>Extract-Transform-Load</kwd>
        <kwd>Provenance</kwd>
        <kwd>Interactive Data Transformation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The growing availability of data on the Web and in organizations brings the
opportunity to reuse existing data to feed new applications or analyses. In order
to reuse existing data, users must perform data transformations to repurpose
data for new requirements. Traditionally, data transformation operations have
been supported by data transformation scripts organized inside ETL
(ExtractTransform-Load) environments or by ad-hoc applications. Currently, users
developing data transformation programs follow a typical software development
cycle, taking data samples from datasets, and developing the transformation
logic, testing and debugging. These approaches are problematic in emerging
scenarios such as the Linked Data Web where the reduction of the barriers for
producing and consuming new data brings increasing challenges in coping with
heterogeneous, high-volume, and dynamic data sources. In this scenario, the
process of interactive data transformation emerges as a solution to scale data
transformation.</p>
      <p>Recently platforms, such as Google Re ne3, are exploring user interface and
interaction paradigms for de ning data transformations. These platforms allow
users to operate over data using Graphical User Interface (GUI) elements instead
of scripts for data transformation. By providing a set of pre-de ned operations
and instant feedback mechanism for each iteration of the data transformation
process, this model de nes a powerful approach to data transformation and
curation. However, despite its practical success, the assumptions behind platforms
such as Google Re ne have not been modeled nor formalized in the literature.
As the need for curation increases in data abundant environments, the need to
model, understand, and improve data curation and transformation processes and
tools emerges as a critical problem. A foundational model of IDT that (a) brings
out the underlying assumptions that IDT platforms use, (b) makes explicit how
IDT relates to provenance, and (c) helps IDT platforms be comparable and
thereby helping in de ning interfaces for their interoperability, would be highly
useful.</p>
      <p>The process of data transformation is usually positioned as one element in
a connected pipeline which needs to be integrated with di erent processes, to
form a work ow. Currently, the set of transformation operations present in
IDT platforms are not materialized in a way that could allow its use in contexts
outside the IDT environment. A transformation process normally involves a
data object that is being transformed, and the transformation procedure itself.
Provenance is the contextual metadata describing the origin or source of that
data. Prospective provenance provides mechanisms to describe generic work ows
which could be used to materialize (future) data transformations. Retrospective
provenance captures past work ow execution and data derivation information to
provide useful context for what had happened up to the present state and time.</p>
      <p>This paper investigates the two complementary perspectives described above,
and an approach for expressively capturing and persisting provenance in IDT
platforms. A provenance extension for the Google Re ne platform is
implemented and used to validate the proposed solution. The supporting provenance
model focuses on the maximization of interoperability, using the three-layered
data transformation model proposed in [?] and uses Semantic Web standards to
persist provenance data.</p>
      <p>The contributions of this work include: (a) a formal model of IDT; (b) the
use of an ontology-based provenance model for mapping data transformation
operations on IDT platforms; (c) the veri cation of the suitability of the proposed
provenance model for capturing IDT transformations, using Google Re ne as the
IDT platform; and (d) the validation of the system against the set of
require</p>
    </sec>
    <sec id="sec-2">
      <title>3 http://code.google.com/p/google-re ne/</title>
      <p>Prototypical User Actions
Data Curation Application Operation selection</p>
      <p>KBwoforokpfelorawtisons/ Selection Parameter selection
(extensional semantics)</p>
      <p>Selection
Operation
composition</p>
      <p>Storage
Prospective provenance</p>
      <p>Materialization</p>
      <p>Data</p>
      <p>Schema-level Data
Instance-level Data</p>
      <p>Input
Data Transformation
Generation Cycle
Data Input
(Re)Definition
of User Actions
Transformation
Execution
Data</p>
      <p>Output
Inductive assessment
of user intention</p>
      <p>Output
Data Transformation Program</p>
      <p>Parameter α Operation A
Parameter β</p>
      <p>...</p>
      <p>Parameter ω</p>
      <p>Operation B</p>
      <p>...</p>
      <p>Operation Z</p>
      <p>Retrospective provenance</p>
      <p>Materialization
Output</p>
      <p>Data Specific Transformation</p>
      <p>ODuatptaut trDawantsoafroskrpfmelocawitfiiocn
ments, from the literature, used for evaluating provenance management provided
by IDT platforms.</p>
      <p>Interactive Data Transformations (IDT)</p>
      <sec id="sec-2-1">
        <title>IDT Operations Overview</title>
        <p>IDT is de ned as the application of a pre-de ned set of data transformation
operations over a dataset. In IDT, after a transformation operation has been
selected from the set of available operations (operation selection), users usually
need the input of con guration parameters (parameter selection). In addition,
users can compose di erent set of operations to create a data transformation
work ow (operation composition). Operation selection, parameter selection,
and operation composition are the core user actions available for interactive data
transformation as depicted in gure ??.</p>
        <p>In the IDT model, the expected outputs are a data transformation program
and a transformed data output ( gure ??). The transformation program is
generated through an iterative data transformation program generation cycle where
the user selects and composes an initial set of operations, executes the
transformation, and assesses the suitability of the output results by an inductive analysis
over the materialized output. This is the point where users decide to rede ne
(reselect) the set of transformations and con guration parameters. The
organization of operations as GUI elements minimizes program construction time by
minimizing the overhead introduced by the need to ensure programming
language correctness and compilation/deployment.</p>
        <p>The transformation generation cycle generates two types of output: a data
output with a data speci c transformation work ow, and a data transformation
program which is materialized as a prospective provenance descriptor. A
provenance descriptor is a data structure showing the relationship between the
inputs, the outputs, and the transformation operations applied in the process.
While the provenance-aware data output is made available for further changes,
the prospective provenance work ow is inserted into the KB of available
workows, and can be later reused.</p>
        <p>In the next section, we shall present an algebra for Interactive Data
Transformation.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Foundations - An Algebra for Provenance-based Interactive</title>
      </sec>
      <sec id="sec-2-3">
        <title>Data Transformation (IDT)</title>
        <p>Despite the practical success of IDT tools, such as Google Re ne, for data
transformation and curation, the underlying assumptions and the operational
behaviour behind such platforms have not been explicitly brought out. Here, we
present an algebra of IDT, bringing out the relationships between the inputs,
the outputs, and the functions facilitating the data transformations.</p>
      </sec>
      <sec id="sec-2-4">
        <title>De nition 1: Provenance-based Interactive Data Transformation En</title>
        <p>gine G A provenance-based Interactive Data Transformation Engine, G,
consists of a set of transformations (or activities) on a set of datasets generating
outputs in the form of other datasets or events which may trigger further
transformations.</p>
        <p>G is de ned as a tuple,</p>
        <p>G = &lt; D; (D [ V ); I; O; ; ;
&gt;
where
1. D is the non-empty set of all datasets in G,
2. D is the dataset being currently transformed,
3. V is the set of views in G (V may be empty),
4. I is a nite set of input channels (this represents the points at which user
interactions start),
5. O is a nite set of output channels (this represents the points at which user
interactions may end),
6. is a nite alphabet of actions (this represents the set of transformations
provided by the data transformation engine),
7. is a nite set of functions that allocate alphabets to channels (this
represents all user interactions), and
8. = &lt; D O ! &gt; is a function, where, in a modal transformation engine,
(D; O) 2 (O) is the dataset that is the output on channel O when D is
the dataset being currently transformed.</p>
      </sec>
      <sec id="sec-2-5">
        <title>De nition 2: Interactive Data Transformation Event An Interactive Data</title>
        <p>Transformation Event is a tuple,</p>
        <p>PT E = &lt; Di; Ftrans; (Do [ V ); Ttrans &gt;
where
{ Di is the input dataset for the transformation event,
{ Do is the dataset that is the result of the transformation event,
{ V is a view or facet that is a result of the transformation event,
{ Di [ Do [ V D
{ Ftrans is the transformation function applied to Di (applied element-wise),
and
{ Ttrans is the time the transformation took place.</p>
        <p>De nition 3: Run A run can be informally de ned as a function from time
to dataset(s) and the transformation applied to those dataset(s). Intuitively, a
run is a description of how G has evolved over time.</p>
        <p>So, a run over G is a function, P : t ! &lt; D; f &gt; where t is an element in the
time domain, D is the state of all datasets and views in G, and f is the function
applied at time, t.</p>
        <p>A system R over G, i.e. R(G) , is a set of all runs over G. We say that
&lt; P; t &gt; is a point in system R if P 2 R.</p>
        <p>P captures our notion of \prospective provenance".</p>
        <p>De nition 4: Trace Let =&lt; P; t &gt;2 R(G). The trace of , denoted by, !,
is the sequence of pairs &lt; ri; ti &gt; where ri is the i th run at time ti. The set
of all traces of G, T r(G), is the set f ! j 2 R(G) g. An element of T r(G) is a
trace of G. A trace captures our notion of retrospective provenance.
3
3.1</p>
        <p>Provenance-based Data Transformation</p>
      </sec>
      <sec id="sec-2-6">
        <title>Provenance Model</title>
        <p>Community e orts towards the convergence into a common provenance model led
to the Open Provenance Model (OPM)4. OPM descriptions allow
interoperability on the level of work ow structure. This model allows systems with di erent
provenance representations to share at least a work ow-level semantics. OPM,
however, is not intended to be a complete provenance model, demanding the
complementary use of additional provenance models in order to enable uses of
provenance which requires higher level of semantic interoperability. This work
targets interoperable provenance representations of IDT and ETL work ows
using a three-layered approach to represent provenance. In this representation, the
bottom layer represents the basic work ow semantics and structure provided
by OPM, the second layer extends the work ow structure provided by OPM</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://openprovenance.org/</title>
      <p>with Cogs[?], a provenance vocabulary that provides a rich type structure for
describing ETL transformations in the provenance work ow, and voidp [?], a
provenance extension for the void5 vocabulary, that allows data publishers to
specify the provenance relationships of the elements of their datasets. The third
layer consists of a domain speci c schema-level information of the source and
target datasets or classes/instances pointing to speci c elements in the ETL
process. In our implementation, the third layer contains the mapping of instances
to source code elements. Figure ?? shows how we applied the model (and
architecture) to Google Re ne, a popular IDT platform.
3.2</p>
      <sec id="sec-3-1">
        <title>Provenance Capture</title>
        <p>There are two major approaches for representing provenance information, and
these representations have implications on their cost of recording. These two
approaches are: (a) The (Manual) Annotation method: Metadata of the derivation
history of a data are collected as annotation. Here, provenance is pre-computed
and readily usable as metadata, and (b) The Inversion method: This uses the
relationships between the input data, the process used to transform and to derive
the output data, giving the records of this trace.</p>
        <p>The Annotation method is coarser-grained and more suitable for
slowlychanging transformation procedures. For more highly-dynamic and time-sensitive
transformations, such as IDT procedures, the Inversion method is more suitable.
We map the provenance data to the three-layered data transformation model
provenance model as described in section ??.</p>
        <p>After choosing the representation mechanism, the next questions to ask are:
(a) what data transformation points would generate the provenance data salient
to our provenance needs, and (b) what is the minimal unit of a dataset to attach
provenance to. For our system, we choose an Interactive Data Transformation
Event to capture these two questions and this is made up of the data object</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5 http://vocab.deri.ie/void/guide</title>
      <p>being transformed, the transformation operation being applied to the data, the
data output as a result of the transformation, and the time of the operation (as
stated in section ??).</p>
      <p>User</p>
      <p>User Interactions
User</p>
      <p>Queries
Queries</p>
      <p>Provenance Graph
Provenance
Storage
Layer</p>
      <p>Data
Transformation</p>
      <p>Engine
Provenance
Event Capture</p>
      <p>Layer
Provenance
Representation</p>
      <p>Layer</p>
      <p>Tt Di Do Ft</p>
      <p>Interceptor
Data Transformation</p>
      <p>Engine</p>
      <p>Google Refine
Keys:
Tt: Time of Operation
Di: Input Data or Dataset
Do: Output Data of Transformation
Ft: Applied Transformation Operation</p>
      <p>Tt Di Do Ft</p>
      <p>Object-toJavaClass
Introspector
Event-to</p>
      <p>RDF
Statements
Mapper
Provenance
Data in RDF
Provenance
Storage
(A) Process Flow</p>
      <p>(B) Provenance Event Capture Sequence Flow
1. The Provenance Event Capture Layer, which consists of the following layers
and operations' sequence ( gure ??(B)): (i) The Interceptor Layer: Here,
user interactions are intercepted and provenance events extracted from these
interactions; (ii) Object-to-JavaClass Introspector: The inputs to this layer
are the Transformation Operation chosen to transform the data. We employ
the Java language re ection mechanism to elicit the full class path of the
operation performed. Eliciting the full class path is useful for the following
reasons: (a) It allows us to have a pointer to the binary of the programs doing
the transformation, and (b) This pointer to the program binary allows the
connection between the full semantics of the program and the data layer. The
outputs of this layer are the full class paths of the transformation operations;
(iii) Event-to-RDF Statements Mapper: it receives the provenance event and
is responsible for the generation of the RDF predicates.
2. These events are then sent to the Provenance Representation Layer,
which encodes the captured events into RDF using the provenance
representation described in section ??.
3. These events, represented in RDF, are then sent to the Provenance
Storage Layer, which stores them in its Knowledge Base (KB).
4</p>
      <sec id="sec-4-1">
        <title>Google Re ne: An Exemplar Data Transformation/Curation System</title>
        <p>Google Re ne6 (GRe ne) is an exemplar Interactive Data Transformation
system, and some of its mechanisms include: (i) ability to import data into GRe ne
from di erent formats including tsv, csv, xls, xml, json, and google spreadsheets
; (ii) GRe ne supports faceted browsing7, such as: Text Facets, Numeric Facets,
and Text Filters; (iii) Editing Cells, Columns, and Rows, using GRe ne's editing
functions; and (iv) The provision of an extension framework API.
4.1</p>
        <sec id="sec-4-1-1">
          <title>Making Google Re ne Provenance-Aware</title>
          <p>Some of the design decisions that must be made when making an application
provenance-aware is to ask \what" type of provenance data to capture, \when"
to collect the said provenance data, and \what" type of method to use for the
capture.</p>
          <p>As regards to \when" to collect provenance data: (a) provenance data can
be collected in real-time, i.e. while the work ow application is running and the
input dataset(s) are being processed and used, or (b) ex-post (after-the-fact),
i.e. provenance data is gathered after a series of processing events or a sequence
of activities has completed.</p>
          <p>As regards to \what" type of method to use for collection, provenance data
collection methods fall into three types: (a) Through \User annotation": A
human data entry activity where users enter textual annotations, capturing, and
describing the data transformation process. User-centered metadata is often
incomplete and inconsistent [?]. This approach imposes a low burden on the
application, but a high burden on the humans responsible for annotation; (b) An
automated provenance instrumentation tool can be provided that is inserted into
the work ow application to collect provenance data. This places a low burden
on the user, but a higher burden on the application in terms of process cycles
and/or memory, and (c) Hybrid method: This method uses an existing
mechanism, such as a logging or an auditing tool, within the work ow application to
collect provenance data.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6 http://code.google.com/p/google-re ne/</title>
    </sec>
    <sec id="sec-6">
      <title>7 http://code.google.com/p/google-re ne/wiki/FacetedBrowsingArchitecture</title>
      <p>We built an automated instrumentation tool to capture provenance data from
user operations as they use the system ( gure ??). This approach incurs very
little burden on the user.
4.2</p>
      <sec id="sec-6-1">
        <title>Usage Scenarios and Data Transformation Experimentation using Google Re ne</title>
        <p>Here, we describe how we have used GRe ne to capture data transformation
provenance events, how these events have been represented using our provenance
models (described in section ??), and how we have made use of the IDT algebra
(in section ??). We converted the contents of FilmAwardsForBestActress8 into
a GRe ne project called \actresses".</p>
        <p>We have applied our de nition of a Run (as stated in section ??) to actual
implementations in the Usage Scenarios described below. Also here, we see the
Google Re ne system as an implementation of an IDT, G (from De nition 1 of
section ??).</p>
        <p>Edit Usage Scenario If we want to change the entry in row 2 (of gure ??) from
* "'1955 [[Meena Kumari]] "[[Parineeta (1953 lm)|Parineeta]]""' as "'Lolita"
to \1935 John Wayne" and would like our system to keep a provenance record
of this transaction, we can achieve that, in GRe ne, by viewing the column as a
\Text Facet" and applying the GRe ne's \Edit" operation on row 2. Figure ??
shows us the before and after pictures of our Edit operation.
The Events-to-RDF Statements Mapper automatically generates the RDF
statements using the following mechanisms. A transformation function, e.g. \Edit",
from GRe ne is automatically mapped to a type (an \rdf:type") of opmv:Process.
What the operation gets mapped to in the Cogs ontology depends on the
attribute of the data item that was the domain of the operation. For example, if
the data item attribute is a Column, this gets mapped to a Cogs
\ColumnOperation" class, while if the item attribute is a Row, this gets mapped to a
Cogs \RowOperation" class. Since this is a transformation process, we made
use of \cogs:TransformationProcess" class. voidp has a single class of
ProvenanceEvent and every transformation operation is mapped to
\voidp:ProvenanceEvent".</p>
        <p>The system's Object-to-JavaClass Introspector is used to elicit the actual
Java class responsible for the transformation. We have made use of a Cogs
property, \cogs:programUsed", to specify the full path to this Java class. It
allows the generated provenance data to communicate with the program
semantics, in this way the intensional program semantics is linked up with the
provenance extensional semantics. The data item that is actually made use of
in the transformation process is mapped to \opmv:Artifact". To give us the
process ow that is part of the transformation, we used two opmv properties: (a)
\opmv:wasDerivedFrom", which tells us from which artifact this present data
item is derived from, and (b) \opmv:wasGeneratedBy", which tells us from which
process this data item is generated from. In order to store temporal information
of when the derivation took place, we used \opmv:wasGeneratedAt", an opmv
property that tells us the time of transformation.</p>
        <p>The result of the transformation operation as RDF statements is below:
@prefix id: &lt;http://127.0.0.1:3333/project/1402144365904/&gt; .
id:MassCellChange-1092380975 rdf:type opmv:Process,
cogs:ColumnOperation, cogs:TransformationProcess, voidp:ProvenanceEvent ;
opmv:used &lt;http://127.0.0.1:3333/project/1402144365904/</p>
        <p>MassCellChange-1092380975/1_0&gt; ;
cogs:operationName "MassCellChange"^^xsd:string;
cogs:programUsed "com.google.refine.operations.cell.</p>
        <p>MassEditOperation"^^xsd:string;
rdfs:label "Mass edit 1 cells in column ==List of winners=="^^xsd:string.
&lt;http://127.0.0.1:3333/project/1402144365904/MassCellChange-1092380975/1_0&gt;
rdf:type opmv:Artifact ;
rdfs:label "* '''1955 [[Meena Kumari]]</p>
        <p>'[[Parineeta (1953 film)|Parineeta]]''''' as '''Lolita'''"^^xsd:string.
http://127.0.0.1:3333/project/1402144365904/MassCellChange-1092380975/1_1&gt;
rdf:type opmv:Artifact ;
rdfs:label "* '''John Wayne'''"^^xsd:string;
opmv:wasDerivedFrom &lt;http://127.0.0.1:3333/project/1402144365904/</p>
        <p>MassCellChange-1092380975/1_0&gt;;
opmv:wasGeneratedBy &lt;http://127.0.0.1:3333/project/1402144365904/</p>
        <p>
          MassCellChange-109238097
          <xref ref-type="bibr" rid="ref5">5&gt;;
opmv:wasGeneratedAt "2011</xref>
          -11-16T11:2:14"^xsd: dateTime.
        </p>
        <p>In the next section, we will describe the requirements we used for analysis.
5</p>
        <sec id="sec-6-1-1">
          <title>Analysis and Discussion</title>
          <p>Our approach and system were evaluated in relation to the set of requirements
listed in [?][?] and enumerated below:
1. Decentralization: deployable one database at a time, without requiring
co-operation among all databases at once. Coverage: High. Justi cation:
Use of Semantic Web Standards and vocabularies (OPMV + Cogs + voidp)
to reach a decentralized/interoperable provenance solution.
2. Data model independency: should work for data stored in at le,
relational, XML, le system, Web site, etc., model. Coverage: High. Justi
cation: Minimization of the interaction between the data level and the
provenance level. Data is connected with its provenance descriptor by a provenance
URI and the provenance representation is normalized as RDF/S.
3. Minimum impact to existing IDT practice: Provenance tracking is
invisible to the user. Coverage: High. Justi cation: The provenance capture is
transparent to the user. Provenance is captured by a lightweight
instrumentation of the IDT platform, mapping program structures to the provenance
model.
4. Scalability: to situations in which many databases cooperate to maintain
provenance chain. Coverage: High. Justi cation: Usage of Semantic Web
standards and vocabularies for provenance representation allows for the
cooperation of multiple platforms for provenance management.
6</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>Related Work</title>
          <p>We categorize related work into two categories: (i) provenance management for
manually curated data [?], and (ii) interactive data transformation models and
tools [?][?]. Buneman et al. [?] propose a model for recording provenance of data
in a single manually curated database. In their approach, the data is copied from
external data sources or modi ed within the target database creating a
copypaste model for describing user actions in assimilating external datasources into
curated database records.</p>
          <p>Raman and Hellerstein [?] describe Potters Wheel, an interactive data
cleaning system which allowed users to specify transformations through graphic
elements. Similarly, Wrangler [?] is an interactive tool based on the visual speci
cation of data transformation. Both systems are similar to Google Re ne: [?],
however, provides a more principled analysis of the interactive data transformation
process, while [?] focuses on new techniques for specifying data transformations.
The IDT and provenance model proposed in this work can be directly applied
to both systems.
7</p>
        </sec>
        <sec id="sec-6-1-3">
          <title>Conclusion</title>
          <p>The world contains an unimaginably vast amount of data which is getting ever
vaster ever more rapidly. These opportunities demand data transformation
efforts for data analysis and re-purposing. Interactive data transformation (IDT)
tools are becoming easily available to lower barriers to data transformation
challenges. Some of these challenges will be solved by developing mechanisms useful
for capturing the process ow and data lineage of these transformation processes
in an interoperable manner. In this paper, we provide a formal model of IDT, a
description of the design and architecture of an ontology-based provenance
system used for mapping data transformation operations, and its implementation
and validation on a popular IDT platform, Google Re ne. We shall be making
our provenance extension to Google Re ne publicly accessible very soon.
8</p>
        </sec>
        <sec id="sec-6-1-4">
          <title>Acknowledgements</title>
          <p>This work was supported by the EnAKTing project, funded by EPSRC project
number EP/G008493/1 and by the Science Foundation Ireland under Grant No.
SFI/08/CE/I1380 (Lion-2).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Buneman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cheney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vansummeren</surname>
          </string-name>
          ,
          <article-title>A Provenance Model for Manually Curated Data</article-title>
          ,
          <source>International Provenance and Annotation Workshop (IPAW)</source>
          , p.
          <fpage>162</fpage>
          -
          <lpage>170</lpage>
          ,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Kandel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Paepcke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heer</surname>
          </string-name>
          , Wrangler: Interactive Visual Speci - cation
          <source>of Data Transformation Scripts ACM Human Factors in Computing Systems (CHI)</source>
          ,
          <year>2011</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>V.</given-names>
            <surname>Raman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hellerstein</surname>
          </string-name>
          ,
          <article-title>Potter's Wheel: An Interactive Data Cleaning System</article-title>
          ,
          <source>In Proceedings of the 27th International Conference on Very Large Data Bases</source>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Freitas</surname>
          </string-name>
          , B. Kampgen,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. O</given-names>
            <surname>'Riain</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Curry: Representing Interoperable Provenance Descriptions for ETL Work ows</article-title>
          ,
          <source>In Proceedings of the 3rd International Workshop on Role of Semantic Web in Provenance Management (SWPM</source>
          <year>2012</year>
          ),
          <source>Extended Semantic Web Conference (ESWC)</source>
          , Heraklion, Crete,
          <year>2012</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>T.</given-names>
            <surname>Omitola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gutteridge</surname>
          </string-name>
          , I. Millard,
          <string-name>
            <given-names>H.</given-names>
            <surname>Glaser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gibbins</surname>
          </string-name>
          , and
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Shadbolt: Tracing the Provenance of Linked Data using voiD</article-title>
          .
          <source>In The International Conference on Web Intelligence, Mining and Semantics (WIMS'11)</source>
          ,
          <year>2011</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>E.</given-names>
            <surname>Deelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shields</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Taylor: Work ows and e-Science: An overview of work ow system features and capabilities</article-title>
          .
          <source>In Future Generation Computer Systems</source>
          ,
          <volume>25</volume>
          (
          <issue>5</issue>
          ) pp.
          <fpage>528</fpage>
          -
          <lpage>540</lpage>
          ,
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Newhouse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Schopf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Richards</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Atkinson: Study of User Priorities for e-Infrastructure for e-Research (SUPER)</article-title>
          .
          <source>In UK e-Science Technical Report Series Report UKeS-2007-01</source>
          ,
          <year>2007</year>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>