<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IWSG</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards Traceability in Data Ecosystems using a Bill of Materials Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Iain Barclay, Alun Preece, Ian Taylor</string-name>
          <email>BarclayIS@cardiff.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dinesh Verma</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Crime and Security Research Institute, Cardiff University</institution>
          ,
          <addr-line>Cardiff</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IBM TJ Watson Research Center</institution>
          ,
          <addr-line>1110 Kitchawan Road, Yorktown Heights, NY 10598</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>12</volume>
      <fpage>12</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>-Researchers and scientists use aggregations of data from a diverse combination of sources, including partners, open data providers and commercial data suppliers. As the complexity of such data ecosystems increases, and in turn leads to the generation of new reusable assets, it becomes ever more difficult to track data usage, and to maintain a clear view on where data in a system has originated and makes onward contributions. Reliable traceability on data usage is needed for accountability, both in demonstrating the right to use data, and having assurance that the data is as it is claimed to be. Society is demanding more accountability in data-driven and artificial intelligence systems deployed and used commercially and in the public sector. This paper introduces the conceptual design of a model for data traceability based on a Bill of Materials scheme, widely used for supply chain traceability in manufacturing industries, and presents details of the architecture and implementation of a gateway built upon the model. Use of the gateway is illustrated through a case study, which demonstrates how data and artifacts used in an experiment would be defined and instantiated to achieve the desired traceability goals, and how blockchain technology can facilitate accurate recordings of transactions between contributors.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Scientists and researchers increasingly assemble and use
rich data ecosystems[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in their experimentation. As these
ecosystems expand in capability and leverage data from a
diverse combination of internal sources, partners and third
party data suppliers, it is becoming necessary for users and
curators of data to have reliable traceability on its origins and
uses. This can be important to provide accountability[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], such
as proving ownership or legitimate usage of the source data,
as well as being able to identify quality or supply problems
and alert users to problems or to seek redress when things go
awry.
      </p>
      <p>Using a gateway to provide traceability on data used within
experiments offers mechanisms for demonstrating where data
and assets derived from the data are used, as well as aiding
understanding where data contributing to a system has come
from. By coupling the traceability trail with distributed ledger
or blockchain technology, it is possible to provide a distributed
store that can record digital data or events in a way that makes
them immutable, non-repudiable and identifiable, thereby
leading to a trustworthy record of fact.</p>
      <p>Research into manufacturing, agricultural and food
industries, where the need for traceability of products and their
Copyright © 2021 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
component parts is well-established, has informed the design
and development of a gateway which enables data ecosystems
to be described in terms of sub-assemblies of their constituent
data components and supporting artifacts, in a Bill of Materials
(BoM) format. Artifacts in a BoM might include data licenses,
software descriptions and versions, and lists of staff or other
human resources involved in producing the outputs. When the
system described by the BoM is run, the BoM is instantiated,
queried for the locations of data sources and populated with
any dynamic values for the data or artifacts of each run,
generating a Bill of Lots (BoL). The BoM and BoL together
provide a record of the static and dynamic elements of the
system for an invocation at a particular point in time. This
allows for later inspection of the data and the supporting
environment, and provides a means for scientists to trace
data and artifact usage through and across experiments - for
example, identifying all uses of a particular IOT sensor, all
runs using a particular version of a machine learning model,
or all uses of data generated by a particular researcher.</p>
      <p>
        A pilot gateway, dataBoM, has been developed to allow
scientists to describe data ecosystem as a Bill of
Materials, containing pipelines of assemblies detailing sets of data
sources and artifacts, and to instantiate the BoM into a BoL
for each run of an experiment. The dataBoM gateway has
been developed using GraphQL[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which facilitates the rapid
development of cross-platform applications and web services
which scientists can use to generate and query BoMs and
populate and store BoL records. Integration of the dataBoM
gateway with blockchain or distributed ledger technologies
can provide dynamic behaviour in data acquisition, as well
as providing a permanent audit trail of both the data used and
its supporting environment.
      </p>
      <p>The remainder of this paper is structured as follows:
Section II discusses the context in which the BoM model for
data ecosystem traceability has been derived; the architecture
and implementation of the dataBoM gateway is discussed in
Section III, with Section IV describing a case study illustrating
how a scientist could use the pilot gateway to conduct research
using data from several sources to identify traffic congestion.
Section V considers areas for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>II. REQUIREMENTS</title>
      <p>
        In manufacturing industries it has been standard practice
since the late twentieth century to track product through the
life-cycle from its origin as raw materials, through component
assembly to finished goods in a store, with the
relationships and information flows between suppliers and customers
recorded and tracked using supply chain management (SCM)
processes[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In agri-food industries, traceability through the
supply chain is necessary to give visibility from a product
on a supermarket shelf, back to the farm and to the batch of
foodstuff, as well as to other products in which the same batch
has been used.
      </p>
      <p>Describing data ecosystems in terms of the data supply
chain provides a mechanism to identify data sources and the
assets which contribute to the development of the data
components, or which are produced as the results of intermediate
processes. As new assets are created and used in other systems
- perhaps by other parties - the supply chain mapping can be
extended to give traceability on the extended data ecosystem.</p>
      <p>
        A definition for traceability is provided by Opara[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], as
”the collection, documentation, maintenance, and application
of information related to all processes in the supply chain
in a manner that provides guarantee to the consumer and
other stakeholders on the origin, location and life history of a
product as well as assisting in crises management in the event
of a safety and quality breach.”
      </p>
      <p>
        Further helpful terminology is provided by Kelepouris,
Pramatari and Doukidis[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] when discussing the traceability of
information in terms of the direction of analysis of the supply
chain. Petroff and Hill[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] define Tracing as the ability to work
backwards from any point in the supply chain to find the origin
of a product (i.e., ‘where-from’ relationships) and Tracking
as the ability to work forwards, finding products made up
of given constituents (i.e., ‘where-used’ relationships). Thus,
an effective traceability solution should support both tracing
and tracking; providing effectiveness in one direction does not
necessary deliver effectiveness in the other[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Jansen-Vullers, van Dorp, and Beulens[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and van Dorp[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
discuss the composition of products in terms of a Bill of
Materials (BoM) and a Bill of Lots (BoL). The BoM is the
list of types of component needed to make a finished item of
a certain type, whereas the BoL lists the actual components
used to create an instance of the item. In other words, the
BoM might specify a sub-assembly to be used, and the BoL
would identify which exact batch the sub-assembly used in
the building of a particular instance of a product was part of.
Furthermore, a BoM can be multi-level, wherein components
can be used to create sub-assemblies which are subsequently
used in several different product types.
      </p>
      <p>The notion of using a BoM to identify and record
component parts of assets in an IT context is already established,
with US Department of Commerce working on the NTIA
Software Component Transparency initiative to provide a
standardised Software BoM1 format to detail the sub-components
1https://www.ntia.doc.gov/SoftwareTransparency
in software applications. The intent is to give visibility on
the underlying components used in software applications and
processes such that vulnerable out-of-date modules can easily
be identified and replaced. Tools such as CycloneDX2, SPDX3,
and SWID4 are defining formats for identifying and tracking
such sub-components.</p>
      <p>
        As well as the data used and efforts through standards[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
and research made to secure its provenance in workflows[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ],
there are many supporting assets which can be considered
useful supplementary information when recording the
characteristics of a data ecosystem, which Singh, Cobbe and Norval[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
have described as providing decision provenance. Hind, et
al, describe a document based on a Supplier’s Declaration of
Conformity[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] as a suitable vehicle for providing an overview
of an AI system, detailing the purpose, performance, safety,
security, and provenance characteristics of the overall system.
At the component level, Gebru et al explore the benefits of
developing and maintaining Datasheets for Data[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], which
replicates the specification documents that often accompany
physical components, and Mitchell et al propose a document
format for AI model specifications and benchmarks[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Schelter, Bo¨se, Kirschnick, Klein and Seufert[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] describe a system
to automatically document the parameters of machine learning
experiments by extracting and archiving the metadata from
the model generation process, which would be appropriate
information to store alongside the data used in a system.
      </p>
      <p>
        Members of the scientific community are familiar with
the use of workflow systems, such as Node-RED[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and
Pegasus WMS[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], to define and execute the processes for
their experiments. The BoM model proposed herein is intended
to augment a workflow by providing a means to add contextual
traceability as the workflow progresses, such that it can be
archived, and the supporting conditions retrieved and inspected
later. Workflow blocks typically describe a job or a service,
and do not allow other contributing artifacts to be described.
The proposed BoM model describes a rich set of information
per node, which can better represent the data supply chain and
associated documents and payloads that are contained at each
stage. By maintaining a BoM model alongside a workflow,
researchers can populate and capture a record of the data for
each run, as well as the supporting artifacts for each run,
giving traceability of the data and the circumstances in which
it was obtained and used. In practical terms, a function could
be written to populate the BoL with dynamic data, and invoked
at appropriate points in the workflow.
      </p>
      <p>
        Distributed ledger technologies, such as those afforded by
blockchain platforms[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], provide a means of recording
information and transactions between parties who do not have
formal trust relationships[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], such as inter-organisational or
commercial data sharing entities. The design of a blockchain
system ensures that data written cannot be changed,
providing a level of immutability and non-repudiation which
2https://cyclonedx.org
3https://spdx.org”
4https://www.iso.org/standard/65666.html
is well suited to keeping an auditable record of events and
transactions which occur between parties. Furthermore, the
use of a public blockchain platform, such as the Ethereum
Project[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], provides an archival resource which remains in
existence long after the resources of a project have been
retired. State-of-the-art blockchain platforms, including the
Ethereum Project, allow for the deployment of so-called
smart contracts, which can be considered to be “autonomous
agents stored in the blockchain, encoded as part of a creation
transaction that introduces a contract to the blockchain”[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
Such smart contracts enable blockchain platforms to facilitate
non-repudiable dynamic behaviours alongside their immutable
storage capabilities.
      </p>
    </sec>
    <sec id="sec-3">
      <title>III. A DATA TRACEABILITY GATEWAY</title>
      <p>In this section the design and implementation of dataBoM,
a gateway capable of supporting levels of tracking and tracing
appropriate for providing traceability in multi-party
decentralised data ecosystems, is described. The solution uses a
model based on a Bill of Materials scheme, where data and
supporting materials are treated as constituent components of
a deployed system, which is instantiated into a unique Bill of
Lots each time the deployment is run.</p>
      <sec id="sec-3-1">
        <title>A. Conceptual Model</title>
        <p>The dataBoM gateway employs a BoM model, such that
each experiment utilising the system is described in terms
of its data supply chain. The BoM consists of a collection
of assemblies, with each assembly being an aggregation of
contributing input components and an output component.</p>
        <p>An assembly will typically have at least one data input,
and can produce new data as its output. Data output from
one assembly can be used as a data input in a subsequent
assembly within the current BoM, or used in other systems
by being referenced in their BoM. To reflect this, data inputs
and outputs are defined as data sources.</p>
        <p>Assemblies can also contain artifacts, which are pertinent
software components, ML models, and documentation such
as licenses, staff lists, policy documentation, etc.. Including
artifacts in assemblies in the BoM definition ensures that each
BoL retains a full record of its heritage and dependencies.</p>
        <p>An assembly can produce a new artifact as its output; for
example, an assembly which described the training of an AI
model would produce the trained model as its output. The
trained model would then be considered an artifact, which
could be used as an input to other assemblies.</p>
        <p>Figure 1 shows two assemblies that are chained to produce
a data component (Data 1’) and an artifact (Artifact 2’)
as outputs. Such a BoM could be used by a scientist to
describe a simple AI model training process containing two
assemblies. Assembly 1 represents the data labelling process,
and Assembly 2 the model training process. Data 1 is an input
data source, which could be training data. Artifact 1 might be
a roster of the staff employed to label the data, and the central
data source, Data 1’ (which, as illustrated, is both the output of
the data labelling assembly and the input to the model training
assembly) could be a labelled data set. In the second assembly,
Artifact 2 would be relevant to the model training process, for
example the parameters used in training. The output artifact,
Artifact 2’, would be the trained model. Note that both the
intermediate output, Data 1’ and the final output, Artifact 2’
could be further used as inputs by other processes and specified
as inputs to subsequent assemblies.</p>
        <p>The BoM defines a map of the structure of the system
by providing a record of the connections between the
assemblies, and provides a framework to enumerate a system’s
data sources and artifacts as well as any static data that
applies to the contained data sources or artifacts. This static
information could include a location for access to the data,
for example, a Digital Object Identifier (DOI) or an API
URL, and metadata specifying acceptable data threshold levels
or response requirements for active quality of service (QoS)
monitoring.</p>
        <p>Each time the process described by the BoM is run, the
application code for the process will instantiate a new BoL
for the given BoM. In order to provide on-going traceability,
a shadow data item is created for each data source and artifact
in the BoM when it is instantiated in a BoL. The shadow
items in the BoL are used to maintain a record of the dynamic
elements of each run.</p>
        <p>By storing and then later referencing the assemblies, data
sources and artifacts in a BoM, and all the instantiations of the
BoM in each BoL, along with the shadow data, it is possible
to derive an overview of the history of the data lifecycle of the
system, such that any item can be traced back to its origins
or tracked forward to find all its consumers.</p>
        <p>One of the roles for the data source elements specified
in the BoM is to store the means to access the data when
the experiment is run. In many cases this will be via a url
parameterised dynamically at runtime - the static entities of the
url could be stored in the data source as part of the BoM, with
the dynamic parameters and the results stored in the shadow
data item of the BoL. The intent of the design is that there is
flexibility of type, so any metadata could be stored in the BoM
and retrieved and interpreted in the application process. Uses
of this metadata could include storing encrypted information,
which is unencrypted and subsequently used by the client
application. Further, the metadata could include information
to initiate an asynchronous data request and an endpoint to
which the data should be delivered. The intent is to provide a
flexible storage slot for static data about the data, which can
be retrieved, interpreted and used by the client application.
Experimentally, it has been possible to use the dataBoM
gateway pilot to store and retrieve an encoded blockchain
contract address and function interface from a data source,
and use this information to initiate a blockchain transaction
from the client application to retrieve data at runtime. Such a
transaction could be used to provide immutable proof of a data
request, or for gateway users to have a means to access
thirdparty data on a pay-per-use basis, which is discussed further
in Section VI.</p>
      </sec>
      <sec id="sec-3-2">
        <title>B. The dataBoM Gateway</title>
        <p>
          The dataBoM gateway provides a working implementation
of the conceptual data ecosystem BoM model[
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] and enables
researchers to declare BoMs to describe the data components
of their experiments, and instantiate BoLs to preserve
contextual records for each run to provide traceability.
        </p>
        <p>The architecture of the dataBoM gateway is shown in
Figure 2. The gateway is to be offered as a web service, with
interactions between the gateway and researchers conducted
through a web interface or via an API.</p>
        <p>
          The pilot version of the dataBoM gateway stores data in
a MongoDB5 database, such that queries can be written to
provide traceability on data sourcing and data use for any
BoM. Further development of the gateway will explore the
off-loading of the archival of the BoMs and BoLs to
commonsbased decentralised storage, such as IPFS[
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], with indexing
secured on a public ledger or blockchain. This will serve
to preserve records beyond the lifetime of the gateway, and
provide an immutable record of events, suitable for later audit
or inspection.
        </p>
        <p>
          The dataBoM gateway is initially hosted on an intranet,
and it is envisaged that future versions of the gateway will
be migrated to public facing web services, or serverless[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]
environments, such as AWS Appsync6, to provide a robust and
reliable service.
        </p>
        <p>5https://www.mongodb.com
6https://aws.amazon.com/appsync/</p>
        <p>The gateway server is written in Node.js7, using Apollo
GraphQL Server8, which acts as an abstraction layer above
the gateway’s Mongo DB database store.</p>
        <p>GraphQL allows developers to specify a data schema, and
define queries and mutations, which are interfaces to allow
reading and writing of the data, respectively. The GraphQL
data schema, queries and mutations are public interfaces,
which hide the details of the underlying data storage from
users of the interfaces. The server’s data store does not have
to match the GraphQL schema, as the server code which
implements the queries and mutations performs the mapping
to read and write the correct data to its database. GraphQL
is intended to provide an efficient transfer of data between
client and server, as queries can be written to request only
the data needed. Furthermore, the gateway’s API can be
enhanced by extending the queries and mutations offered,
without implications for existing users.</p>
        <p>The GraphQL interface is self-documenting, and can be
queried by client application developers to find out the data
structures and queries and mutations available to them.</p>
        <p>The dataBoM gateway offers access to its GraphQL server
via an https end-point for API access.</p>
      </sec>
      <sec id="sec-3-3">
        <title>C. Integration with Client Applications</title>
        <p>To take advantage of the traceability capabilities provided
by the dataBoM gateway, scientists should use the supplied
API to define a BoM for their experiments, detailing the
assemblies, data sources and artifacts required in their processes,
passing the desired parameters and retaining the identifiers
which are returned by the API calls in order to chain entities
together - for example, when creating a data source item, the
identifier that is returned should be retained so that it can be
used as a parameter when creating an assembly.</p>
        <p>Once the BoM is defined, the researcher should instantiate
the BoM whenever they run their experiment, and then use
the API from their application code to query the experiment’s
BoM for static factors such as the locations of data assets, with
any dynamic state arising during experimentation (eg. data
values) being written to the BoL via the API as the experiment
progresses.</p>
        <p>Use of the API requires the researcher to integrate a
GraphQL client library with their application code or workflow
scripts, and support is available for popular web and mobile
platforms, including Python, Node.js, iOS and Android.</p>
        <p>The steps in the integration would typically include:</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Define data sources, artifacts and assemblies in BoM</title>
      <p>Use BoM’s ID to instantiate a new BoL for a new run
Access data source metadata for data location or endpoint
On receipt of data, populate data source shadow in BoL
In this way, the BoM and the BoL can combine to
generate an evidence trail of the dynamic data values and the
static components of the data and supporting artifacts which
contributed to each run of an experiment.</p>
      <p>7https://nodejs.org/en/
8https://www.apollographql.com/docs/apollo-server/
Section IV, below, describes a case study implementation, to
provide further insight and explanation of dataBoM integration
and usage.</p>
    </sec>
    <sec id="sec-5">
      <title>IV. CASE STUDY</title>
      <p>By way of illustration of the use of the dataBoM gateway,
consider a simple software application which serves to provide
a ‘traffic congestion score’ for a fixed location, e.g., Hyde
Park Corner, depending on how much traffic the application
determines is currently at the location. This simple process has
a single assembly, Traffic Scene Analysis, an input data source</p>
      <sec id="sec-5-1">
        <title>Location Photo, an ML model artifact Congestion Model and</title>
        <p>an output data source Congestion Score (Figure 3).
In defining the BoM for the Hyde Park Corner (HPC)
congestion rating process, the scientist should give each element a
name and an optional description, and declare static elements,
such as the URL to be used to retrieve a live photo from
the location of interest. Encoding this simple single assembly
process as a BoM through the gateways’s API gives a data
model as shown in Listing 1, which is the result of a GraphQL
query on the BoM’s entry.
”bom”: f
”name”: ”HPC Congestion”,
”description”: ”Determine congestion levels on Hyde Park Corner”,
”assemblies”: [
f
”name”: ”Traffic Scene Analysis”,
”description”: ”Determine congestion at Hyde Park Corner”,
”inputData”: [
f
”name”: ”Traffic Scene”,
”dataAccess”: ”https://xyz.com/00001.06514.jpg”
g
g
g
],
”outputData”: [</p>
        <p>f
],
”inputArtifacts”: [
f
”name”: ”Result”
”name”: ”Congestion Model”
g</p>
        <p>In the application code for the experiment, the BoM should
be instantiated via its identifier to generate a new BoL for
the run. As the code runs, it should refer to its BoM (via the
instantiated BoL) to get locations for data it needs to access,
and write any dynamic information to its BoL for permanent
archival.</p>
        <p>In the HPC congestion scoring example, the data source for
the traffic scene holds a static URL for a live camera. The
scientist’s code would retrieve this information through the
dataBoM API and access the photo, and (if desired) store a
permanent copy of the photo to its own archives, writing a
reference to the location of the archived copy to the shadow
data item, such that it will be saved as part of the archival of
the BoL. The resultant congestion score should also be written
to the BoL, by referencing the appropriate data source item.</p>
        <p>Thus, each data source and artifact in every BoL would
have any dynamic values recorded and stored in a database as
a persistent record of the run, so that each of the Assemblies
in the BoL would have traceable input and output data values
which could be accessed at a later date.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>V. DISCUSSION</title>
      <p>There are a number of interesting directions in which future
development of the dataBoM gateway could be taken.
Interaction with the gateway is currently provided by a GraphQL
API, which provides good integration with the application
code at runtime, however, initial definition of the BoM and
its elements would be more intuitive if it were faciliated
through a visual UI. Thus, the BoM could be authored using a
visual interface via a web browser, with the runtime invocation
and interaction with the BoL remaining an API-driven task.
There is a similar opportunity to add a visual interface to the
overview of each experiment logged by the gateway. Such an
interface would provide a means to explore the composition
of the data and artifact elements of each experiment, and help
to satisfy the traceability goals of the gateway, by providing
a convenient means of exploring the nodes in the BoM and
each BoL.</p>
      <p>Integration of the dataBoM gateway with the workflow
manager systems that are popular in the research community will
facilitate smoother integration of the gateway into experiment
workflows, and help to foster acceptance of the benefits of
the BoM model in providing traceability in scientific
dataecosystems.</p>
      <p>There is scope to extend and deepen the integration of
the gateway and its BoM and BoL models with blockchain
technologies, such as the programmable smart contracts
provided by the Ethereum blockchain platform. By associating
smart contracts with the data sources and artifacts from the
BoM model, novel dynamic behaviour in data ecosystems can
be explored. Such dynamic behaviours might include runtime
selection of the most appropriate data source sets, along with
automatic remuneration and sanctioning, based on dynamic
measures of data quality. Further development of the dataBoM
gateway could provide a means by which scientists are able
to share data and artifacts with their peers, and a blockchain
platform might underpin this. Related to blockchain integration
is motivation to explore traceability on the human side of the
experimental process, using Decentralised Identifiers9 (DIDs)
to associate researchers or crowd-workers with components of
the system and to provide a means to trace their activity and
the data and artifacts they are associated with.</p>
    </sec>
    <sec id="sec-7">
      <title>VI. CONCLUSION</title>
      <p>The dataBoM gateway provides scientists and developers
with a means to map the overall structure of the
components that make up complex data ecosystems used in their
experiments. By going beyond the data, and considering other
contributing factors such as the software and hardware which
produces or manages the data, licenses which govern the use
and sharing of the data, and policies which contributed to the
generation of the data, the development of a BoM for each
system provides a mechanism to archive the ecosystem for
each experiment. Instantiating the BoM into a BoL each time
the system runs augments the static parts list with a dynamic
and traceable view into every invocation of the system, such
that the data inputs, data outputs and any artifacts which
are used or produced by the system can be archived, readily
identified and traced back to their source. Similarly, future
users of produced data and artifacts, such as models, can be
identified, which could prove to be very important if errors
are later found and are notifiable. Storing metadata capable of
identifying smart contracts on the blockchain further enables
immutable recording of the action and timing of requests for
data provision, along with the potential for encoding quality
of service requirements, and providing automatic payment for
services.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENT</title>
      <p>This research was sponsored by the U.S. Army Research
Laboratory and the UK Ministry of Defence under Agreement
Number W911NF-16-3-0001. The views and conclusions
contained in this document are those of the authors and should
not be interpreted as representing the official policies, either
expressed or implied, of the U.S. Army Research Laboratory,
the U.S. Government, the UK Ministry of Defence or the UK
Government. The U.S. and UK Governments are authorized
to reproduce and distribute reprints for Government purposes
notwithstanding any copyright notation hereon.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. I. S.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. d. F. B.</given-names>
            <surname>Lima</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B. F.</given-names>
            <surname>Lo</surname>
          </string-name>
          <article-title>´scio, “Investigations into data ecosystems: a systematic mapping study</article-title>
          ,
          <source>” Knowledge and Information Systems</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>42</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Diakopoulos</surname>
          </string-name>
          , “
          <article-title>Accountability in algorithmic decision making,” Communications of the ACM</article-title>
          , vol.
          <volume>59</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>56</fpage>
          -
          <lpage>62</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Byron</surname>
          </string-name>
          , “
          <article-title>Graphql: A data query language</article-title>
          .” [Online]. Available: https://code.facebook.com/posts/1691455094417024/ graphql-a
          <article-title>-data-query-language</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Lambert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Cooper</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Pagh</surname>
          </string-name>
          , “
          <article-title>Supply chain management: implementation issues and research opportunities,” The international journal of logistics management</article-title>
          , vol.
          <volume>9</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L. U.</given-names>
            <surname>Opara</surname>
          </string-name>
          , “
          <article-title>Traceability in agriculture and food supply chain: a review of basic concepts, technological implications, and future prospects</article-title>
          ,
          <source>” Journal of Food Agriculture and Environment</source>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>106</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kelepouris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Pramatari</surname>
          </string-name>
          , and G. Doukidis, “
          <article-title>Rfid-enabled traceability in the food supply chain,” Industrial Management &amp; data systems</article-title>
          , vol.
          <volume>107</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>200</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. N.</given-names>
            <surname>Petroff</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Hill</surname>
          </string-name>
          , “
          <article-title>A framework for the design of lot-tracing systems for the 1990s,” Production and Inventory Management Journal</article-title>
          , vol.
          <volume>32</volume>
          , no.
          <issue>2</issue>
          , p.
          <fpage>55</fpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Jansen-Vullers</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. A. van Dorp</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Beulens</surname>
          </string-name>
          , “Managing traceability information in manufacture,”
          <source>International journal of information management</source>
          , vol.
          <volume>23</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>395</fpage>
          -
          <lpage>413</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>C. Van Dorp</surname>
          </string-name>
          ,
          <article-title>“A traceability application based on gozinto graphs,”</article-title>
          <source>in Proceedings of EFITA 2003 Conference</source>
          ,
          <year>2003</year>
          , pp.
          <fpage>280</fpage>
          -
          <lpage>285</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Missier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Belhajjame</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Cheney</surname>
          </string-name>
          , “
          <article-title>The w3c prov family of specifications for modelling provenance metadata</article-title>
          ,”
          <source>in Proceedings of the 16th International Conference on Extending Database Technology. ACM</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>773</fpage>
          -
          <lpage>776</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Davidson</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Freire</surname>
          </string-name>
          , “
          <article-title>Provenance and scientific workflows: challenges and opportunities</article-title>
          ,”
          <source>in Proceedings of the 2008 ACM SIGMOD international conference on Management of data. ACM</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>1345</fpage>
          -
          <lpage>1350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cobbe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Norval</surname>
          </string-name>
          , “
          <article-title>Decision provenance: Harnessing data flow for accountable systems</article-title>
          ,
          <source>” IEEE Access</source>
          , vol.
          <volume>7</volume>
          , pp.
          <fpage>6562</fpage>
          -
          <lpage>6574</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hind</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mehta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mojsilovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Nair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. N.</given-names>
            <surname>Ramamurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Olteanu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Varshney</surname>
          </string-name>
          , “
          <article-title>Increasing trust in ai services through supplier's declarations of conformity</article-title>
          ,” arXiv preprint arXiv:
          <year>1808</year>
          .07261,
          <year>2018</year>
          . [Online]. Available: https://arxiv.org/pdf/
          <year>1808</year>
          .07261.pdf
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Gebru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Morgenstern</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vecchione</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Vaughan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Daumee´ III, and</article-title>
          K. Crawford, “Datasheets for datasets,” arXiv preprint arXiv:
          <year>1803</year>
          .09010,
          <year>2018</year>
          . [Online]. Available: https: //arxiv.org/abs/
          <year>1803</year>
          .09010
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zaldivar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Vasserman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hutchinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Spitzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. D.</given-names>
            <surname>Raji</surname>
          </string-name>
          , and T. Gebru, “
          <article-title>Model cards for model reporting</article-title>
          ,”
          <source>in Proceedings of the Conference on Fairness, Accountability, and Transparency. ACM</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>220</fpage>
          -
          <lpage>229</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schelter</surname>
          </string-name>
          , J.-H. Bo¨se, J. Kirschnick,
          <string-name>
            <given-names>T.</given-names>
            <surname>Klein</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Seufert</surname>
          </string-name>
          , “
          <article-title>Automatically tracking metadata and provenance of machine learning experiments</article-title>
          <source>,” in Machine Learning Systems Workshop at NIPS</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17] “
          <article-title>Node-red: Flow-based programming for the internet of things</article-title>
          .” [Online]. Available: https://nodered.org/
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Deelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Vahi</surname>
          </string-name>
          , G. Juve,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rynge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Callaghan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Maechling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mayani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. F.</given-names>
            <surname>Da Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Livny</surname>
          </string-name>
          et al.,
          <article-title>“Pegasus, a workflow management system for science automation,” Future Generation Computer Systems</article-title>
          , vol.
          <volume>46</volume>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>35</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nakamoto</surname>
          </string-name>
          et al.,
          <article-title>“Bitcoin: A peer-to-peer electronic cash system</article-title>
          ,”
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>G.</given-names>
            <surname>Wood</surname>
          </string-name>
          , “
          <article-title>Ethereum: A secure decentralised generalised transaction ledger,” Ethereum project yellow paper</article-title>
          , vol.
          <volume>151</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>32</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tapscott</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Tapscott</surname>
          </string-name>
          , “
          <article-title>How blockchain will change organizations,” MIT Sloan Management Review</article-title>
          , vol.
          <volume>58</volume>
          , no.
          <issue>2</issue>
          , p.
          <fpage>10</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Luu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.-H.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Olickel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Saxena</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hobor</surname>
          </string-name>
          , “
          <article-title>Making smart contracts smarter</article-title>
          ,”
          <source>in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. ACM</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>254</fpage>
          -
          <lpage>269</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>I.</given-names>
            <surname>Barclay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Preece</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Taylor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Verma</surname>
          </string-name>
          , “
          <article-title>A conceptual architecture for contractual data sharing in a decentralised environment</article-title>
          ,” arXiv preprint arXiv:
          <year>1904</year>
          .03045,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>J.</given-names>
            <surname>Benet</surname>
          </string-name>
          , “
          <article-title>Ipfs-content addressed, versioned, p2p file system</article-title>
          ,
          <source>” arXiv preprint arXiv:1407.3561</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>E.</given-names>
            <surname>Jonas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schleier-Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sreekanti</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-C. Tsai</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Khandelwal</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Pu</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Shankar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Carreira</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Krauth</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Yadwadkar</surname>
          </string-name>
          et al.,
          <article-title>“Cloud programming simplified: A berkeley view on serverless computing</article-title>
          ,” arXiv preprint arXiv:
          <year>1902</year>
          .03383,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>